Discussion about this post

User's avatar
ST's avatar
Jul 22Edited

Perhaps the preferable structure is not assessment on behalf of labs, but insurers. As you note the fundamental problem you're trying to solve is internalizing negative externalities (expected costs of harms imposed on third-parties). But you have a principal-agent problem. And also an information asymmetry.

The Insurance Institute for Highway Safety falls into the broad IVO category, but is funded entirely by the insurance industry and not producer of evaluted product. And like AI, there's a consumer information asymmetry.

There's still colorable criticism lodged against IIHS, but auto insurers directly benefit from accuracy and do not benefit from systematically shading assessments one way or the other.

TBS when IIHS was founded the auto insurance industry was already fairly mature. A meaningful barrier in AI is that cybersecurity and other insurers have not fully adjusted to the new risk gradients and moreover standardization is probably a while off. Then again, could this be a sufficient condition for that happen?

I assume there'd also be some friction in the degree of access to models without lab cooperation. But the point remains insurers have no reason to underestimate risk.

A FT article this week said US cyber coverage went down in 2024 and so did premiums(!). That's not "AI risk" per se but the insurers quoted say it's inarguable AI-enabled risks are growing. So maybe not a great example. And of course terrorism risk insurance required a government reinsurance program (TRIA) to prevent a complete collapse in that market.

Inside The Black Box's avatar

The coercive-choices-stay-with-government split is the load-bearing claim, and last month tested it. Commerce pulled two frontier models on a single reported jailbreak. The lab called it "previously known, minor vulnerabilities"; the administration called it a refused safety request; the disagreement was never resolved in public. No IVO consensus would have been waited for there. Auditors also work because independence is backed by law and liability, and this design deliberately trades that legal shield for earned credibility. So what makes an administration honor a measurement it dislikes, when it already overrode the lab's own assessment?

No posts

Ready for more?