Perhaps the preferable structure is not assessment on behalf of labs, but insurers. As you note the fundamental problem you're trying to solve is internalizing negative externalities (expected costs of harms imposed on third-parties). But you have a principal-agent problem. And also an information asymmetry.
The Insurance Institute for Highway Safety falls into the broad IVO category, but is funded entirely by the insurance industry and not carmakers. There's still colorable criticism lodged against IIHS, but auto insurers directly benefit from accuracy and do not systematically benefit from shading assessments one way or the other.
When IIHS was founded the adjunct insurance industry was already fairly mature. A meaningful barrier in AI is that cybersecurity and other insurers I don't believe have fully adjusted to the new risk gradients and moreover standardization is probably a while off. I assume there'd also be some friction in the degree of access to models without lab cooperation. But the point remains insurers have no reason to underestimate or underprice risk. And the policy is for insurance users, not so much for developers.
A FT article this week said US cyber coverage went down in 2024 and so did premiums(!). That's not "AI risk" per se but the insurers quoted say it's inarguable AI-enabled risks are growing. So maybe not a great example. And of course terrorism risk insurance required a government reinsurance program (TRIA) to prevent a complete collapse in that market.
The CRA analogy nails the incentive problem. A second failure sits under it: even an auditor with no reason to go easy can only inspect a fraction of the model. Anthropic's own July 6 global-workspace paper puts the reportable part at less than a tenth of a model's activity, the rest running automatically and unreadable to its makers' best tools. An IVO certifies the slice it can see. The CRAs could at least read the whole loan tape.
Sorry but I cant get past the blanket statement "vaccines are safe" because looking at your reference for that statement, I see that the study you linked only looked at vaccines on the US immunization schedule up till the date of Novemeber 9. 2020. So your reference does not support your blanket statement. Vaccines are a medical product like any other, they do not possess magical qualities of inherent safety nor inherent qualities of harm, they must reviewed rigorously on a case-by-case basis. Perhaps you have forgotten about the AstraZeneca Covid-19 vaccine already?
Perhaps the preferable structure is not assessment on behalf of labs, but insurers. As you note the fundamental problem you're trying to solve is internalizing negative externalities (expected costs of harms imposed on third-parties). But you have a principal-agent problem. And also an information asymmetry.
The Insurance Institute for Highway Safety falls into the broad IVO category, but is funded entirely by the insurance industry and not carmakers. There's still colorable criticism lodged against IIHS, but auto insurers directly benefit from accuracy and do not systematically benefit from shading assessments one way or the other.
When IIHS was founded the adjunct insurance industry was already fairly mature. A meaningful barrier in AI is that cybersecurity and other insurers I don't believe have fully adjusted to the new risk gradients and moreover standardization is probably a while off. I assume there'd also be some friction in the degree of access to models without lab cooperation. But the point remains insurers have no reason to underestimate or underprice risk. And the policy is for insurance users, not so much for developers.
A FT article this week said US cyber coverage went down in 2024 and so did premiums(!). That's not "AI risk" per se but the insurers quoted say it's inarguable AI-enabled risks are growing. So maybe not a great example. And of course terrorism risk insurance required a government reinsurance program (TRIA) to prevent a complete collapse in that market.
Then I'll be looking forward to that piece! :)
The CRA analogy nails the incentive problem. A second failure sits under it: even an auditor with no reason to go easy can only inspect a fraction of the model. Anthropic's own July 6 global-workspace paper puts the reportable part at less than a tenth of a model's activity, the rest running automatically and unreadable to its makers' best tools. An IVO certifies the slice it can see. The CRAs could at least read the whole loan tape.
Okay, so there would need to be competitive incentives for more accurate safety assessments. But what could that look like?
I'm working on that piece now: happy to share what I have privately.
Sorry but I cant get past the blanket statement "vaccines are safe" because looking at your reference for that statement, I see that the study you linked only looked at vaccines on the US immunization schedule up till the date of Novemeber 9. 2020. So your reference does not support your blanket statement. Vaccines are a medical product like any other, they do not possess magical qualities of inherent safety nor inherent qualities of harm, they must reviewed rigorously on a case-by-case basis. Perhaps you have forgotten about the AstraZeneca Covid-19 vaccine already?
Apologies, please ping me to read your response tomorrow, but just on this comment: random selection implies no incentives to improve, right?