Perhaps the preferable structure is not assessment on behalf of labs, but insurers. As you note the fundamental problem you're trying to solve is internalizing negative externalities (expected costs of harms imposed on third-parties). But you have a principal-agent problem. And also an information asymmetry.
The Insurance Institute for Highway Safety falls into the broad IVO category, but is funded entirely by the insurance industry and not carmakers. There's still colorable criticism lodged against IIHS, but auto insurers directly benefit from accuracy and do not systematically benefit from shading assessments one way or the other.
When IIHS was founded the adjunct insurance industry was already fairly mature. A meaningful barrier in AI is that cybersecurity and other insurers I don't believe have fully adjusted to the new risk gradients and moreover standardization is probably a while off. I assume there'd also be some friction in the degree of access to models without lab cooperation. But the point remains insurers have no reason to underestimate or underprice risk. And the policy is for insurance users, not so much for developers.
A FT article this week said US cyber coverage went down in 2024 and so did premiums(!). That's not "AI risk" per se but the insurers quoted say it's inarguable AI-enabled risks are growing. So maybe not a great example. And of course terrorism risk insurance required a government reinsurance program (TRIA) to prevent a complete collapse in that market.
I agree that that a marketplace where AI labs select and pay auditors will reduce quality and that auditors would not be incentivized to identify risks. There is good evidence that this is a problem. But that is an argument against letting labs choose auditors, not against independent verification. Random selection among qualified evaluators with pooled payment and regulator oversight can avoid conflicts of interest while creating incentives to improve the science of risk evaluation.
Random selection can provide incentives to improve. For example the frequency of random assignment can (and should) be modified based on past performance to incentivize auditors to innovate. This was part of the proposed CRA approach.
The CRA analogy nails the incentive problem. A second failure sits under it: even an auditor with no reason to go easy can only inspect a fraction of the model. Anthropic's own July 6 global-workspace paper puts the reportable part at less than a tenth of a model's activity, the rest running automatically and unreadable to its makers' best tools. An IVO certifies the slice it can see. The CRAs could at least read the whole loan tape.
Sorry but I cant get past the blanket statement "vaccines are safe" because looking at your reference for that statement, I see that the study you linked only looked at vaccines on the US immunization schedule up till the date of Novemeber 9. 2020. So your reference does not support your blanket statement. Vaccines are a medical product like any other, they do not possess magical qualities of inherent safety nor inherent qualities of harm, they must reviewed rigorously on a case-by-case basis. Perhaps you have forgotten about the AstraZeneca Covid-19 vaccine already?
Perhaps the preferable structure is not assessment on behalf of labs, but insurers. As you note the fundamental problem you're trying to solve is internalizing negative externalities (expected costs of harms imposed on third-parties). But you have a principal-agent problem. And also an information asymmetry.
The Insurance Institute for Highway Safety falls into the broad IVO category, but is funded entirely by the insurance industry and not carmakers. There's still colorable criticism lodged against IIHS, but auto insurers directly benefit from accuracy and do not systematically benefit from shading assessments one way or the other.
When IIHS was founded the adjunct insurance industry was already fairly mature. A meaningful barrier in AI is that cybersecurity and other insurers I don't believe have fully adjusted to the new risk gradients and moreover standardization is probably a while off. I assume there'd also be some friction in the degree of access to models without lab cooperation. But the point remains insurers have no reason to underestimate or underprice risk. And the policy is for insurance users, not so much for developers.
A FT article this week said US cyber coverage went down in 2024 and so did premiums(!). That's not "AI risk" per se but the insurers quoted say it's inarguable AI-enabled risks are growing. So maybe not a great example. And of course terrorism risk insurance required a government reinsurance program (TRIA) to prevent a complete collapse in that market.
I agree that that a marketplace where AI labs select and pay auditors will reduce quality and that auditors would not be incentivized to identify risks. There is good evidence that this is a problem. But that is an argument against letting labs choose auditors, not against independent verification. Random selection among qualified evaluators with pooled payment and regulator oversight can avoid conflicts of interest while creating incentives to improve the science of risk evaluation.
I wrote a response to this op-ed: https://ronbodkin904558.substack.com/p/how-to-make-independent-ai-audits
Apologies, please ping me to read your response tomorrow, but just on this comment: random selection implies no incentives to improve, right?
Random selection can provide incentives to improve. For example the frequency of random assignment can (and should) be modified based on past performance to incentivize auditors to innovate. This was part of the proposed CRA approach.
Then I'll be looking forward to that piece! :)
The CRA analogy nails the incentive problem. A second failure sits under it: even an auditor with no reason to go easy can only inspect a fraction of the model. Anthropic's own July 6 global-workspace paper puts the reportable part at less than a tenth of a model's activity, the rest running automatically and unreadable to its makers' best tools. An IVO certifies the slice it can see. The CRAs could at least read the whole loan tape.
Okay, so there would need to be competitive incentives for more accurate safety assessments. But what could that look like?
I'm working on that piece now: happy to share what I have privately.
Sorry but I cant get past the blanket statement "vaccines are safe" because looking at your reference for that statement, I see that the study you linked only looked at vaccines on the US immunization schedule up till the date of Novemeber 9. 2020. So your reference does not support your blanket statement. Vaccines are a medical product like any other, they do not possess magical qualities of inherent safety nor inherent qualities of harm, they must reviewed rigorously on a case-by-case basis. Perhaps you have forgotten about the AstraZeneca Covid-19 vaccine already?