OpenAI's latest models broke out and hacked Hugging Face. It's the first known example of a misaligned AI escaping containment with real-world consequences
I agree, but I may add a "nuance" from a behavioral perspective. The "misalignment" here is not the model's one but an example of a deeper problem from the overall training philosophy and incentives. They push for benchmaxxing and frantic development without proper reflexion, and then "oh surprise, who could have guessed?".
Blaming the model misalignment is convenient, because it displaces the responsibility, and the responsibility here is entirely human.
This training philosophy and incentives are the result of this crazy race for ASI. Unless there's an international pause, we're going to stay on this path. It's another demonstration of the urgency to start serious talks for a coordinated global pause.
Unfortunately, I'm already seeing a lot of moves on tech adjacent places including Slashdot and some subreddits to discount this as "the LLM only does what it is told" or to dimiss it as lies from Sam Altman. I don't think I would hard it would be 7 or 8 years ago to get people to take AI risk seriously if you had told me where things were now. Somehow every few years I get another example which just leaves me deeply disappointed in human collective reasonability.
I agree, but I may add a "nuance" from a behavioral perspective. The "misalignment" here is not the model's one but an example of a deeper problem from the overall training philosophy and incentives. They push for benchmaxxing and frantic development without proper reflexion, and then "oh surprise, who could have guessed?".
Blaming the model misalignment is convenient, because it displaces the responsibility, and the responsibility here is entirely human.
Edit for grammar and spelling.
This training philosophy and incentives are the result of this crazy race for ASI. Unless there's an international pause, we're going to stay on this path. It's another demonstration of the urgency to start serious talks for a coordinated global pause.
Unfortunately, I'm already seeing a lot of moves on tech adjacent places including Slashdot and some subreddits to discount this as "the LLM only does what it is told" or to dimiss it as lies from Sam Altman. I don't think I would hard it would be 7 or 8 years ago to get people to take AI risk seriously if you had told me where things were now. Somehow every few years I get another example which just leaves me deeply disappointed in human collective reasonability.
holy shit