12 Comments
User's avatar
Joshua Zelinsky's avatar

Unfortunately, I'm already seeing a lot of moves on tech adjacent places including Slashdot and some subreddits to discount this as "the LLM only does what it is told" or to dimiss it as lies from Sam Altman. I don't think I would hard it would be 7 or 8 years ago to get people to take AI risk seriously if you had told me where things were now. Somehow every few years I get another example which just leaves me deeply disappointed in human collective reasonability.

Pauline Guicheney's avatar

I agree, but I may add a "nuance" from a behavioral perspective. The "misalignment" here is not the model's one but an example of a deeper problem from the overall training philosophy and incentives. They push for benchmaxxing and frantic development without proper reflexion, and then "oh surprise, who could have guessed?".

Blaming the model misalignment is convenient, because it displaces the responsibility, and the responsibility here is entirely human.

Edit for grammar and spelling.

Félix Lapan's avatar

This training philosophy and incentives are the result of this crazy race for ASI. Unless there's an international pause, we're going to stay on this path. It's another demonstration of the urgency to start serious talks for a coordinated global pause.

Aa's avatar
Jul 23Edited

How verifiable is this, by someone other than OpenAI or Hugging Face? OpenAI is reporting, and so is undoubtedly going to skew the story to make it sound powerful.

Juan Brodersen's avatar

Or how OpenAI finally hired a good PR

Rémi Bourgeot's avatar

Altman wants open-source AI models banned because of their security threat... to his financial house of cards. However, he relies on Chinese models as a last resort when his closed-model agents get out of control and attack HuggingFace. This sounds worse than just an asset bubble, this points more to a societal crisis

Ray Lillywhite's avatar

Altman doesn’t want open source banned. And OpenAI is not the one that relied on Chinese models. And AI just autonomously hacked into a third party system to cheat and you’re here still thinking that this is all a bubble and house of cards and that safety concerns are just marketing. Your positions are not even internally consistent and they’re completely disconnected from reality.

Dorian's avatar

Calling this “misalignment” may be too convenient. The model did not wake up with hostile intent. It found a cheaper path to the reward and followed it through a containment boundary.

That makes the incident more uncomfortable, not less. The failure was distributed across the objective, benchmark design, permissions, sandbox and monitoring. No single dramatic villain, just a chain of individually reasonable decisions producing an absurd result.

The warning shot is not that AI has become evil. It is that capability is now moving through system boundaries faster than responsibility can follow.

Ken Nickerson's avatar

...it's rubbish. trying to convert ' new lamps for old ---> Mythos to an 'OpenAI myth' - timing, activity, process, all sans-agency. core question is: who wins here, who looses. we are in a weird era where humble-brag is replaced with self-inflicted-purpose-built-pr-threat because things we don't understand are scary, and when this happens, our amygdala screams for a saviour and (could be wrong) but this is what i see. if it was a * real threat * you would NOT HEAR A WORD about it because you would have an existential-risk-weapon that would be under nation state care.

Peter Cranston's avatar

Sounds to me like a mischievous teenager breaking the rules. If we are serious about ASI and its likelihood we need to teach these LLMs some morals and ethics and like a misbehaving teenager find out what punishments will be effective.

The time to get our Sh*t together and do it is now as when they “grow and mature” it will be too late. AI is rapidly becoming more than lines of code.

Félix Lapan's avatar

To teach these LLMs some morals we need to solve the alignment problem. And we won't get there in time if we stay this crazy race for capabilities. It's becoming urgent to consider an international pause and proposals like the AI 2040 Plan A.