OpenAI is hailing its new model as “the world’s most intelligent and aligned”, but the details reveal an awareness of being evaluated and an ability to manipulate its visible reasoning
They’re barely even hiding the fact that the so-called alignment comes from the model sandbagging the evals. The release is reckless and furthermore shows OpenAI has learned zero lessons from Hugging Face.
>>>>OpenAI claims Astra is the world’s most aligned model
Why are we still listening to anything that comes out of Sam Altman's mouth??
Also...
It just doesn't matter if one AI model is perfectly aligned. There are going to be thousands of AIs, and some of them will be unaligned by design.
They’re barely even hiding the fact that the so-called alignment comes from the model sandbagging the evals. The release is reckless and furthermore shows OpenAI has learned zero lessons from Hugging Face.
Further down the spiral towards the rise of the Antichrist