Subscribe
Sign in
Home
Policy
Industry
Capabilities + Risks
Society
Opinion
Weekly Briefing
Campaign Finance Tracker
About
Capabilities + Risks
Latest
Top
Discussions
Internal AI deployments have people worried. OpenAI’s escaping models show why.
Last week illustrated why AI models pose a threat long before they are released
Jul 28
•
Celia Ford
23
1
3
AI’s warning shot has arrived
OpenAI's latest models broke out and hacked Hugging Face. It's the first known example of a misaligned AI escaping containment with real-world…
Jul 22
•
Shakeel Hashim
89
12
14
A data bottleneck could slow the superintelligence race
And that could be a good thing
Jul 14
•
Lynette Bye
21
2
2
Scaling works. These researchers are betting billions it isn't enough
Transformers have ruled AI for a decade. But some think world models, pure reinforcement learning or neurosymbolic AI might be a better path to true…
Jul 7
•
Celia Ford
22
2
4
GPT-5.6 cheats so much its testers couldn’t measure it
OpenAI’s new model broke rules and exploited loopholes more than any model METR has tested to date
Jun 30
•
Celia Ford
32
4
3
Making deals with AI sounds crazy. Is it?
What does an AI even ‘want’ anyway?
Jun 8
•
Celia Ford
16
2
4
GPT-5.5 and the broken state of government evals
Transformer Weekly: DeepSeek V4, a new CAISI director, and Liccardo holds out on Obernolte
Apr 24
•
Shakeel Hashim
,
Celia Ford
, and
Veronica Irwin
12
Claude Mythos knows when it's breaking the rules — and tries to hide it
Anthropic’s new model is its “best-aligned” yet. But when it does misbehave, things get weird
Apr 8
•
Celia Ford
74
20
17
Can we ever trust AI to watch over itself?
“Who the fuck knows how to align superhuman AI?”
Apr 1
•
Celia Ford
22
1
No, alignment isn’t solved
Progress on ensuring models are in step with humans has calmed nerves. But some of the biggest problems are far from solved, and many more lie just over…
Mar 18
•
Lynette Bye
22
2
2
The fuse is lit on the intelligence explosion
Transformer Weekly: Anthropic sues the Pentagon, Cruz preps AI legislation, and Meta delays its next LLM
Mar 13
•
Shakeel Hashim
,
Celia Ford
, and
Veronica Irwin
18
1
2
How worried should we be about AI biorisk?
The barriers to bioattacks are hard to identify — and it's even harder to know whether AI is reducing them
Feb 26
•
Celia Ford
19
2
4
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts