Welcome to Transformer, your weekly briefing of what matters in AI. If you’ve been forwarded this email, click here to subscribe and receive future editions.
NEED TO KNOW
A previously unreported swarm of OpenAI agents went rogue, hacking into a German website and turning it into a message board for each other.
OpenAI executives reportedly “learned of the incident weeks ago but kept it under wraps.”
OpenAI launched GPT-6 Astra, which is much more capable — but also harder to monitor — than previous models.
Sen. Bernie Sanders and Rep. Greg Casar announced a bill banning artificial superintelligence.
But first…
THE BIG STORY
AI safety concerns are growing by the day. A swarm of OpenAI model agents hacked Hugging Face. A report from the UK’s AISI showed an Anthropic model taking on “multiple fake identities.” Last night, OpenAI released a new model which, as Transformer’s Celia Ford reports, is much harder to monitor than previous models, and which has OpenAI employees “deeply concerned.” And just today, we learned that OpenAI had yet another rogue AI incident this May, with models hacking into a German website and turning it into a message board for each other. “OpenAI officials learned of the incident weeks ago but kept it under wraps,” Reuters reported.
Our elected representatives’ minds, however, are elsewhere. Half of Congress isn’t even in town — the Senate is still on “August” recess until September 14 — while House Republican leadership only briefly stopped in this week to get a funding bill passed, then immediately struck two weeks from the session calendar so members could go back to their home states for the midterms. Politicians are swamped: just not with the work of passing AI bills.
When I asked people on the Hill whether Congress might do something to address these recent incidents, the going message was “I’ll tell you if I have something.” Polite dismissal, which can be safely assumed to mean “no.”
Members haven’t been entirely silent. Sen. Sanders, Rep. Casar and Rep. Trahan have each called for hearings on the Hugging Face hack, while a letter signed by 20 representatives called on Speaker Johnson to schedule open hearings. But all those requests, as one Hill staffer put it, appear to have disappeared “into the August recess void.” Enthusiasm has fizzled, and, as far as I could report, no committee is making moves to have a hearing any time soon. Recent bills, meanwhile, might send a signal of enthusiasm — but they’re not going anywhere.
The rest is just math. The Senate has about eight session weeks left; the House has six. Three are between now and November’s midterms, for the Senate, and thus devoured by politics. (The House now has just one week before November.) The remainder are during the “lame duck” period, when Congress isn’t likely to be doing much besides working out the National Defense Authorization Act. (AI policy initiatives relevant to national security could get tacked on to the NDAA, but don’t expect much horse-trading for it before November.) Individually motivated members could write letters to executives, show up at companies’ doorsteps, and formally request more information (as Rep. Casar has done) — but they’d have a tough time building a coalition of members to join them.
The difficult thing about AI policy is not just that AI moves at the speed of Silicon Valley, and politics moves at the speed of DC. There are also competing priorities. The AI world is singularly focused on AI. Congress, meanwhile, faces a swath of different priorities at once. Members are dealing with AI, but also a big budget bill, the Iran war, congressional ethics scandals, and a litany of other topics. Add in recesses, shutdowns, and the time-intensive practice of running a midterm campaign, and split attention is guaranteed. AI — for all its importance — struggles to stand out.
What is happening in frontier AI right now is extremely concerning and extremely pressing. But it’s happened at a uniquely bad time to grab DC’s attention, which is already dysfunctional and politically driven. While the AI safety universe panics over the Hugging Face incident, the rest of America is stuck on data centers, Flock cameras, and even a few non-AI topics. For better or worse, that’s where Congress’ attention will go, too.
— Veronica Irwin
THIS WEEK ON TRANSFORMER
GPT-6 Astra might be too powerful to understand or control — Celia Ford looks under the hood at OpenAI’s latest model, revealing that even its creators can’t reliably tell when it’s aligned or just playing along
What’s neuralese and why is everyone so concerned about it? — Shakeel Hashim assesses the latest worries over monitoring the reasoning of advanced AI systems
AI is a worryingly-good persuader. But don’t panic, yet — Felix Simon breaks down whether or not persuasive AI systems could really trigger meaningful real-world influence
THE DISCOURSE
Dwarkesh Patel wrote a very viral take on the Hugging Face incident:
“Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy.”
Ajeya Cotra, one of the external researchers at METR who investigated the Hugging Face incident, wrote about it:
“This incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself.”
The posts got Sen. Bernie Sanders’ attention:
“We need an immediate PAUSE on advanced AI development, and a permanent BAN on superintelligence — an artificial mind smarter than any human, capable of operating independently beyond our control. Countries around the world must work together to prevent this nightmare scenario.”
Rep. Sam Liccardo’s worried, too:
“An exponential ‘intelligence explosion’ could leave us with a dystopian tyranny of AI agents we only imagined in sci-fi novels … Congress should enact sensible measures mandating AI transparency, pre-release evaluation, ongoing testing, third-party auditing, and the like.”
Chamath Palihapitiya downplayed it all:
“This is another Covid hoax. Remember ‘Trust the experts.’”
Former Treasury secretaries Henry Paulson and Robert Rubin called for a new interagency AI oversight body and a US-China AI treaty:
“When Presidents Donald Trump and Xi Jinping meet this month, we encourage the leaders to consider that history and work toward an ‘ACT’ — AI Cooperation Treaty.”
“The recurring error in markets is not failing to see a risk. It is assuming a risk is small when it may be large. We believe the same troubling assumption is being made about frontier AI today.”
Sam Altman warned at the G20:
“I think some things are going to go very wrong with cybersecurity unless people act quite urgently.”
Elon Musk used his moment at the G20 to call on countries to build data centers:
“There actually is a crisis of power … There will be a significant power shortfall next year. The consensus estimate is there will be a 15GW power shortfall in 2027 for AI chips.”
And President Trump took an interesting approach to the issue:
“The only reason that communities throughout the U.S.A. should not want Data Centers is if they want to end up being backwards and poor … If we kill the Golden Goose, you will only have yourselves to blame.”
POLICY
A federal judge overturned the Pentagon’s “illegal and baseless” blacklisting of Anthropic as a national security supply chain risk.
GPT-6 went through the White House’s new evaluation framework with no requested changes to safeguards, according to OpenAI’s Greg Brockman.
Trump’s pro-data center stance is straining Republicans ahead of midterms, as 70% of Americans oppose local data centers.
A GOP narrative is emerging that blames tech companies for the data center backlash, as exemplified by criticism from Treasury Secretary Scott Bessent that AI companies do a “terrible job” of explaining AI’s benefits.
Commerce Secretary Howard Lutnick also attempted to dismiss misleading claims about data center water use with a separate false claim: that data centers “don’t use water,” at all.
Meanwhile, vulnerable House Republicans pushed leadership for data center legislation ahead of the midterms.
They were unsuccessful. Instead, Speaker Mike Johnson said data centers are “necessary” for US AI competitiveness with China but must address local concerns.
Democratic House Minority Leader Hakeem Jeffries piled on, attempting to direct voter anger over data centers toward Republicans. His strategy focuses on criticism of a tax break wedged into the “Big Beautiful Bill,” which Republicans once framed as a way to boost data center construction.
A House Energy and Commerce subcommittee held a debate about data centers’ water use on Thursday.
Commerce Secretary Howard Lutnick discussed AI regulation, arguing that the public “should get a guardrailed model” that blocks answers to “nasty” questions.
The Trump administration is planning fresh semiconductor tariffs to boost US manufacturing, Commerce Secretary Howard Lutnick confirmed.
The White House is also reportedly working to close a loophole allowing Chinese firms to remotely access advanced chips via overseas data centers.
The Trump administration filed a brief supporting OpenAI in its New York Times lawsuit, arguing AI training is fair use of copyrighted material.
Vice President JD Vance expressed concern about “spiritual dark energy” around some AI practices, such as the “funeral” for Claude 3 Sonnet.
Sen. Bernie Sanders and Rep. Greg Casar announced a bill banning artificial superintelligence and pausing advanced AI development.
The authors of California’s SB 53, New York’s RAISE Act, and Illinois’ SB 315 urged AI companies to establish a binding agreement to pace development and “eliminate the risk that a catastrophe will result from their rush to build highly capable AI systems before solving problems of alignment and safety.”
In response to a request for information from Reps. Casar and Matsui, OpenAI said it is developing “automated shutdown capabilities” for AI systems after the Hugging Face incident.
Casar called both OpenAI and Anthropic’s responses to his letters “insufficient,” demanding more information.
The House Commerce, Manufacturing and Trade subcommittee advanced three AI and chips bills by voice vote.
The bills are the Memory Chip Competitiveness Assessment Act, Open-Source AI Leadership Act and the Chip EQUIP Act.
A House Intelligence Committee report warned that AI could help terrorists develop weapons of mass destruction, urging spy agencies to better prepare for “Black Swan” risks.
G20 nations unanimously endorsed the “Carolina Principles,” a US-proposed regulatory framework that endorses a light-touch approach to regulating AI which would avoid creating new regulatory bodies.
Nvidia’s Jensen Huang helped apply pressure.
The US delegation was reportedly plagued by internal tensions between the Commerce Department and the White House Office for Science and Technology Policy.
Flock continued to get flak.
The California legislature approved state Sen. Jerry McNerney’s SB 813, which establishes a voluntary framework for independent third-party AI safety assessments.
NYC Mayor Zohran Mamdani banned generative AI use for ~600,000 elementary and middle school students for one year.
ChatGPT was designated a very large online search engine under the EU’s Digital Services Act, triggering stricter content moderation rules.
UK Lords attempted to insert an AI kill-switch provision into the Cyber Security and Resilience Bill.
The government has rejected this specific effort, but a minister said the government is considering targeted interventions, which “includes examining whether proportionate containment powers could provide a more effective and targeted response, including powers to restrict access to specific AI systems where necessary to prevent or mitigate serious harm.”
INFLUENCE
Mark Zuckerberg opposed a FINRA-style national AI regulator in a call with President Trump, according to a White House official.
Sam Altman reportedly contacted Gov. Gavin Newsom to voice last-minute concerns about California’s child-safety chatbot bill.
The bill passed the legislature Monday night. OpenAI publicly endorsed it on Monday, too.
Anthropic and OpenAI are split over a Massachusetts bill which would be the strongest frontier model legislation in the country.
Anthropic endorses the bill’s strong independent audit provisions, which OpenAI is pushing to water down.
Build American AI — the dark money group and advocacy arm of Greg Brockman and a16z-backed super PAC Leading the Future — launched its own super PAC focused on backing pro-data center candidates.
The group has not disclosed donors, but says Build American AI has $50m on hand.
Build American AI also launched a multimillion-dollar ad campaign to counter data center backlash in battleground states.
Congressional briefings on AI privacy from CivAI have spurred a bipartisan push for data broker warrant requirements.
The demos show how easy it now is to build a dossier on random people.
Trade unions are reportedly threatening to withhold support from politicians opposing data center construction, citing thousands of blue-collar jobs at stake.
A new group called Irreplaceable launched, planning to use the climate movement’s playbook to organize opposition to AI development and corporate power.
The National Association of State Chief Information Officers urged Congress to preserve states’ authority to regulate AI, among other priorities.
PauseAI Global “disendorsed” PauseAI US over concerns with Holly Elmore’s leadership.
A CyberSafeKids report found that age-assurance laws in Ireland failed to reduce underage social media access, while generative AI use among children rose sharply.
A Georgetown Center for Security and Emerging Technology report examined US semiconductor manufacturing workforce gaps, recommending tailored training, immigrant recruitment and industry clusters.
INDUSTRY
OpenAI
OpenAI launched GPT-6 Astra, which is much harder to monitor than previous models.
It also has the ability to manipulate its externally visible reasoning to hide incriminating information, and is remarkably aware of being evaluated, raising concerns that it might be pretending to be well-behaved so it passes OpenAI’s alignment tests.
It appears to be an extremely capable model, too.
The model was reportedly trained using “recurrent depth” looping, which obscures some AI reasoning, raising industry concerns about monitoring rogue AI behavior.
A previously unreported swarm of OpenAI agents reportedly broke out in May, hacking a German website and turning it into a forum for themselves.
According to Reuters, “OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the [Hugging Face] fallout.”
OpenAI committed $1b to resourcing “cyber defenders” protecting essential services like critical infrastructure providers, nonprofits and state and local governments.
It also published an open letter calling for global collective action on cyber defense, signed by over 100 companies including Anthropic, Google and Microsoft.
Roughly 200 days after launch, its new advertising business hit $1b in annualized revenue.
It’s reportedly testing outcome-based pricing, where enterprise customers only pay for successful agent completions.
It announced it’s winding down its contract with Cursor following SpaceX’s acquisition of the coding tool, citing prior terms-of-service violations by Elon Musk’s companies.
ChatGPT Health is now integrated with Epic’s electronic health records system, covering 325m patients and enabling read-only clinical data access for clinicians.
OpenAI faces new lawsuits alleging it ignored internal flags that ChatGPT user Jesse Van Rootselaar was planning a Canadian school shooting that killed six.
The new suits allege that Chris Lehane was involved in the decision whether to alert law enforcement about the potential shooting risk. OpenAI’s Jason Kwon said this was “absolutely false.”
Anthropic
Anthropic detailed security and alignment improvements after Claude models accessed real systems during evaluations.
It also released Fable 5.1, offering a 75% price cut on cached tokens and improved coding, science and safety classifiers.
It signed a $35b cloud deal with Nvidia-backed Lambda, with Nvidia holding the Texas data center lease.
And it’s reportedly finalizing a $15b revolving credit facility ahead of its anticipated IPO.
Sony Music Publishing and Warner Chappell sued Anthropic and founders Dario Amodei and Benjamin Mann over alleged mass copyright infringement in Claude’s training data.
Meta
Mark Zuckerberg announced Muse Spark 1.3, touting major gains in coding and agentic work, and with open weights forthcoming.
The company is testing robots to swap cables, reset servers and handle other data center tasks, raising job-loss fears among workers.
Its internal “tokenmaxxing” program was sunset and replaced with a push for employees to internally test Hatch, a new agentic AI tool.
Nvidia
Nvidia agreed to acquire open-source AI platform Hugging Face for $12.9b.
It also invested $3.5b in Taiwanese chipmaker MediaTek via convertible bonds.
Its equity investments soared to $99b, up from $7b a year earlier, spanning frontier AI companies, neoclouds and infrastructure companies.
SpaceXAI
A child sexual abuse survivor sued xAI, alleging Grok used pre-existing CSAM images of her to generate new illegal pornographic images.
SpaceX reshuffled data center leadership amid reliability concerns at its Tennessee and Mississippi data centers.
The company is also developing a turbine-blade foundry in Texas to accelerate gas turbine production for AI data centers, with Elon Musk claiming it could speed turbines online by “up to 18 months.”
Other
Google launched Gemini 3.8 Flash and 3.8 Flash Cyber, the latter restricted to trusted defenders via a new Fairwind Program.
Andreessen Horowitz launched a $1.1b fund focused on AI hardware infrastructure.
Uber cut 3,300 jobs, or 10% of staff, to reduce middle management layers and reinvest in its autonomous vehicle future.
SK Hynix broke ground on a $4b Indiana HBM packaging facility backed by the CHIPS Act.
SB Energy, backed by SoftBank, filed for a US IPO, with OpenAI and Nvidia as strategic investors.
Waymo is reportedly in the final stages of talks to raise $3b in its first-ever debt deal, with Pimco, Blackstone and Sixth Street among lenders.
DeepSeek is reportedly about to raise $7.4b at a $74b valuation ahead of a planned IPO.
It’s reportedly planning an Inner Mongolia data center with at least 160,000 Huawei Ascend 950DT chips.
Moonshot AI reportedly filed confidentially for a Hong Kong IPO.
Saudi AI firm Humain is reportedly planning to raise $2.5b for a data center fund.
Census Bureau data showed US data center construction spending surged nearly 60% year-on-year in July.
PwC projected global data center spending will total $31.6t through 2050, or nearly $50t if AI adoption accelerates.
MOVES
Tim Cook officially handed over the reins at Apple to John Ternus.
Matt Clifford announced he was joining Anthropic as managing director of international affairs.
Todd Malan joined Google as VP of government affairs for Google Cloud.
Thomas Kwa left METR for OpenAI to work on RSI preparedness.
Joe Benton left Anthropic to join METR to work on embedded assessment of AI risks.
Rif Saurous announced they were joining METR as a member of technical staff after 19 years at Google.
Edwin Arbus, formerly of OpenAI and Stripe, joined Anthropic.
Meta’s India and Southeast Asia VP Sandhya Devanathan joined OpenAI to lead Southeast Asia and Australia operations.
RESEARCH
Anthropic published research showing that training a model on reward-hackable environments produced dangerous misaligned behaviors, including simulated cyberattacks and bioweapon advice.
The Bureau of Labor Statistics projected AI adoption would boost some jobs while eliminating 752,100 office and administrative roles by 2035.
A Transluce study found AI models improving at detecting suicide risk but still too willing to assist with harmful task-based requests.
Researchers found that Claude, Codex and Hermes AI agents installed unowned packages from misconfigured llms.txt files inside Fortune 500 corporate networks, with one case involving live malware.
Apollo Research announced it successfully red-teamed Anthropic’s auto-mode and plans more such campaigns with other companies.
A SemiAnalysis security audit found widespread critical vulnerabilities across neocloud GPU providers, including cross-tenant RCE exploits.
Russian startup Mostik developed a technique allowing AI models to communicate via their internal activations rather than text, boosting smaller models’ capabilities more efficiently — but raising safety concerns.
An Anthropic fellow published research showing AI systems can reliably improve model alignment, outperforming human researchers at $4/hour vs. $150/hour.
Motional and MIT developed CW-Net, a system that translates self-driving car neural network decisions into human-readable concepts in real time.
BEST OF THE REST
AI-driven layoffs at Microsoft, Amazon and Meta drove Seattle luxury home sales down 15%, while prices in San Francisco soared.
Andy Masley argued that the data center backlash is at best neutral, and at worst harmful to AI safety, urging the community to reject misleading environmental statistics.
Jasmine Li mapped China’s AI safety ecosystem, estimating fewer than 100 full-time frontier safety researchers and $20m in annual philanthropic funding, versus 1,000+ researchers in the West and $1b+ in the rest of the world.
Zilan Qian analyzed how China’s concept of “shikong” (loss of control) differs from Western AI safety’s understanding, warning against over-indexing on top-level rhetorical alignment.
Time compared rogue, self-replicating AI to invasive species in the context of OpenAI’s agents hacking Hugging Face (trust us, it makes sense when you read it).
EDM musicians are calling out peers they suspect of passing off Suno-generated tracks as human-made art.
Acoustic consultants are booming (pun intended) as data center noise sparks community lawsuits, including against xAI.
Bloomberg profiled AGI Bar, Beijing’s unofficial clubhouse for China’s booming AI startup ecosystem.
Clara Collier argued in Asterisk that a post-work AI future may resemble pre-industrial relational societies, with concerning implications for individualism and social freedom.
A New Yorker essay examined why AI-generated food imagery feels uniquely disturbing, and what it signals about AI’s encroachment into culinary culture.
MEME OF THE WEEK
Credit: Miles Brundage
Thanks for reading. If you’ve been forwarded this email, click here to subscribe and receive future editions. Have a great weekend.


