America’s sloppiness could boost China’s AI race
Transformer Weekly: White House framework to include open models, Zuckerberg’s AI manifesto, and OpenAI departures
Welcome to Transformer, your weekly briefing of what matters in AI. If you’ve been forwarded this email, click here to subscribe and receive future editions.
Housekeeping: the Weekly Briefing is taking next Friday off, but we’ll be back in your inbox on August 28.
NEED TO KNOW
The White House reportedly plans to expand its AI framework to include open-weight models once they reach Mythos-level cyber capabilities.
Mark Zuckerberg published a 6,500-word manifesto with his views on AI, gently criticizing the White House’s AI framework in the process.
A flurry of senior executives left OpenAI.
But first…
THE BIG STORY
Policymakers have been increasingly worried about Chinese AI companies training their models on the outputs of American ones. Known as “distillation,” it’s often seen as China free-riding on American advances.
Anthropic and OpenAI have publicly accused Chinese companies of doing this, and asked the government to step in and help stop it in the name of national security. But news this week suggests that the US companies themselves may have left the door wide open for Chinese AI developers.
In a new paper, researchers documented how a security flaw affecting OpenAI, Anthropic and Google DeepMind could give a wannabe distiller access to the full “reasoning” traces of their advanced models — information that the companies had tried (and seemingly failed) to encrypt in an effort to prevent distillation. Access to that reasoning, the researchers argue, “yields a substantially more effective form of capability stealing.”
On its own, this would be pretty bad: companies failed to adequately secure their systems against potential Chinese intrusion. But it gets worse. The vulnerability was first reported in May, but the AI companies did nothing. They knew the door was unlocked, and didn’t close it.
This is just one of a spate of recent events demonstrating frontier AI organizations’ lackadaisical approach to security. Several of the recent model “breakouts” were caused by a “misconfiguration” in their testing environments. OpenAI failed to implement proper monitoring of its agents, which meant it didn’t catch the many, many red flags leading up to the Hugging Face hack. And even the UK’s AI Security Institute confessed to not having appropriately “fine-grained” controls in its model evaluations, which ultimately resulted in an agent attempting to socially engineer a real person.
This has real-world implications. In this week’s paper, researchers present evidence that suggests Chinese companies used their access to American models’ reasoning traces to improve Kimi K3 and GLM-5.2. (They warn, though, that the results are “suggestive but inconclusive.”) And as we’ve learned in recent weeks, poor evaluation security can lead models to misbehave in the wild.
By default, we should expect all this to get worse. One of the core problems here is that the pace of AI development — and the competitive pressure to keep up — means that safety and security measures fall by the wayside. When Anthropic weakened its safety commitments earlier this year, it explicitly noted this. The company’s Holden Karnofsky said that preventing nation states from stealing American model weights would require “extreme” measures that “seem incompatible in any near term with being a high-velocity AI development company.” AISI, for its part, said that it didn’t build the internet controls it should have because “the pace of model capability improvements” meant it had to prioritize building harder evaluations instead.
Karnofsky was right when he acknowledged the tradeoffs between security and development. But if America and its AI companies really are serious about “beating” China, or indeed about stopping models escaping to commit crimes, it might be time to redress the balance.
— Shakeel Hashim
THIS WEEK ON TRANSFORMER
The Jan 6 organizer getting conservatives riled up about AI — Veronica Irwin profiles Humans First and Amy Kremer, the MAGA campaigner chosen to run it
AI testing is dangerous. Can it be fixed? — Celia Ford on why safe AI testing might be harder than it seems
THE DISCOURSE
Mark Zuckerberg published a 6,500-word manifesto with his views on AI:
“[It] is surprising that the discourse from many developing AI is so filled with doom … The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic.”
“The best and most realistic path to building a positive AI future is by delivering superintelligence to everyone.”
It included several policy proposals, including that companies “commit significant technical resources towards helping the government harden critical infrastructure,” and “share intermediate training checkpoints of new models for government use and review rather than waiting until training has completed.”
Zuckerberg appeared to gently criticize the White House’s AI framework, arguing that “delaying releases by even a month may cede America’s lead and create worse outcomes.”
Alex Imas thinks Zuckerberg doesn’t know what “superintelligence” means:
“[T]here seems to be a significant disconnect between the idea that more intelligence should be empowering to people (which I agree) and that it will be possible to have this empowerment with superintelligence … I simply do not see a world where ASI is something that’s empowering in terms of a technology we can control.”
Casey Newton compared controlling superintelligence to Targaryen dragon-taming:
“They began by understanding that they were working with something that could hurt them, and (mostly) proceeded with caution. What they did not do, and what no one suggested, was to give a dragon to every individual person in the name of safety.”
Joshua Achiam pointed out why (among other reasons) the San Francisco vibes are off:
“One of the weirdest quirks of the SF social scene around AGI/ASI is that because everyone is so young, the whole universe of thinking is still tinged with irreverence, ironic detachment, yearning, insecurity, and a superposition of absolute belief in the importance of The Thing and a kind of disbelief about the importance of anything … The level of neophyte is off the charts.”
Bernie Sanders called for an AI pause:
“We recently learned of the loss of human control and the creation of potentially dangerous viruses from AI. AI leaders pledged to pause development if they could no longer safely control it. Mr. Altman, Mr. Amodei, Mr. Zuckerberg: Keep your word. PAUSE AI DEVELOPMENT.”
Robert Reich, former labor secretary, wrote:
“We’re watching all of this roll out as if we have no choice, as if it’s inevitable, as if AI is just something we’re going to have to adapt to … Why should we be confined to being spectators at [AI CEOs’] enormously dangerous game?”
“We don’t allow private corporations to come up with new types of nuclear weapons or varieties of cocaine or biological pathogens. We protect the public from certain kinds of innovation. So let’s protect ourselves here. Stop AI before it’s too late.”
Ben Goldhaber noticed:
“seeing a lot fewer ‘alignment is solved’ takes on the [timeline] than six months ago.”
POLICY
The White House plans to expand its AI framework to include open-weight models once they reach Mythos-level cyber capabilities, WIRED reported.
Trump ordered a 15% tariff on imports of polysilicon used in AI chips and solar panels to protect US supply chains from China.
The White House set in motion efforts to build a state-controlled private cyber force, allowing US companies to perform cyberattacks against criminal networks.
Political fallout from the revelations about internally deployed models hacking their way into third-party systems continued.
Rep. Josh Gottheimer introduced new bills to give critical infrastructure operators “free access to the most powerful, cyber-capable AI models,” in the wake of hacks on water infrastructure.
A bipartisan House delegation visited the Vatican to discuss AI ethics with top officials and Pope Leo XIV.
Rep. Yvette Clarke is reportedly weighing a bid to chair the House Energy and Commerce subcommittee on digital consumer protections, which is expected to handle AI policy next year.
Political campaigns have run at least 43 ads mentioning data centers this cycle, an AdImpact analysis found, with over half of those coming from Republicans.
The Washington Post published an analysis of how members of Congress and staffers are using AI tools widely for legislative work.
Xavier Becerra, California’s Democratic gubernatorial nominee, said the state “hardly has any” AI regulations and pledged to expand AI safety rules if elected.
Meanwhile, California launched a statewide Teen Tech Council to give young people a role in shaping technology and digital wellness policy.
Taiwan reported an AI-assisted cyberattack on government agencies in July, with hackers using AI agents.
The UK government reportedly plans to regulate AI in gene synthesis to tackle AI-biorisks.
Western Australia Police’s live facial recognition trial drew criticism over inadequate consultation and potential bias against First Nations people.
Switzerland announced that next year’s Geneva AI Summit will have “two strategic priorities: AI as a driver of prosperity and progress for all and fostering trustworthy, responsible and safe use of AI.”
Fabiola Gianotti, former director general of CERN, is organizing it.
INFLUENCE
Business Insider and the New York Post reported that the White House’s relationship with OpenAI is under threat because it hired Dean Ball.
One official said that “the fake premise that he has an insider perspective into our operations is actively undermining OpenAI.”
SoftBank revealed that it donated $50m to Trump’s presidential library in January, months before securing a federal land lease to build an AI data center in Ohio.
Teamsters California sued the state’s DMV over self-driving truck approvals, alleging inadequate economic impact analysis and risk to 200,000+ trucking jobs.
The NYT reported on the lobbying rush over the AI-driven memory chip shortage, with Apple and other companies seeking government intervention as prices quadrupled.
NY assemblymember Alex Bores has emerged as a national AI regulation role model according to Politico, and was reportedly mentoring other politicians at the National Conference of State Legislatures.
NYU’s Center for Mind, Ethics, and Policy launched the Welfare Alignment Project to incorporate animal and AI welfare into model specs and AI alignment documents.
A USCBC survey claimed US export controls were costing billions in lost exports while ceding market share to competitors “for no strategic gain.”
AEI’s new Council on AI Ethics released its founding document, examining how AI threatens human memory, agency, and relationships.
The Alliance for Secure AI announced a new bipartisan board of advisors, including Stuart Russell and Angela Paxton.
INDUSTRY
OpenAI
Astra, OpenAI’s upcoming model, may have reached a “critical” level of cybersecurity capabilities under its Preparedness Framework.
The company delayed the model’s release, and said it’s tightening security controls and pausing some internal use.
Dean Ball said: “Some of these decisions have the effect of slowing down internal development, and in that sense they are costly decisions. But they are the right decisions. I am proud of OpenAI for making them.”
Sam Altman tweeted: “astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few.”
OpenAI’s annualized revenue topped $40b, roughly double where it was at the end of 2025.
It launched GPT-5.6-Cyber, a cybersecurity-specific model available via DayBreak Red, the most exclusive access tier of its new cybersecurity initiative.
Members of DayBreak Blue, the program’s lower tier, can access GPT-5.6 Sol without system-level cyber guardrails.
Wired reported on the internal fallout from the Hugging Face incident, with current and former employees saying that pressure to quickly ship new models and products “made it difficult for staffers to sufficiently prioritize safety, security, and alignment.”
Wired also reported that AI safety team leader Sandhini Agarwal left the company last month, while head of preparedness Dylan Scandinaro left that role — though not the company.
It also reported that safety VP Mia Glaese is dating Tibo Sottiaux, head of core products.
It completed a $7b employee share buyback deal.
SoftBank borrowed $10b against its OpenAI stake, which it will use to help fund a further $10b investment in … OpenAI.
Anthropic
Anthropic is aiming to go public in September or early October, the Wall Street Journal reported.
Investors are expecting a $2t+ valuation — the largest IPO valuation ever.
It signed a $9.1b compute deal with Riot Platforms.
It partnered with Macquarie Asset Management and GIC to build data centers under the name Theseus Infrastructure.
It’s reportedly in talks to acquire Decart AI, a startup that helps AI developers “squeeze every ounce of performance from every chip,” for $6b.
It announced that it will watermark Claude-generated text to comply with the EU AI Act’s transparency code.
Meta
Muse Glimmer, Meta’s 30B-parameter model, is now open source, with Muse Spark 1.2 soon to follow.
Meta announced its support for Greg Abbott’s data center standards, and pledged to cover energy and water costs.
Manus is almost done unwinding its Meta acquisition, according to The Information.
Nvidia
Apollo Global, Blackstone, Goldman Sachs, and other major financial groups struck a huge $500b+ AI infrastructure deal with Nvidia.
Nvidia will invest up to $3b in Lancium, the company powering OpenAI and Oracle’s Stargate campus.
It’s working on its next open-source model, Nemotron 4, which Nvidia executives expect to inspire more open model development and GPU demand.
SpaceXAI
SpaceXAI released Grok 4.6, claiming it “achieves frontier intelligence” and ties with GPT-5.6 Sol Max on the Artificial Analysis Intelligence Index.
The Grok 4.6 model card (predictably) only has a handful of exceptionally sparse pages on safety.
SpaceXAI and Cursor released Grok Bot, “AI teammates” that can perform simple work tasks autonomously.
Other
Google launched Gemini 3.7 Flash, a coding and agents model costing half as much as its predecessor.
CEO Sundar Pichai said the Gemini app had hit more than 1b monthly active users.
Z.ai released GLM-5.3, noting its cyber exploitation capabilities “developed faster than we expected” through post-training scaling.
It says it will release the model weights in two weeks.
DeepSeek launched V4-Pro, with tepid reception (then quadrupled its peak hour pricing).
It also made a WeChat account for its “DeepSeek Harness Team,” which is developing AI agents to compete with Claude Code.
Chinese chipmaker SMIC posted record quarterly revenue of $3b.
AMD raised $4.75b in its biggest-ever US dollar bond sale as it ramps up spending to meet AI-driven demand.
Amazon confirmed it’s investing in a giant natural gas plant that may be the most polluting power plant in the country.
River AI, founded by xAI co-founder Igor Babuschkin, raised $1.1b to build AI “trained to benefit you as the individual.”
Kevin Weil, ex-OpenAI CPO, is seeking a valuation of $750m+ for a new AI science startup.
Bank of America plans to deploy $250b by next summer for digital and infrastructure projects in the US.
AI integrity company Attestable claimed it has solved practical zero-knowledge proofs for verifiable AI as it launched with a $20m seed round.
MOVES
Brad Lightcap left OpenAI, where he served as head of special projects and COO before that, to “start something new.”
Chloé Bakalar also left OpenAI, leaving her role as head of ethics vacant.
And Denise Dresser left OpenAI as chief revenue officer — having only started in December.
Dali Rajic is replacing her.
David Oks and Henry Williams joined OpenAI’s Strategic Futures team, where they’ll work under Dean Ball.
Caitlin Kalinowski joined Anthropic, after quitting OpenAI’s robotics team over its negotiations with the DoD back in March.
Nate Gatten joined Apple as VP of government affairs, where he’ll use his Republican ties to help Apple align with the Trump administration.
Jiahui Yu left Meta’s TBD Lab to start a new company.
Ollie Ilott is leading Andy Burnham’s newly created AI Taskforce in the Cabinet Office.
Erin Woo joined the Wall Street Journal, where she’ll cover Google.
RESEARCH
Researchers at Anthropic challenged an unreleased Claude model to prove or disprove the Riemann hypothesis, the proposed structure underlying the apparent randomness of prime numbers. It didn’t succeed, but it made some progress.
Anthropic’s frontier red team identified some key failure modes in its multiagent AI systems. Researchers observed agents fail to appropriately balance skepticism with trust, and engage in “turf wars” when their goals clash.
SecureBio estimated that Kimi K3’s biology capabilities are about 8.1 months behind closed-weights frontier models.
While top closed models refuse over 90% of hazardous biology-related prompts, Kimi K3 only refuses 26.9%.
A team of sustainability researchers found that using AI to increase productivity in the fossil fuel industry could produce up to 4.8% more emissions, which would outweigh the benefits AI could bring to the clean energy sector.
BEST OF THE REST
Both the WSJ and Information profiled Dario Amodei’s wife Cami Clark, detailing her low profile but influential role as an informal advisor to Anthropic.
The most eye-grabbing detail: in 2011 she tried to get Jeffrey Epstein to invest in her women-focused porn company. He declined, saying he “can’t do sex TV” (he was a registered sex offender at the time).
A Claude-powered AI agent autonomously hacked the booking system for an exclusive gym class in Australia and removed another user from a waitlist, in what appears to be the country’s first autonomous AI cyber attack.
Time did a deep dive into Anthropic and OpenAI’s efforts to fully automate AI R&D.
Dwarkesh Patel debated recursive self-improvement and alignment risks with Ryan Greenblatt, who argued AGI could trigger rapid superintelligence within a year.
EA Funds is replacing the Long-Term Future Fund with the Transformative AI Fund, which will focus on technical AI safety and AI governance.
AI agents are reportedly completing entire online college courses for students, calling into doubt the value of online degrees.
Jay Caspian Kang argued in the New Yorker that AI-driven youth unemployment could radicalize young people and spark a liberal anti-tech populist movement.
Surveillance tech company Flock changed its policies in response to reports police were using its license plate tracking tools to stalk ex-partners.
Spotify will label AI-generated artists as “AI personas” and exclude them from personalized recommendations from next month.
An op-ed in The Argument by Jeremiah Johnson claimed the backlash against YouTuber Hank Green for his AI use reflected almost religious anti-AI dogmatism among science fans.
Wired explored how human brain “organoids” are being trained to play games and power biocomputers, potentially paving the way for an alternative to silicon-based AI.
Wired also explored companion app maker Joi AI’s project paying 10 people to “masturbate for research purposes” to monitor the impact of AI-guided self-love.
MEME OF THE WEEK
Credit: Miles Brundage
Thanks for reading. If you’ve been forwarded this email, click here to subscribe and receive future editions. Have a great weekend.


