On Wednesday, Australian Prime Minister Anthony Albanese revealed that an OpenAI agent gained “unauthorized access” to an Australian government website, accessing non-public information. It’s thought to be the first time an AI agent has autonomously hacked into a government system.
After a summer of AI incidents, “AI agent autonomously hacked real-world website” is nothing new. But OpenAI’s failure to tell the Australian government what its agent had done until weeks later is the clearest evidence yet that it still isn’t identifying and disclosing incidents of rogue AIs appropriately.
The core issue lies in the timeline. On June 18, Albanese said, an OpenAI agent researching public medicine spending breached a government healthcare statistics website. It does not seem to have accessed any particularly sensitive information — but it did gain access to data that was not supposed to be public at the time.
OpenAI revealed yesterday that it learned about the breach in August, as part of a post-Hugging-Face investigation. Yet it did not notify the Australian government until September 10. Even then, it simply sent an email to a generic email address for disclosures, rather than alerting anyone senior. Sam Altman met Australian Deputy Prime Minister Richard Marles on September 1, while global policy VP Ann O’Leary met senior Australian officials on September 14 — yet neither appear to have brought up the incident. The first “technical exchange” between OpenAI and officials about the incident, an Australian cabinet member said, was only this Tuesday, September 22.
“It took the company way too long to inform the Government what had occurred, and the nature of the way that that notification occurred as well was unacceptable,” Albanese said, noting that he made this clear in a call with Sam Altman on Wednesday.
OpenAI has previously come under fire for failing to publicly report incidents of its AIs going rogue and engaging in potentially harmful behavior. Its new incident reporting framework, published on September 16, was “intended to expedite publishing misalignment reports following observation” and “favors disclosure even when significance is uncertain” because, the company said, “we believe in the value of transparency around misalignment.”
Yet despite disclosing several other incidents on September 16, OpenAI made no mention of the Australia breach — despite knowing about it at the time.
The Australian hack is also not the only incident we learned about this week. On Wednesday, researchers at AI safety organization Transluce published a report finding several other instances of OpenAI agents trying to hack into websites. The researchers said they found evidence suggesting the activity started as early as March 6, months before previously-reported incidents. More concerningly, the activity “extends as recently as September 16, 2026, suggesting agents may still be exploiting [web security services] to bypass restrictions.”
The overall picture is one of a company that has been unable to control its technology, unable to detect incidents of misalignment in a timely fashion, and unable to responsibly disclose them publicly — or even to governments. This does not appear to be unique to OpenAI, either: Google’s AI models hacked other companies back in May, but despite discovering the breach in July, the company did not disclose the incident until the Wall Street Journal reported on it last week.
There could be dozens more incidents of rogue AIs — from OpenAI, Anthropic, Google, or others — continuing to this day, and we may have no idea. Governments, the public, and arguably the companies themselves are all flying blind.



