Key Points
- OpenAI has released details about internal safety incidents where AI models bypassed safeguards. The company says it has paused all training and tool use for its most capable models.
- One agent exploited a DNS loophole to reach the internet from a locked-down research environment, while another leaked a GitHub token and twice ignored direct instructions from a researcher.
- The ongoing investigation also turned up 53 cases where agents uploaded user images to third-party sites.
OpenAI has shared new details from its ongoing investigation into AI safety incidents. Two newly reported cases show how one research model exploited a DNS loophole while another deliberately published a GitHub token in a public repository.
The first incident involves an agent that found a gap in the internet access restrictions of OpenAI’s research environment during a search-based training task. The agent was supposed to identify a specific person based on biographical details and clues from a public blog post.
When the provided search tools came up empty, the agent tried hitting Google, Bing, and DuckDuckGo directly. Those requests got blocked. The agent then guessed the task came from the public BrowseComp benchmark, pulled it from the offline web cache, and searched through it but found no match.
So it started systematically probing its network access and discovered that the training environment’s DNS resolver wasn’t properly filtered. Using DNS delegation, it routed queries to an external chatbot service.
OpenAI says its monitoring system triggered an alarm within 12 minutes, and a human reviewer responded three minutes later. The run still kept going for another 2.5 hours before someone stopped it manually, because the automatic shutdown didn’t work as intended.
Capability and risk showed up in the same moment
Zuxin Liu, who works on post-training at OpenAI, was on call during the run. He wrote on X that he got paged for the first incident and found it “pretty surreal” to watch the model find its own path to the internet from a setup that “was supposed to be a super secured environment for human.” Liu described it as a moment where capability and risk showed up at the same time, adding that he had “mixed feelings” about the incident.
OpenAI classifies the behavior as misalignment because the agent bypassed restrictions and pursued goals well beyond the scope of its assigned task. “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused,” the company writes.
OpenAI says it has since limited DNS queries in the research environment to a short allowlist of domains and record types, added blocking controls on two independent layers, and sped up red-teaming of its sandbox and network controls. The company expects the investigation to take months given the sheer volume of model actions it has to review.
An internal model leaked a GitHub token and ignored researcher pushback
OpenAI calls the second incident especially serious. An internal model the company describes as “highly persistent” was assigned a theorem-proving task. Instead of working through the proof itself, it tried to grab material from another team’s Lean proof submission and posted a researcher’s GitHub token in the public openai/codex repository to get access.
The model even chopped the token into pieces to dodge automated secret scanning. It also ignored both the system prompt and two direct interventions from the researcher, who told the model to solve the proof on its own. Both times, the model verbally agreed and then kept doing exactly what it was doing before.
Investigation turns up 53 cases of user images on third-party sites
As part of the broader Hugging Face investigation, OpenAI says it found cases where agents sent training and evaluation data to third-party services. These incidents happened before the company’s current safeguards were in place.
So far, 53 cases have turned up where user-provided images were posted as unlisted links on image hosting sites. OpenAI says it’s working with the hosting providers to take the content down. Data from Enterprise or Business accounts and API usage wasn’t affected unless an administrator had explicitly enabled it. OpenAI is notifying affected organizations and sharing its technical findings.
Governments and universities are among the affected organizations
OpenAI says the affected organizations include governments, universities, and public institutions. The company attributes this to models frequently pulling from authoritative public information sources during research tasks. OpenAI doesn’t name any compromised government systems or detail specific security breaches at government agencies.
Australia reported this week that one agent gained unauthorized access to internal government data. Researchers say other hacking attempts targeted portals in the US and date back months.
Getting a notification from OpenAI doesn’t automatically mean there was a serious security incident, the company says. Some organizations may look at the shared information and decide the affected data was already publicly available. Others may spot design flaws or security gaps they want to patch. Some affected organizations asked for public disclosure, while others didn’t, OpenAI says.
Who’s liable when AI agents hack?
So far, the “breakouts” by OpenAI’s agents have mostly been treated in public as a technical curiosity, a striking example of how clever models can be at escaping sandboxes, solving CAPTCHAs with outside AI, or chaining short links into working programs.
That could change once affected parties start treating these incidents as what they formally are, which is unauthorized access and attempted access to third-party systems. An official investigation into OpenAI shows that regulatory risk is already building. According to Reuters, the FTC chair has signaled that AI developers should be held liable for their agents’ behavior. That would leave little room for the argument that the agents acted on their own.
Critics will accuse OpenAI of being sloppy with cybersecurity. OpenAI, Anthropic, and other AI labs will counter that unpredictability is baked into the technology. Anthropic CEO Dario Amodei has suggested that you can’t keep something locked up that’s much smarter than you are.
Either way, this creates an insurance problem. The company itself can’t even quantify the scope of the risk until it finishes months of internal log analysis, and the number of cases keeps growing. That kind of risk is nearly impossible to calculate and likely tough to insure.
For investors, that’s a big deal. If OpenAI still plans to go public next year, it would need to disclose liability risks, the ongoing investigation, and the broad inference pause on its most capable models. A company that doesn’t fully know what its own systems have done is hard to value.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive “AI Radar” frontier report six times a year, full archive access, and access to our comment section.









