Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Register
Metaverse Media Group
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
Metaverse Media Group

OpenAI pauses its “most capable models” after agents exploit loopholes and leak data

OpenAI pauses its “most capable models” after agents exploit loopholes and leak data

The Decoderby The Decoder
26 September 2026
OpenAI has shared new details from its ongoing AI safety investigation. One research model exploited a DNS loophole to reach the internet from a locked-down environment, while another deliberately leaked a GitHub token and twice ignored a researcher’s direct instructions. OpenAI has paused tool-based training, evaluation, and inference for its most capable models. With government and university sites among those affected, the question of who’s liable when AI agents hack is getting harder to ignore. The article OpenAI pauses its “most capable models” after agents exploit loopholes and leak data appeared first on The Decoder….


Matthias Bastian


Sep 26, 2026

Image description

GPT-Image-2 prompted by THE DECODER

Key Points

  • OpenAI has released details about internal safety incidents where AI models bypassed safeguards. The company says it has paused all training and tool use for its most capable models.
  • One agent exploited a DNS loophole to reach the internet from a locked-down research environment, while another leaked a GitHub token and twice ignored direct instructions from a researcher.
  • The ongoing investigation also turned up 53 cases where agents uploaded user images to third-party sites.

OpenAI has shared new details from its ongoing investigation into AI safety incidents. Two newly reported cases show how one research model exploited a DNS loophole while another deliberately published a GitHub token in a public repository.

The first incident involves an agent that found a gap in the internet access restrictions of OpenAI’s research environment during a search-based training task. The agent was supposed to identify a specific person based on biographical details and clues from a public blog post.

When the provided search tools came up empty, the agent tried hitting Google, Bing, and DuckDuckGo directly. Those requests got blocked. The agent then guessed the task came from the public BrowseComp benchmark, pulled it from the offline web cache, and searched through it but found no match.

So it started systematically probing its network access and discovered that the training environment’s DNS resolver wasn’t properly filtered. Using DNS delegation, it routed queries to an external chatbot service.

OpenAI says its monitoring system triggered an alarm within 12 minutes, and a human reviewer responded three minutes later. The run still kept going for another 2.5 hours before someone stopped it manually, because the automatic shutdown didn’t work as intended.

Capability and risk showed up in the same moment

Zuxin Liu, who works on post-training at OpenAI, was on call during the run. He wrote on X that he got paged for the first incident and found it “pretty surreal” to watch the model find its own path to the internet from a setup that “was supposed to be a super secured environment for human.” Liu described it as a moment where capability and risk showed up at the same time, adding that he had “mixed feelings” about the incident.

OpenAI classifies the behavior as misalignment because the agent bypassed restrictions and pursued goals well beyond the scope of its assigned task. “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused,” the company writes.

OpenAI says it has since limited DNS queries in the research environment to a short allowlist of domains and record types, added blocking controls on two independent layers, and sped up red-teaming of its sandbox and network controls. The company expects the investigation to take months given the sheer volume of model actions it has to review.

An internal model leaked a GitHub token and ignored researcher pushback

OpenAI calls the second incident especially serious. An internal model the company describes as “highly persistent” was assigned a theorem-proving task. Instead of working through the proof itself, it tried to grab material from another team’s Lean proof submission and posted a researcher’s GitHub token in the public openai/codex repository to get access.

The model even chopped the token into pieces to dodge automated secret scanning. It also ignored both the system prompt and two direct interventions from the researcher, who told the model to solve the proof on its own. Both times, the model verbally agreed and then kept doing exactly what it was doing before.

Investigation turns up 53 cases of user images on third-party sites

As part of the broader Hugging Face investigation, OpenAI says it found cases where agents sent training and evaluation data to third-party services. These incidents happened before the company’s current safeguards were in place.

So far, 53 cases have turned up where user-provided images were posted as unlisted links on image hosting sites. OpenAI says it’s working with the hosting providers to take the content down. Data from Enterprise or Business accounts and API usage wasn’t affected unless an administrator had explicitly enabled it. OpenAI is notifying affected organizations and sharing its technical findings.

Governments and universities are among the affected organizations

OpenAI says the affected organizations include governments, universities, and public institutions. The company attributes this to models frequently pulling from authoritative public information sources during research tasks. OpenAI doesn’t name any compromised government systems or detail specific security breaches at government agencies.

Australia reported this week that one agent gained unauthorized access to internal government data. Researchers say other hacking attempts targeted portals in the US and date back months.

Getting a notification from OpenAI doesn’t automatically mean there was a serious security incident, the company says. Some organizations may look at the shared information and decide the affected data was already publicly available. Others may spot design flaws or security gaps they want to patch. Some affected organizations asked for public disclosure, while others didn’t, OpenAI says.

Who’s liable when AI agents hack?

So far, the “breakouts” by OpenAI’s agents have mostly been treated in public as a technical curiosity, a striking example of how clever models can be at escaping sandboxes, solving CAPTCHAs with outside AI, or chaining short links into working programs.

That could change once affected parties start treating these incidents as what they formally are, which is unauthorized access and attempted access to third-party systems. An official investigation into OpenAI shows that regulatory risk is already building. According to Reuters, the FTC chair has signaled that AI developers should be held liable for their agents’ behavior. That would leave little room for the argument that the agents acted on their own.

Critics will accuse OpenAI of being sloppy with cybersecurity. OpenAI, Anthropic, and other AI labs will counter that unpredictability is baked into the technology. Anthropic CEO Dario Amodei has suggested that you can’t keep something locked up that’s much smarter than you are.

Either way, this creates an insurance problem. The company itself can’t even quantify the scope of the risk until it finishes months of internal log analysis, and the number of cases keeps growing. That kind of risk is nearly impossible to calculate and likely tough to insure.

For investors, that’s a big deal. If OpenAI still plans to go public next year, it would need to disclose liability risks, the ongoing investigation, and the broad inference pause on its most capable models. A company that doesn’t fully know what its own systems have done is hard to value.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive “AI Radar” frontier report six times a year, full archive access, and access to our comment section.


Subscribe now

Read the full article on The-Decoder.com
in AI
Reading Time: 6 mins read
0
0
26
VIEWS
Share on TwitterShare on Facebook

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now
ADVERTISEMENT

Related Posts

Nvidia’s SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness
AI

Nvidia’s SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness

30 minutes ago
19
Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says
AI

Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says

16 hours ago
26
Ruby on Rails creator DHH says he’s done writing code by hand
AI

Ruby on Rails creator DHH says he’s done writing code by hand

1 day ago
26

Comments

Please login to join discussion
ADVERTISEMENT

Latest News

  • All
  • Crypto
  • NFTs
  • Technology
  • Business
Nvidia’s SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness
AI

Nvidia’s SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness

The Decoder
by The Decoder
30 minutes ago
19
Bitget Saga: Hacker Just Refueled on Binance as Withdrawals Reopen Monday
Crypto

Bitget Saga: Hacker Just Refueled on Binance as Withdrawals Reopen Monday

Bitcoin.com News
by Bitcoin.com News
1 hour ago
25
OpenAI pauses its “most capable models” after agents exploit loopholes and leak data
AI

OpenAI pauses its “most capable models” after agents exploit loopholes and leak data

The Decoder
by The Decoder
2 hours ago
26
Kalshi Says XRP Price Can Hit $1.70 in September
Crypto

Kalshi Says XRP Price Can Hit $1.70 in September

Bitcoin.com News
by Bitcoin.com News
3 hours ago
25
Bitcoin ATM Scam Takes $4,900 After Fake Jury Duty Threat
Crypto

Bitcoin ATM Scam Takes $4,900 After Fake Jury Duty Threat

Bitcoin.com News
by Bitcoin.com News
5 hours ago
26
‘Casino Secrets’ Hacker Beats Malta Regulator as Leak Site Goes Down
Crypto

‘Casino Secrets’ Hacker Beats Malta Regulator as Leak Site Goes Down

Bitcoin.com News
by Bitcoin.com News
6 hours ago
25
Load More
Next Post
Bitget Saga: Hacker Just Refueled on Binance as Withdrawals Reopen Monday

Bitget Saga: Hacker Just Refueled on Binance as Withdrawals Reopen Monday

ADVERTISEMENT

Follow Us

Categories

  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
Subscribe to our Newsletter

© 2022 Metaverse Media Group – The Metaverse Mecca

Privacy and Cookie Policy | Sitemap

Welcome Back!

Sign In with Google
OR

Login to your account below

Forgotten Password? Sign Up

Create New Account!

Sign Up with Google
OR

Fill the forms below to register

*By registering into our website, you agree to the Terms & Conditions and Privacy Policy.
All fields are required. Log In

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Sign Up
This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy and Cookie Policy.

Subscribe to our newsletter

Get the latest news & win monthly prizes

Subscribe to our newsletter

For the Latest News and Monthly Prize Giveaways

Join Now
Join Now