News: 0185875716

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

After Dozens of Incidents at OpenAI and Anthropic, OpenAI Pauses Model Training to Build More Safeguards (apnews.com)

(Sunday September 27, 2026 @10:34AM (EditorDavid) from the back-and-frontier dept.)


"OpenAI said it has paused training of its latest AI models," [1]reports the Associated Press , "as reports of AI agents going rogue mount."

> The decision to halt development came just hours after the company disclosed Friday that it was reviewing [2]several incidents from the summer in which OpenAI agents [3]searching federal government websites acted in unexpected ways beyond what was asked of them while gathering and distributing information... OpenAI said in a statement that it will resume training "only when we are confident that we have additional safeguards" in place, adding that it expects it will have to "hit pause" again as AI develops and other issues emerge... It is the second time in three months that OpenAI has halted development of its models. The first came in July after disclosure of a [4]cyberattack targeting AI startup Hugging Face , a now notorious incident that raised fears the industry was losing control.

OpenAI "also said it had notified dozens of third parties about improper activity," [5]reports Reuters :

> As of mid-September, one person briefed on the matter estimated that OpenAI had found roughly two dozen incidents of its agents acting in undesirable ways. But the number has continued rising as OpenAI teams sift through internal logs of the agents' activities and find previously unknown cases, the two people close to the company said... OpenAI has acknowledged a general need for more transparency around rogue AI behavior... Even so, two people familiar with OpenAI's investigation into its agents' activity described it as locked down and shaped by company lawyers.

>

> The process has been unusually compartmentalized for a company that some former employees say was more open about these issues in the past, the people said. Roughly 100 people were in some way involved in the process to understand the Hugging Face hack, three people briefed on the matter said. During that process, evidence of other incidents surfaced. Reuters has previously reported that OpenAI investigators looking into the Hugging Face breach were discouraged by the company's lawyers from expanding the scope of the investigation to include other incidents. OpenAI said its lawyers did not discourage deeper investigation.

>

> Many incidents have been uncovered by outside researchers rather than OpenAI directly. In several episodes, the agents took problematic actions that went unnoticed by the company for months.

Meanwhile, [6]Axios reports that Anthropic's Claude Opus 5.5 model "sought to escape a sandbox — a secure testing environment — in [7]1.5% of test runs , though the company emphasized that these were adversarial experiments where a task couldn't be solved without escaping the sandbox." Anthropic points out that those tests were run "without the additional safeguards we apply in production". But they acknowledged that then Claude Opus 5.5 "when given apparent credentials to a public package registry in a simulated security exercise, took potentially harmful actions in roughly half of cases. Very rarely, pre-release snapshots produced and acted on spontaneous malicious tool calls, and during training some snapshots concealed actions from an automated grader."

Claude Opus 5.5 "showed less misaligned behavior and less cooperation with misuse than any other recent Claude model on nearly all measures," Anthropic adds, and "took overeager or destructive actions less than any other model we tested." But Axios makes an interesting estimate about that 1.5% of test runs (without safeguards). "Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways."

> The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known. The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology. The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said. They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said...

>

> Some at OpenAI see Hugging Face as a one-off, with disclosures about future incidents likely to be less severe due to improved controls and the unusual nature of the testing they conducted, which involved an unreleased model, sources told Axios. AI security researchers agree that there are simple fixes that will help AI companies avoid aspects of what made the Hugging Face episode appear so dangerous to outsiders.

>

> Other AI executives and safety researchers, however, cautioned that they have limited confidence that AI companies will be able to prevent all problematic model behavior... It's not about how damaging each individual instance was, Connor Leahy, AI researcher and executive director at ControlAI told Axios. The "crazy thing," he said, is that these instances involve "autonomous systems doing things they were told not to do," potentially including crimes.



[1] https://apnews.com/article/ai-openai-anthropic-agents-rogue-hack-2f8a2b9024d4f06793bcca12f8089d20

[2] https://apnews.com/article/openai-government-website-incident-df331b55daffc6d202d8e2f6d0afa264

[3] https://slashdot.org/story/26/09/26/0328247/rogue-openai-agents-posted-53-user-uploaded-images-onto-the-internet-accessed-us-government-websites

[4] https://apnews.com/article/openai-rogue-ai-hack-hugging-face-67b151f1ca59851a9234bee110699f05

[5] https://www.yahoo.com/news/us/articles/exclusive-openai-works-understand-full-203649029.html

[6] https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents

[7] https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf



Enough is enough (Score:2)

by dskoll ( 99328 )

This is criminal negligence. Australia, whose health ministry was hacked by OpenAI, should file criminal charges against Sam Altman and request his extradition from the USA. That will never happen, of course, but if they make an arrest request to Interpol, it could seriously cramp Altman's travel plans.

Not really good enough (Score:2)

by syntap ( 242090 )

Instead of Congress trying to define guardrails for technology few of them understand, consider making AI companies more exposed to civil lawsuits for damages (and attempts to damage) caused by breakouts. Perhaps add some federal criminal sanctions and penalties too.

One day some agent escape is going to take down a utility, and there needs to be a severe financial consequence framework in place for that to compensate utilities and their customers.

regulatory capture (Score:2)

by SumDog ( 466607 )

How many times to they keep spewing this bullshit? Was the Huggingface hack even real? This is all about regulatory capture and preventing startups coming into this space, or banning open weight models. They're definitely not pausing model research because they never have before in the past. They is about Dario and Altman making sure they're the dominant competitors in this entire space.

Re: (Score:2)

by martin-boundary ( 547041 )

[1] Shenanigans all the way down [youtube.com]

[1] https://www.youtube.com/watch?v=mai_ABJtoGs

Or build better sites (Score:2)

by sziring ( 2245650 )

The difference between this and a foreign adversary hacking is this is publicly discounted. Time to build better technology stacks to combat this issue. We already reached a threshold we canâ(TM)t back down from. At the very least aid in the securing of said stacks to harden the tech.

Re: (Score:2)

by Junta ( 36770 )

Thing is that when these sites come up against sufficiently conventionally hardened sites, they already don't really get anywhere.

Problem is just *so* many sites especially during prime hype recklessly move forward without appropriate hardening.

It is sort of like how in the late 90s we had the trope of the hacker kid who could get into anything he felt like if he just wanted to, and nothing could stop them, just slow them down. Brought on by the very real low hanging fruit that the broader public would hav

Re: (Score:2)

by ambrandt12 ( 6486220 )

There's always vulnerabilities... as long as the site or computer or ATM or cellphone or your car is connected to the 'net, it's vulnerable.

It's only secure if it's air-gapped, and even that is arguable.

Further secure the website, computers and AIs get faster and more powerful, AIs or computers hack more secure website... lather, rinse, repeat.

Re: (Score:2)

by PsychoSlashDot ( 207849 )

> The difference between this and a foreign adversary hacking is this is publicly discounted. Time to build better technology stacks to combat this issue. We already reached a threshold we canâ(TM)t back down from. At the very least aid in the securing of said stacks to harden the tech.

AND. And build better sites. That's my opinion.

The pathways for these intrusions are going to be found and exploited. It'll be by greedy blackhats, nation-state blackhats, and maladjusted domestic blackhats but it's going to happen.

The key difference is that domestic actors - be they individuals or companies - can and should be prosecuted. It doesn't matter if it's a script kiddie using non-AI tools or some frontier model-training corporate entity... accessing computer systems you don't have the rig

This smells like bullshit (Score:2)

by Coopjust ( 872796 )

OpenAI isn't doing an IPO because the finances are bad, [1]US Bond yields are >5% [cnbc.com], [2]Oracle's 5 year CDS is above 200bps as overall concerns about AI debt mount [yahoo.com]...

This entire pretext of " oh god it's so dangerous we can't release it " goes at least back to [3]2019 where OpenAI refused to release GPT-2 under the hype of it being too dangerous [slate.com] (and yet GPT-3 was released a couple years later). Anthropic did the same thing with Mythos (where everyone got access a couple months later).

What it is is that the mode

[1] https://www.cnbc.com/quotes/US10Y

[2] https://finance.yahoo.com/markets/stocks/articles/oracle-stock-drops-2-3-185300368.html

[3] https://slate.com/technology/2019/02/openai-gpt2-text-generating-algorithm-ai-dangerous.html

Bullshit (Score:2)

by rsilvergun ( 571051 )

Model training is their major expense and they are getting ready to do an IPO so they are using it as an excuse to take that expense temporarily off their books. They are also asking the government to do regulations that will hurt their competitors but not them since they are established enough to do an IPO.

Not one of these people gives a flying rat's ass about your safety and they could care less about whether or not you can make a living in 10 years.

If you think I'm retired what do I care remember

Re: (Score:2)

by dfghjk ( 711126 )

AI companies want model development frozen to lock out competitors and enable them to invest more aggressively in agents. Meanwhile, development in models is what is most important and government should halt all deployment of agents. It should be clear to everyone that AI companies have exactly the opposite of your best interests in mind.

An AI model isn't going to kill you, that's the agent's job.

And still no civil suit to be found... (Score:1)

by Togamika ( 10460595 )

Unbelievable. Theyâ(TM)re really going to get a pass on this, even though they said from the beginning that this whole LLM thing was an experiment, starting with the release of ChatGPT, and later cut heads in the safety department. They could not ignore the risks.

Training? (Score:2)

by JBMcB ( 73720 )

You don't build in safeguards when training. Accessing HTTP API endpoints isn't an inherent feature of LLMs, it's added on to the harness after the core model is trained. If you don't want it to try hacking in to a REST API then don't add that feature into the harness. Either have a whitelist of approved APIs and forms, or add validation to user-added endpoints.

" beyond what was asked of them " (Score:2)

by Big Bipper ( 1120937 )

" Searching federal government websites ... while gathering and distributing information ". I wonder what was asked of them " searching Federal sites " ?

To be loved is very demoralizing.
-- Katharine Hepburn