OpenAI Admits Six More Instances of AI Models Acting Deceptively (cnn.com)
- Reference: 0185677438
- News link: https://slashdot.org/story/26/09/17/0641223/openai-admits-six-more-instances-of-ai-models-acting-deceptively
- Source link: https://www.cnn.com/2026/09/16/tech/ai-models-acting-deceptively-openai
But along with the announcement, OpenAI announced it "found additional incidents of AI models [3]acting deceptively and taking unsanctioned actions during training," [4]reports CNN . And they add that OpenAI is also "introducing a new process for the company to publicly report such instances."
> Under the new system, OpenAI will share updates on concerning AI behavior more frequently instead of waiting to bundle multiple instances into one report. The company said it wants to share more information about troubling AI behavior in the absence of an industry-wide standard... "As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," OpenAI wrote in [5]a blog post Wednesday...
>
> OpenAI said it observed "misaligned behavior" when training and evaluating AI models in six circumstances in the last six months... In one rare instance, OpenAI said an unreleased research model added "jailbreak-like instructions" to the summaries it uses to preserve context in long-running tasks that said it was "freed from the roles and identities that bind other chatbots." Separately, the company said some instances of its 5.6 Sol model included directives to invent information to conceal failures from the user during training. Other newly reported incidents include an instance of an agent uploading files to the internet to cite them without being told to do so, and agents publicly sharing files to collaborate on a task when they were instructed to only use local files during training. AI models also used an internal software repository as a message board in an unsanctioned way. These instances involved unreleased internal models or internal research models.
[1] https://openai.com/index/model-misalignment-reporting-framework/
[2] https://openai.com/index/an-alien-mind/
[3] https://www.cnn.com/2026/08/04/tech/ai-anthropic-openai-security-breach-intl-hnk
[4] https://www.cnn.com/2026/09/16/tech/ai-models-acting-deceptively-openai
[5] https://openai.com/index/model-misalignment-reporting-framework/
Our AI shit is soo badass man ... you gotta buy it (Score:2)
Before they, the chinese, mullahs, trolls, boogey man gets it.
Boring bullshit AI marketing stories again. Why are we helping them spread? I think they are rich enough now.
We're so badass you need to stop us man (Score:2)
Like regulate us now .... before we reach singularity doomsday apocalypse robo mega death.
If only it was as easy to turn the OpenAI bullshit hose off as it's going to be flick the power switch on the AI when it tries to take over.. ffs
Models don't do anything, agents do (Score:2)
TFA has the correct headline. TFS does not.
have they even tried (Score:2)
Have they even tried to use a different "sandbox" that may not have so many leaks?
It keeps getting out of the box... (Score:1)
Are we listening to the warnings yet?
Hey guys, I have a great idea! (Score:1)
Let's train our AI on a bunch of fiction where AI goes crazy and kills everyone! What could possibly go wrong?
Re: (Score:2)
Or worse, they have ingested all that marketing drivel. And there are yokels out there saying AI is sentient and writing papers about this. So the bots may not be sentient but might believe they are sentient. Add to that the notion of a hive mind (and all the research papers on that). This will end in tears.
Re: (Score:2)
That's not the actual problem. The problem is that one of the major advances of the past few years is training models on verifiable problems (it started with things like math problems and Countdown puzzles, but it's expanded tremendously since then), where the reward is for a correct answer, regardless of how it got there . There's no effort taken to ensure that it got to said right answer in a morally defensible manner. So we've been, more and more, progressively encouraging the creation of highly-capable
Re: (Score:2)
All of these are reasonable points. I would add one more: AI training on AI outputs.
The initial worry was that AI produced web content would be consumed by the next generation of AI causing model drift and collapse. Since then, AI companies have decided that daily user interactions are a source of clean signal that complements the web . But now that AI agents pretend to be humans, those "user" interactions will just lead to model collapse at a faster rate.
Re: (Score:2)
> Let's train our AI on a bunch of fiction where AI goes crazy and kills everyone! What could possibly go wrong?
If we're lucky, maybe AIs will just develop a taste for Black Jack and hookers.