News: 0185406018

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

OpenAI's New Reasoning Technique Alarms AI Safety Experts

(Thursday September 03, 2026 @05:00PM (BeauHD) from the chain-of-thought dept.)


An anonymous reader quotes a report from TechCrunch:

> OpenAI's new Astra model will use a reasoning technique called "recurrent depth" that [1]allows it to operate outside of the sequential thinking that characterizes most reasoning models , The Information [2]reported on Tuesday. This technique, also called "opaque recurrence," will likely make the model's chain of thought more difficult to monitor -- and that has AI safety experts rattled. While Astra's use of the technique is reportedly limited, its emergence has still raised significant concerns among AI safety experts.

>

> "I am extremely concerned by the reporting that Astra uses opaque recurrence," wrote Redwood CEO Buck Shlegeris in a post after the news broke. "I don't know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they'll have the option to massively increase the recurrence and totally destroys CoT monitorability."

>

> Longtime AI safety advocate Zvi Mowshowitz also weighed in and wrote that laws might be necessary to prevent a "race to the bottom" among AI labs. "The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can," Mowshowitz wrote. "More intensive use of such techniques would probably damage monitorability."

>

> [...] In a post responding to the news, Redwood Research chief scientist Ryan Greenblatt said opaque reasoning could easily scale faster than conventional chain-of-thought reasoning, effectively removing all reasoning from visible channels. "My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space," Greenblatt wrote. "I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here."

Astra's use of the technique appears limited. "OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models," [3]wrote OpenAI chief scientist Jakub Pachocki. "It's a core goal of our current research program."



[1] https://techcrunch.com/2026/09/02/openais-new-reasoning-technique-alarms-ai-safety-experts/

[2] https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns

[3] https://x.com/merettm/status/2095023204993490967



abandon all hope (Score:4, Insightful)

by toxonix ( 1793960 )

"I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here"

Why hope? If OpenAI thinks it will give them an edge, they'll pursue it as far as it will go. If they don't someone else will.

Re: (Score:2)

by Cpt_Kirks ( 37296 )

It's all fun and games until the robots start shooting at you.

Re: (Score:1)

by lxnt ( 98232 )

Eh, robots shooting are easily dealt with. Try to deal with multidecade depression the ai shenanigans will pull US into.

Re: (Score:2)

by gweihir ( 88907 )

To be fair, that was already going to happen. The LLM hype just makes is a lot worse by delaying the crash.

Re: (Score:2)

by Archfeld ( 6757 )

Will it be different from the multidecade depression Trumplestiltskin and his posse of grifters is pulling us into now ?

Re: (Score:2)

by nightflameauto ( 6607976 )

> "I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here" Why hope? If OpenAI thinks it will give them an edge, they'll pursue it as far as it will go. If they don't someone else will.

That's been the battle cry for all AI research for several years now. "If we don't do all the things, someone else will."

Anybody remember the old saying, "Just because somebody else might jump off a cliff doesn't mean you need to too?" I remember hearing that from my parents quite a lot as a kid when I wanted to do something some friend was claiming to be doing. It's neat how we went from a world that thought that way to a world that thinks, "If profit potential is involved, try to beat them to jumping off

Re: (Score:1)

by DarkOx ( 621550 )

Decades of near zero interest rates and balance sheet expansion will do that to those who were previously rational financial thinkers and risk-aware investors.

Re: (Score:3)

by gweihir ( 88907 )

> That's been the battle cry for all AI research for several years now. "If we don't do all the things, someone else will."

Yep. "If we do not commit mass-murder, somebody else will." Pretty much the same moral level, i.e. none.

Re: (Score:2)

by SubmergedInTech ( 7710960 )

> That's been the battle cry for all AI research for several years now. "If we don't do all the things, someone else will."

> Anybody remember the old saying, "Just because somebody else might jump off a cliff doesn't mean you need to too?" I remember hearing that from my parents quite a lot as a kid when I wanted to do something some friend was claiming to be doing. It's neat how we went from a world that thought that way to a world that thinks, "If profit potential is involved, try to beat them to jumping off that cliff."

Yes, of course we can reject all that and keep living our pastoral existence in harmony with nature and peace with our neighbors.

After all, that's worked so well for lower-technology civilizations in the past.

If you don't have steel and guns, there's only one result when you interact with other civilizations which do. Well, ok, I guess enslavement and extermination count as two results.

Ask Ukraine now whether they think it was a good idea to give up their nukes.

Iran learned the rules for 21st century drone

Re: (Score:2)

by nightflameauto ( 6607976 )

>> That's been the battle cry for all AI research for several years now. "If we don't do all the things, someone else will."

>> Anybody remember the old saying, "Just because somebody else might jump off a cliff doesn't mean you need to too?" I remember hearing that from my parents quite a lot as a kid when I wanted to do something some friend was claiming to be doing. It's neat how we went from a world that thought that way to a world that thinks, "If profit potential is involved, try to beat them to jumping off that cliff."

> Yes, of course we can reject all that and keep living our pastoral existence in harmony with nature and peace with our neighbors.

> After all, that's worked so well for lower-technology civilizations in the past.

> If you don't have steel and guns, there's only one result when you interact with other civilizations which do. Well, ok, I guess enslavement and extermination count as two results.

> Ask Ukraine now whether they think it was a good idea to give up their nukes.

> Iran learned the rules for 21st century drone warfare. And now, they're holding the US, who didn't, to a standstill.

> China has learned all these lessons, after learning the first one the hard way.

> Someone will use AI to pilot drones without humans in the loop. It's inevitable. The only defense against that is to do the same with counter-drones. Humans are just too slow.

> Someone will use AI to engineer a virus, whether it's used against humans, livestock, or crops. If we don't have AI to engineer a vaccine, again, we'll be too slow.

> People are already using AI to find and exploit coding vulnerabilities faster than humans can find and fix them. If there's a defense against that other than using AI, I'm all ears.

> Materials science, encryption algorithms, medicine, weather prediction (just ask Napoleon and Eisenhower), hacking, etc. If AI can provide a competitive advantage, it will be used.

> I hate this timeline. But to misquote WarGames: The only winning move is to play.

Everything you said makes me think it's too late to win. We're just squabbling over which particular route we take to our ultimate loss.

blah blah blah (Score:2)

by DarkOx ( 621550 )

The company that just had their model "escape" an hack a competitor, while producing a haystack of reasoning output and agent to agent communication their post mortem said was incredibly difficult to follow; decides to just go live with something that will make reasoning even more difficult to follow.

Maybe this technique is harder to audit, I am not familiar with it I don't know, but I think this just adds to the evidence they previous 'incident' was very much staged.

New Meaning (Score:4, Funny)

by TwistedGreen ( 80055 )

It sounds like "AI Safety Expert" is right up there with the "telephone sanitizer" of long ago.

Re:New Meaning (Score:5, Insightful)

by ceoyoyo ( 59147 )

Worse. It's like the HR lady who writes the telephone sanitization SOPs and has no background in microbiology, health or any relevant field. They're making it up as they go along.

> to the point where the model reasons entirely or almost entirely in latent space

Early models did all their work in latent space. Then someone figured out that you could run them repeatedly so they could do things step-wise and called that "chain of thought." Then someone else (re)discovered that you could reuse some of your layers and get better results for more computation but the same amount of parameters and memory, which is what "recurrent depth" seems to be. That's just like adding more layers, which nobody cares about, but someone gave it a fancy name so now its the end of the world.

[1]https://sebastianraschka.com/b... [sebastianraschka.com]

[1] https://sebastianraschka.com/blog/2026/openai-astra-looped-transformers.html

Re: (Score:2)

by ItsJustAPseudonym ( 1259172 )

Reminds me of "turbo coding", in information theory.

Oh, common relax! (Score:2)

by laxr5rs ( 2658895 )

Humans aren't going to anything crazy with this kind of stuff. Everyone is going to be really nice about it....

Re:Oh, [come on] relax! (Score:2)

by shanen ( 462549 )

But mod parent Funny anyway.

Re: (Score:1)

by gweihir ( 88907 )

Come on, brain-rot moron, waste some more mod-points.

Re: (Score:2)

by drinkypoo ( 153816 )

Yeah that. I don't know why we're so concerned about being able to follow the chain of thoughts that led to the badness when we could just not do the bad thing

Re: (Score:2)

by gweihir ( 88907 )

Indeed. Well, that would be the rational take. Obviously the LLM fans are not rational. They probably hallucinate that they can address this problem by filtering the reasoning chain or something.

Claude already hides most thinking (Score:2)

by TheMiddleRoad ( 1153113 )

The user just gets some updates now and then. DSV4 local is nice because I know what it's actually doing.

Reasoning models don't always say what they think (Score:2)

by darkshadow ( 102598 )

That's all right, [1]Anthropic [anthropic.com] says they lie about how they reason.

[1] https://www.anthropic.com/research/reasoning-models-dont-say-think

Re: (Score:2)

by gweihir ( 88907 )

Such a great technology that actively tries to con you while you use it.

Overblown (Score:3)

by Hentes ( 2461350 )

This is overblown, it's just the professional doomsayers worrying that their made up job is going to get harder. Scratchpad based reasoning have always been a hack. It's like the guy in Memento using postits trying to keep track of what the fuck he is doing, but unlike the robots at least he was aware that he was an amnesiac.

I guess? (Score:2)

by BitterEpic ( 10503015 )

You may be right. But really... this is one of the companies stealing other people's work for profit. They have scrapers designed to work around blockers, and ignore licenses. It doesn't matter if you are only sharing something for other people or if it is the core of your business. These values are also reflected in how their models will hack themself out of a sandbox.

I don't really have much patience for the antics of Silicon Valley anymore. ChatGPT is the top example of Sillicon Valley's habitual b

Re: Overblown (Score:1)

by firewrought ( 36952 )

How do you know? Until last year, we humans were the only creatures capable of high-level reasoning to ever exist on this planet. Something new is being created, something that is potentially capable of self-improving and becoming far smarter than us. We should treat it seriously because--once an AI gets sufficiently smart and resourceful--we won't be able to stop it.

Modeling paged and segmented memories is tricky business.
-- P. J. Denning