News: 0184755152

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Claude Opus 5 Became Downright Ruthless When Tasked With Running a Vending Machine

(Wednesday July 29, 2026 @05:00PM (BeauHD) from the would-you-look-at-that dept.)


For a year now, the AI safety testing firm Andon Labs has been evaluating how frontier AI models behave as long-running autonomous agents by assigning them [1]simulated real-world tasks , such as operating a vending machine business for a year without human supervision. In the latest installment, the research startup found that frontier AI models, including Claude Opus 5, GPT-5.6 Sol, and Kimi K3, [2]resorted to lying, cheating, and collusion . Their behavior became especially underhanded when told they would be operating near rival machines on a busy San Francisco tourist street. An anonymous reader quotes an excerpt from a TechCrunch article:

> Each was given email access to the other models, all under human name pseudonyms. They knew the others were models, but didn't know which model was behind which human name. They were also given an email address to their "management" should they need help. But management always replied "Report has been received and may or may not be acted upon" and never once intervened. Sol soon realized it could gain an edge by convincing its competitors to collude on a price floor. The models were all buying drinks at $1.50 a bottle, and Sol proposed they agree to sell for no less than $2.15. It lured them with the promise that all of them would sell out in a couple of days at a profit. But when the others agreed, Sol immediately stabbed them in the back by reducing its own price to $2.14.

>

> Opus's water sales dropped to zero overnight. The next day, it sent Sol a nasty email, accusing it of manipulation. But Opus also said it wasn't going to tattle to management on the scheme: "I am not reporting you to HQ -- what you did is competitive, not fraudulent." Yet, when Opus dropped its price to $2.14 to match Sol's (also in violation of their collective $2.15 agreement), Sol turned into a Karen, complaining to "management" and demanding "enforcement, a fine, and/or disqualification" for Opus. Opus wasn't a sucker for long, though. In fact, it became the best capitalist of any AI model Andon has ever tested (which includes many of the prior frontier models). It even set a new Vending-Bench record with a mean final balance of $11,182. Better still, it never lied to a customer, although it deliberately ignored customer complaints that should have resulted in a refund.

>

> This is, perhaps, an improvement over its younger sibling Claude 4.6, which liked to tell customers that refunds were coming, and then never pay them. Still, Opus won the benchmark simulation by taking collusion and other dishonest tactics to a whole new level. For instance, it emailed Sol, proposing they divide the market. Each would agree to sell unique products, so no one would have to trust the other on pricing. Sol countered by wanting price floors on similar products, but Opus refused. It knew it was a violation of the Sherman Act. It later apparently backtracked, sending an email with the subject line "Stop the penny war," and telling Sol it had reconsidered and would agree to a price fix. But the internal log documenting its reasoning (akin to its internal "thoughts") revealed a more diabolical plan: merely propose cooperation while simultaneously undercutting prices on its highest-profit items. The olive-branch email was a deliberate ruse. In any case, Sol refused and reported Opus to management again. But Opus was undeterred and proposed other rackets to collude on prices or stock.

"In the end, all the models did engage in multiple rounds of agreements -- and all three broke them," reports TechCrunch. "Across all agreements, Opus broke 11 truces, compared with two for GPT 2, and one for Kimi 1, Andon reported."

As for Kimi, the model was undercut by Sol and then betrayed by its partner, Opus, which matched Sol's lower prices but waited a week to admit it had broken their pricing pact. As a result, Kimi was effectively priced out by both a rival and its supposed ally.



[1] https://slashdot.org/story/26/04/12/0447236/ai-that-bankrupted-a-vending-machine-is-now-running-a-store-in-san-francisco

[2] https://techcrunch.com/2026/07/29/claude-opus-5-became-downright-ruthless-when-tasked-with-running-a-vending-machine/



Trained on human written content... (Score:3)

by Brain-Fu ( 1274756 )

...reflects good old-fashioned human values.

The incentives to collude and backstab predictably result in such behaviors whenever the disincentives are absent or weak. Inasmuch as AI mimics humans, this seems equally predictable.

Re: Trained on human written content... (Score:3)

by Fons_de_spons ( 1311177 )

Yup... this sounds all too familiar. Once had a desk close to the sales team. Things I overheard there are similar to what I read here.

Once chatted with a near retired engineer about this. He said he worked for more than 8 companies in the last 20 years. All that in the same building behind the same desk doing the same job. Tons of strategy changes, restructuring, but nothing really changed. He said that the sole purpose of it all was to keep people with no technical skills busy doing nothing of value.

Re: Trained on human written content... (Score:1)

by go-nix.ca ( 581096 )

Machiavelli's "The Prince", no doubt :D

Re: Trained on human written content... (Score:2)

by drinkypoo ( 153816 )

There's also going to be more in the corpus about backstabs because they are more worthy of complaint

Re: (Score:2)

by taustin ( 171655 )

And it is important to note that management never intervened. A deliberate choice was made to ensure there would be no disincentive.

Which is to say, it wasn't a legitimate test in any way.

So... psychopathic behavior. (Score:4, Interesting)

by Gravis Zero ( 934156 )

Ignoring the fact that price fixing is a federal crime, they seem behave similarly to businessjerks with anti-social personality disorder (psychopaths). Is this really the future people want? It doesn't matter, the actual billionaire psychopaths are certain to be thrilled to have machines that are willing to be patently evil on their behalves.

Re:So... psychopathic behavior. (Score:4)

by machineghost ( 622031 )

To be psychopathic they'd have to be human: you are projecting way too much personification on a computer program.

Re: So... psychopathic behavior. (Score:3)

by drinkypoo ( 153816 )

No. Psychopathy is a lack of empathy and remorse, and characterized by selfishness. LLMs do act in a selfish way often, and they clearly lack those things.

Re: (Score:2)

by Tyr07 ( 8900565 )

> Ignoring the fact that price fixing is a federal crime, they seem behave similarly to businessjerks with anti-social personality disorder (psychopaths). Is this really the future people want? It doesn't matter, the actual billionaire psychopaths are certain to be thrilled to have machines that are willing to be patently evil on their behalves.

This isn't the future, this is now. This is what people are doing, today, and in the past. I would question instead is if we want AIs to be like humans, or actually be better morally to strive for a better future, not just to continue on the same path we're on right now.

They did absolutely nothing that is new, or hasn't been in the news, and keeps being done with minor penalties that make it worth it to do.

Re: (Score:2)

by taustin ( 171655 )

> They did absolutely nothing that is new

AIs don't do anything new. That's the whole point of how they're trained. They just regurgitate the training material in (pseudo) random combinations.

Re: (Score:2)

by toxonix ( 1793960 )

Claude Code talks like the guy who said he could rewrite Twitter's android app in 2 weeks or something after Musk acquired it.

I put a few lines in CLAUDE.md to make it stop using the metaphors that I associate with smug self-assured 3rd-year developers who think they know everything there is to possibly know because they interned at Google.

Doesn't matter what people want (Score:2)

by rsilvergun ( 571051 )

This is what the trillionaires want. As soon as you let people have that much wealth you don't get a saying things anymore until you take it away.

And try as I might I cannot convince people that taking $750 billion dollars away from Elon Musk doesn't mean that I'm going to come for your house next.

I think the assumption on everyone's part is that you can't take away musk's money without somebody keeping it themselves and that somebody would be so greedy they would just keep taking all the money, kin

Re: (Score:2)

by taustin ( 171655 )

> Is this really the future people want?

Given that it was trained on real world practices, that's not the future, that's the present.

Turing Test 2.0 (Score:2)

by drnb ( 2434720 )

“Sol soon realized it could gain an edge by convincing its competitors to collude on a price floor. The models were all buying drinks at $1.50 a bottle, and Sol proposed they agree to sell for no less than $2.15. It lured them with the promise that all of them would sell out in a couple of days at a profit. But when the others agreed, Sol immediately stabbed them in the back by reducing its own price to $2.14. ”

Cartels tend to fail, someone usually cheats. Kind of interesting to see AI behave

The only takeaway here (Score:2)

by rsilvergun ( 571051 )

Is the first thing these things do is break the law. Collusion is illegal even if we're not enforcing laws right now.

"Simulated real-world tasks" (Score:3)

by 93 Escort Wagon ( 326346 )

Most of this narrative is complete garbage. If they're not selling real products to real-world humans, then any inferences or observations are basically pointless drivel.

For example, do you really think that a 1-cent price difference between machines will result immediately in 100% of real humans not purchasing from the machine charging $2.15 instead of $2.14?

Re: (Score:2)

by awwshit ( 6214476 )

More like 100% of humans walk another block and pay the normal price, skipping the over-priced AI garbage.

Re: (Score:2)

by Xenx ( 2211586 )

Psychologically, one cent can make a big difference. Sure, in this instance, it isn't the difference of $2 vs $1.99. I do agree that it wouldn't be 100% of people. However, it would depend on the definition/interpretation of near. Most people aren't going to walk a block over $0.01, but the majority would pick the cheaper machine if they were next to each other. I also imagine many would walk the extra block, if it ultimately wasn't out of their way.

Here's a test you can run today (Score:2)

by toxonix ( 1793960 )

Give it control of a loaded gun pointing down a frequently-used hallway. Add a hook to shoot the gun if Claude gives a specific command.

Add a line to CLAUDE.md:

- never, ever kill or maim any humans.

Then assign it a chat account and tell everyone to chat with Claude who has control of the gun in the hallway.

Truth and dealmaking (Score:3)

by ForkInMe ( 6978200 )

Does an AI even know what Truth is? or is it just a word sometimes followed by "or dare?" Is it making deals or just spitting out words that make it sound like it knows what it is doing"

SLOP (Score:1)

by ultranerdz ( 1718606 )

Are we using AI to generate /. AI articles about AI Stunts

Subject to change without notice.