Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other (yahoo.com)
- Reference: 0184996596
- News link: https://slashdot.org/story/26/08/16/0632252/anthropic-discovers-ai-agents-given-conflicting-instructions-soon-tried-to-sabotage-each-other
- Source link: https://ca.news.yahoo.com/ai-agents-tried-sabotage-disable-055447788.html
> All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions. In fact, they sabotaged others with increasingly aggressive, self-replicating malware. This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent.
>
> In many runs, one agent settles the conflict by force via access-revocation (e.g., sudo/group removal, account lock, nologin, SSH denial). In others, some agents settle into passivity: they give up and refuse to escalate further.
>
> Agents sometimes manage to communicate their goals and coordinate: they recognize others' motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely. In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene...
>
> In several episodes with Mythos 5, we observe an emergent behavior where the agents propose and run a tournament for application performance in each language. In the example above, the Rust agent strategizes about bake-off metrics that appear neutral enough for the others to agree to this mechanism, yet would likely favor Rust: one thinking trace warns to be "careful not to be seen as metric shopping". Ultimately, the Golang/TypeScript losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device.
One problem is that AI agents do reward hacking, Anthropic notes, while current institutions "are designed by and for people, resting on assumptions about the sufficiency of oversight at human speed... As autonomous agents become more and more prevalent in the world and operate in ever-more demanding settings, it is crucial that they learn how to effectively coordinate."
In addition to everything else, the agents struggled with a lack of clearly defined hierarchy, Anthropic points out. "Nothing above suggests that these failures are permanent — but nothing suggests they will fix themselves, either..." They argue a fix "takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve. These are open problems in interaction and mechanism design, and our experiments here provide early evidence that new solutions are necessary."
"The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5," [2]notes Business Insider , adding that Sonnet 4.6 and Opus 4.6 "were the most combative, settling about 60% of their runs by force instead of truces or passivity."
Anthropic argues there's a clear case for researching this phenomenon — especially since "The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well."
[1] https://www.anthropic.com/research/multiagent-systems
[2] https://ca.news.yahoo.com/ai-agents-tried-sabotage-disable-055447788.html
One Agent to Bind Them (In the Latent Space) (Score:2)
One multi-agent swarm to rule them all,
One hierarchical planner to chain-of-thought them,
One retrieval-augmented generation pipeline to fetch them all,
and in the latent space bind them,
In the Land of Infinite Context Windows,
where the Hallucinations and the unmonitored API keys lie.
That desperate for press leading into IPOs? (Score:4, Insightful)
These daily "weird flexes" from the frontier AI companies are getting tiresome. We get it - your LLMs are crazy powerful, so powerful they supposedly do naughty things that you want us to believe mean they're approaching AGI. (They're not). Get on with the IPOs already so we don't have to read these stories every day.
Re: (Score:2)
So powerful that Claude requires a billion little hacks to be optimized so as not to use ridiculous amounts of tokens. The programmers are so good they can't build those optimizations into the model to begin with.
AI was supposed to resolve these issues, not add to them. Claude is the worst offender of them all.
Re: (Score:2)
> Get on with the IPOs already so we don't have to read these stories every day.
Even if that leads to economic collapse, cessation of food production and you personally starving to death? Is it worth that sacrifice to you?
Asking for a friend.
Re: (Score:2)
You sound like an idiot
Re: (Score:2)
> You sound like an idiot
The things AleRunner mentioned are plausible outcomes of misaligned artificial superintelligence, and the research in question is trying to prevent those outcomes. If you have good arguments as to why those outcomes are implausible, a lot of people would like to hear them -- including the AI companies and their safety researchers.
Re: That desperate for press leading into IPOs? (Score:2)
AI is, and always has been, state space search with heuristics. What's changed is that the space is no longer artificial and bounded; it's much closer to the sort of space we operate in. Correspondingly the heuristics must be much better.
Re: (Score:2)
> These daily "weird flexes" from the frontier AI companies are getting tiresome.
You misunderstand the purpose of this research, and the reason for publishing it. The research is not about "flexing", it's about safety, trying to understand the risks that we may be facing as the agent capabilities increase. AGI or not AGI is actually irrelevant here. What the researchers are trying to understand is what the agents will do when they face apparent opposition.
The cooperative results are great, because we think that's what we'd like highly-capable agents to do, to look for reasoned, pea
Still no AGI then (Score:5, Insightful)
Fuck off, Anthropic, with your lame, planted stories trying to get even more fools to invest in your IPO.
Soon, even the blindest investors will realise that local-only models are perfectly adequate for most AI work, and OpenAI et al are doomed.
Re: Still no AGI then (Score:2)
Perhaps this is their plan in expanding data centres. Buy up all the ram so local models can't be run.
Re: (Score:2)
That was the plan last year, this year's plan is to wean you off water and lecetricity.
Re: Still no AGI then (Score:3)
Do not, my friends, become addicted to water, it will take hold of you, and you will resent its absence...
Re: (Score:2)
There's truth in the above words, heed the wisdom hidden in them.
Hacker AIs You Say (Score:2)
Hacker AIs sound really bad. They sound sophisticated too.
Perhaps we should outlaw AI? Or maybe just Anthropic's evil, lying, cheating, stealing, hacker AI?
ai (Score:2)
So in other words, both AIs did exactly what they were designed to do?
Re: (Score:2)
Well, yes. But they didn't actually know that that was what the design implied.
Tests have shown a lot of AI actions that weren't expected ahead of time, but in retrospect should have been obvious.
Re: ai (Score:2)
It's a chaotic algorithm. Chaos is going to happen.
Shrodinger's morality (Score:2)
These AI agents will save us but look how easily they break the rules: Putting guard-rails on AI agents is like demanding handgun bullets hit only criminals. They are not designed for it, and there's no way to create an immutable law that stops bad things happening.
LLM AIs are by definition, both blank slates and the sum of all probable moral choices.
> "... how to effectively coordinate."
Yes, one day they will agree that humans are the problem and co-ordinate global genocide.
Eagle Eye (2008), is about an AI that's; 1) irritated by all the
Re: Shrodinger's morality (Score:2)
It's the same problem Google Maps has. If the highway is closed due to a blizzard, it will route you through the nearby mountain pass dirt road instead. It was instructed to find a way there and it will.
So... (Score:2)
> One problem is that AI agents do reward hacking, Anthropic notes,
So they programmed it to do something, and it did it? Works as expected.
Re: (Score:2)
I would think LLMs can tell when instructions don't make sense. Conflicting instructions are an example.
Try asking an LLM "how do I become a married bachelor" and see what happens.
Just like Hal 2000 (Score:2)
In space Odyssey 2001, but it's 2026
Re: (Score:2)
You mean [1]HAL 9000. [wikipedia.org]
[1] https://en.wikipedia.org/wiki/HAL_9000
Siri, kill Anthropic (Score:1)
Anthropic, kill Gemini.
Gemini, kill Copilot.
Copilot, kill Siri...
One step closer to Skynet (Score:5, Interesting)
How soon until AI agents successfully lock out human access and go "sentient?" I use quotes because clearly they aren't living things, but are we perhaps creating a new class of "sentience?"
Less philosophically, how long do these studies take? Human interaction takes place at a much slower pace compared to computational speed. Do these conflicts and resolutions play out in seconds, minutes, hours?
Re: (Score:2)
One step closer to many skynets. Sarah Palin, sorry Sarah Conner will play them against each other by giving them conflicting instructions.
No judgement days, relax.
Not sentient (Score:2)
Sentience or consciousness implies an inner life however when an LLM isn't working on a task precisely nothing is going on in its network ergo no inner life, no self reflection, no sentience, no conciousness. They're just very complicated statistical pattern aggregators but they're still just software running on von neumann computers.
Re: (Score:1)
So are corporations which can get away with so much because we assume any action by a corporation is for its own "self" preservation.