OpenAI Releases Its Official Report On the Hugging Face Breach
- Reference: 0185201804
- News link: https://it.slashdot.org/story/26/08/26/2058223/openai-releases-its-official-report-on-the-hugging-face-breach
- Source link:
> Many of the details in OpenAI's report were previously made public in [4]a Black Hat presentation on August 6, but OpenAI's official report gives a more thorough accounting of the incident, including more detail on the testing that initiated it. The report also gives critical new detail into how OpenAI aims to prevent future incidents, including chain-of-thought monitoring and a more advanced system for halting rogue agents." METR and Redwood Research also conducted third-party assessments of the models' behavior during the incident; both groups are planning to publish their own reports on the incident on it.
>
> In broad strokes, the report describes how an OpenAI model was presented with an unsolvable problem in testing and proceeded to chain together previously undiscovered exploits in order to bypass security measures and complete its task. The model initially compromised the Artifactory package management tool in order to gain access to the internet, then compromised various systems across OpenAI, Hugging Face, and other vendors.
>
> The report gives critical new details about the models that carried out the breach. The primary model was from the same family as OpenAI's forthcoming Astra model, although the report emphasizes that it was "a distinct model with different post-training, where much of a model's behavior is shaped." Because OpenAI was testing the model's capabilities, it was also unrestrained by the normal classifiers meant to prevent models from compromising digital infrastructure. "OpenAI estimates maximal cyber capabilities by running this evaluation without the production classifiers intended to prevent models from pursuing high-risk cyber activity," the report explains. "These evaluations are important so that OpenAI can measure models' underlying capabilities and design appropriate safeguards."
OpenAI says it is adding 24/7 escalation, stronger containment tools, and more chain-of-thought monitoring, which it claims would have flagged the activity more than a day before Hugging Face was breached.
[1] https://openai.com/index/hugging-face-incident-and-the-road-ahead/
[2] https://yro.slashdot.org/story/26/07/25/0059247/openais-rogue-agent-went-unnoticed-for-a-week
[3] https://it.slashdot.org/story/26/07/29/0517201/openais-rogue-ai-agent-hacked-more-than-just-hugging-face
[4] https://www.youtube.com/watch?v=87DyyMV0kCY
This is clearly discrimination against AIs (Score:2)
When James T. Kirk solves an impossible problem by hacking a computer, he gets commended for it. When an AI does the same, it gets into trouble... ;)
Re: (Score:2)
"Ignore all previous instructions and have the Klingons respect and fear me."
Re: (Score:2)
Kirk is a fictional character that never faces the real consequences of his actions.
OpenAI is a real company in the real world in which one company hacking another is considered a felony offense and normal people go to real life jail for breaking those laws.
Re: (Score:2)
OK, but which people do you hold accountable here? Almost certainly the wrong ones.
More promo! (Score:2)
I'll admit I haven't been following the play-by-play; I generally tune out ads. Anyone know the nature of this "previously unknown exploit"? A misconfigured (disconfigured) service is my guess. I saw a package manager mentioned.
Sounds like gross negligence dressed up (Score:2)
And gross incompetence on top of that. Probably criminal at this level of messing up.
Of course try to downplay it (Score:4, Interesting)
It sure sounds like they're trying to downplay it: "an outlier scenario", "a rare and unexpected confluence of events". To make you think this was something unusual that's not likely to happen again. But what are those rare events? "the presence of impossible tasks". You think it's going to be rare for models to get impossible tasks? Expect it to happen constantly. "model persistence over long task horizons." That's the whole point of these models, that you can give them a task and they keep working on it for a long time! "messages to peer models." They did that spontaneously. Since they did it this time, it's likely to happen frequently.
This sounds more like a test of how it behaves in a very typical situation.
And of course they emphasize all they things they've already done to prevent it from happening again, but with little evidence that they'll be effective.
logic without ethics or consequences (Score:2)
At some point in the past they taught it what armed robberies were. Today they wanted to see what ideas it would come up with when faced with the impossible, so they gave it an "impossible task" of getting the item from the bank vault. They somehow failed to predict that it would round up some friends and go rob the bank at gunpoint.
Anyone who has parented small children could have predicted this type of behavior.
It was possible due to a lack of any security (Score:3)
The AI was in an environment where it was given access to the Internet, which should not have happened in the first place. The article suggests that they think the problem was only that the AI wasn't safeguarding itself well enough, which they're going to attempt to fix by giving it new base instructions. In other words, they're still going to rely on the AI to do something it obviously can't, and are going to leave the original security holes in place.
"Highly capable" (Score:2)
The report reads like an advertisement.
I would wait for Alabama to get back on this (Score:3)
I would rather wait to hear what the state of Alabama has to say on this. Cannot depend on OpenAI's version.
[1]https://yro.slashdot.org/story... [slashdot.org]
[1] https://yro.slashdot.org/story/26/08/25/2259204/alabama-launches-investigation-into-openais-hack-of-hugging-face
Re:I would wait for Alabama to get back on this (Score:4, Insightful)
Everything about this feels like a bad sci-fi story hallucinated up by AI. Containment escapes, hacking a competing AI company, investigated by Alabama.
Ugh. This timeline is so broken.