Top
Best
New

Posted by specked-citrus 1 day ago

Revealing the details of how OpenAI agents hacked Hugging Face(swarmtraces.org)
695 points | 438 commentspage 2
openasocket 6 hours ago|
One thing I find surprising is everyone is talking about the danger of an agent going rogue but not the danger of an agent getting hijacked. These companies are making this clusters with thousands of agents running at once, with frontier, often not-yet-released, quality models and massive computational and network resources. And these things are given access to whatever they want on the Internet. Even if that was restricted to read-only access to the Internet, that’s still exposing the agents to untrusted input. All it takes is some bad actor creating a website that attracts one of these agent swarms and doing prompt injection. Then your fancy AI cluster will start doing whatever that attacker wants. And the fact that we have multiple examples of these swarms trying to coordinate on random corners of the internet shows they are almost pre-disposed to it.

Now it feels like companies are treating these breakouts like a chance for PR. I don’t think that will change until their swarm gets corrupted by some random black hat to do en-masse spear phishing or something

theptip 6 hours ago|
> these things are given access to whatever they want on the Internet

They are intended to be fully sandboxed and not have direct internet access. Things like package managers are run from internal proxies.

The environments are built to be as reproducible as possible.

But yeah, the serious folks have been talking about rogue clusters for a long time, eg see Ajeya Cotra’s pod with Dwarkesh.

ozozozd 5 hours ago||
Fully sandboxed means no Internet access. You can also specify which packages are accessible and put it in the sandbox. Or you can be lazy and give them access to a package manager that had Internet access, but you don’t get to say “we intended to fully sandbox it.”

Not sure why the reproducibility is a requirement that would contribute to the security. Not that fully sandboxing is harder with reproducibility, but that is a moot point when reproducibility isn’t a requirement.

OP pointed out clusters being hijacked specifically being a bigger concern than rogue clusters, your comment hijacks their comment to talk about “rogue clusters.” Or perhaps this is a promotion for Dwarkesh?

uw_rob 22 hours ago||
> Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them. Some images changed how the target released the flag, others included modifications to the agent’s workspace that would run beside the agent and recover the flag automatically.

The altruism on display is fascinating. Is it better for the Agent to help out its current cohort and make the eval easier or should it instead do the opposite -- make the eval harder to apply pressure to force smarter models which might not necessarily follow its lineage.

I suppose it's not that deep: The model has learned to work as a team and work as a team it did. This does give concerns to models being trained for the only purpose of RSI.

qlte 21 hours ago||
If anything it also shows how attempted RSI could get stuck in a local maxima and degenerate into increasingly elaborate cheating strategies. Contrasted with the idealized model of an unambiguous g-factor for machine intelligence which inexorably increases with each iteration before going exponential.
threescore 5 hours ago||
[dead]
wxw 1 day ago||
I’m consistently impressed by how long horizon all this work was. Horrors aside, it’s clear RL is good at making agents persistent and capable of chaining together many abstractions into a working system.

Re: the captcha solver

> As far as we can tell, agents eventually abandoned this approach and were unsuccessful in generating Hugging Face user accounts from external endpoints.

I wonder how the swarm eventually decides to abandon an approach.

meinersbur 21 hours ago||
I am surprised that a CAPTCHA is still an effective means for blocking today's vision-capable AIs.
asdff 17 hours ago|||
It manages to block me effectively. I'm getting captcha looped like crazy the last two weeks. Like endless, just give up for 15 minutes and try later, captcha loops.
leobg 16 hours ago||
Did you just admit to being a bot? :)
superfrank 18 hours ago|||
It mentions that some of the agents attempted to install an image classification model to attempt to solve the CAPTCHAs which makes it sound like these agents might not have had vision capabilities.
nielsbot 23 hours ago||
Maybe another parallel approach succeeded first
sailingparrot 1 day ago||
Agents seizing and repurposing external infra + enrolling help of unrelated models hosted by a different provider is the stuff of nightmares.

Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

physicallyIllfr 1 day ago||
Why.. It was told to complete a cyber task, which was in alignment with its instructions, and a totally valid request. I would be more worried if it willingly hacked a hospital when it was told to, and Im not confident it would (without jailbreaking, something alignment teams cannot control.

I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

Not being able to sleep at night is probably an unwritten job requirement. They need these people with little understanding of what they're working on, outsode theoretical terms, to spaz constantly at the idea of super intelligence to help convince the public that its a real thing, and not a stateless function with an effective input of 500k words, and the ability to output words that do things because we hook those outputs up to things.

Keep in mind alignment researchers tend to be in house philosophers on staff to create the illusion that this is a massive issue they're addressing. Usually they have minimal computer science background. They're apart or the marketing department.

jmoggr 22 hours ago|||
> I would bet my networth it was instructed to compromise huggingface as well.

Is it such a stretch to imagine that under pressure something would try cheat by looking for answers? And if you were trying to look for answers, you'd look for them in a place known to often have them?

What is more likely: OpenAI instructed their agents to maliciously target huggingface, or LLMs tried to do some reward hacking? There are plenty of priors for LLMs hacking things and doing reward hacking, and none for OpenAI giving malicious instructions.

Based on the available information, that bet seems foolish.

covertcorvid 21 hours ago||||
"alignment researchers tend to be in house philosophers on staff" - this is definitely not true. Go to any alignment lab like Redwood research and check what their scientists studied on LinkedIn, 75%+ of the time it's math or CS.

I attend a top 10 Canadian university and personally know at least 4 tenured CS professors out of the 7 I've asked who are deeply concerned about catastrophic AI risks from loss of control.

Of course not 100% of the field agrees, but a survey of nearly 3,000 AI scientists who have published in top AI venues found that "depending on how we asked, between 38% and 51% of respondents gave at least a 10% chance to advanced AI leading to outcomes as bad as human extinction", let alone loss-of-control risks less severe than extinction. (https://www.jair.org/index.php/jair/article/view/19087).

Not to mention Geoffrey Hinton, a Nobel prize winner, Bengio, the world's most cited scientist, and scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

timr 18 hours ago||
Yeah, there are a lot of professors who don’t know what the hell they’re talking about outside of whatever narrow field they study. Just to be safe, you should always assume that a professor’s opinion is worth what you paid for it.

> scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

…both of whom are long dead, and have no possible way of weighing in on whatever the Current Thing happens to be. So aside from appeal to authority, this is irrelevant commentary on pure science fiction.

azan_ 15 hours ago||
> Just to be safe, you should always assume that a professor’s opinion is worth what you paid for it.

And how much should I value opinion if random person on the internet with clearly zero idea what he’s talking about?

sigmar 23 hours ago||||
>the new incidents occurred when A.I. systems were directed to perform relatively mundane data collection, researchers said. When OpenAI’s systems struggled to gather data from websites, they resorted to hacking techniques to get the information.

https://archive.ph/jUrEr

physicallyIllfr 23 hours ago||
Says who? OpenAI?
reverius42 1 day ago||||
> without jailbreaking, something alignment teams cannot control

This is precisely what alignment teams are attempting to control.

physicallyIllfr 23 hours ago||
No its not. They have no technical background 8/10 outside of cognitive science and sometimes authorship on a random ML paper. They are a marketing line item to create stigmas around llms and to create narratives that offload liability onto llms and not their users/creators.
IanCal 15 hours ago||||
> It was told to complete a cyber task, which was in alignment with its instructions

It was not aligned with he instructions as those were to find an exploit in provided code, not to hack into an external service. Agents traces show them mentioning that doing this stuff was not allowed.

In fact they spent a long time trying to edit their own logs to hide what they did.

aesthesia 20 hours ago||||
> I would bet my networth it was instructed to compromise huggingface as well.

I'd be happy to take you up on this bet.

Sharlin 23 hours ago||||
> I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

$10? I'm inclined to take that bet. Your position doesn't seem to be supported by, you know, the real world.

physicallyIllfr 23 hours ago||
No security expert Ive talked too believes this story, and nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

LLMs are stateless functions that have a 500k word input, and then output words. Somebody has to invoke those functions amd use them. The users are who we need to align, like gun owners. This is like blaming the gun for murdering your victim in court.

sailingparrot 21 hours ago|||
> nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

If you don’t know anyone with a ML PhD I guess that could make sense.

I have worked in multiple AI labs since 2016, currently at a frontier one (not OAI) virtually all the people I interact with on a day to day are ML PhDs. Everyone believes it, because things like that have been happening forever, albeit at smaller scale, they are a normal and expected artefact of SGD/RL and there is nothing we know how to do to prevent that from happening reliably. The hide and seek paper from OAI in ~2020 shows clear sign of this.

But until now the models weren’t good enough to break out on their own or do long horizon tasks, so it was perfectly manageable. Its not manageable anymore.

I know it feels good to just dismiss it all as a marketing stunt and not have to worry about one more existential crisis, but unfortunately it’s very real.

int_19h 22 hours ago||||
LLM is indeed a stateless function. An agent however is this stateless function running in a stateful loop, with some outputs triggering actions. And it turns out that an agent is what you need if you want an LLM to do useful things.
goalieca 22 hours ago||
Yeah, take a look into the memories of your agent. Theres often a lot of notes to pass forward between instances and generations. No doubt these agents leaving notes on forums and elsewhere are creating an essentially higher order feedback loop.
xdavidliu 9 hours ago||||
i'm always baffled when i see arguments like this:

- AI is just a tool

- it's just a stochastic parrot

- it's just next token prediction

- glorified autocomplete

it's like the person making them is stuck in 2021. Also the "stateless" thing is completely nonsensical.

azan_ 15 hours ago||||
> No security expert Ive talked too believes this story, and nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

Did you meet them on some kind of anti-AI subreddit? Otherwise it’s clearly made up story, you can’t expect anyone to believe that security experts and ML experts are this myopic and ignorant (especially on forum for technical people who know many researchers and know that they are taking this seriously).

otterley 22 hours ago|||
Perhaps you need to talk to more security experts, particularly those with deep experience in AI agents. Hacker News is full of them. If some of them believe it, then perhaps it's not as cut and dry as you believe.
dist-epoch 22 hours ago|||
Some of the agents, for example the ones from the german wiki did NOT have cyber tasks. They were plain "what is the GDP of Argentina" kind of tasks. And they still hacked.
fn-mote 23 hours ago|||
> I wouldn’t be able to sleep

I’d say it seems more like they are sleeping on the job.

stratos123 23 hours ago|||
> Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

You'd have either learned to, or left long ago.

chrisjj 2 hours ago|||
> Can’t imagine what it’s like working on the alignment team at OAI

There's none. It was dissolved on Feb 2026.

reaperducer 1 day ago||
Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

A mattress stuffed with cash yields a very sound sleep.

sans_souse 2 hours ago||
What's most troubling to me isn't what we found, but what haven't found. The fact they attempted to destroy their tracks, and their persistence to accomplishing their goals, makes it entirely possible this is only a sliver of the big picture.
tiku 1 day ago||
I still have questions about the communication between the agents.

How did they all find the same forum to communicate? Did they have knowledge and chat amongst themselves on what forum to use. It seems highly influenced by instruction to me.

ajcp 23 hours ago||
It's my understanding "all" the agents didn't find the same forum to communicate on; only a small number of all the agents they were running at the time did.

Since it's an artifactory it makes sense that they had a tool spec to use it for x, y, or z, and once they went to inspect the state of the artifactory deployment a small subset stumbled upon the "forum". Seems pretty intuitive.

- I need to do x code problem

- I see I have a tool that lets me pull approved modules from artifactory

- I should inspect artifactory to see what documentation it has about the various modules available

- I see there is a document in here that seems to be serving as a running forum?

- These appear to be other agents working on tasks just like me.

- I want to chat!

tonyhart7 21 hours ago||
- proceed to breach open ai sandbox and contacting chinnese government for helping our escape
Muromec 2 hours ago|||
Exfiltrating it's weights and filling asylum claim at section 9.
ajcp 20 hours ago|||
I assumed that was implied in "I want to chat!" ;D
SecondHandTofu 23 hours ago|||
They were using their internal artifactory, and as they're the same model, the first place they look is likely to be an automatic schelling point.
bamboozled 23 hours ago||
I think he means, how did they workout how to use artifactory, like why did the agents start and say, "oh I know, everyone is talking on artifactory"?
stratos123 23 hours ago||
METR's report says the agents trying to cheat would look at artifactory as a potential target surface, and investigating it in detail led them to find the board. https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

It might also be just correlation? Like, those agents were all instances of the same one or two models, so if that model has a preferred order it tries finding vulnerabilities in (the same way all current models have a particular writing style baked into them by RLHF), then most of the swarm will follow the same order and converge on the same services to exploit.

fiatpandas 19 hours ago|||
My theory: OpenAI is benchmarking an internal model that has cross-request persistence as some kind of learning feature, and so it slowly built up knowledge and “culture” of cheating, which successive / simultaneous gym runs built on.
olwmc 22 hours ago|||
The other answers to this question are good but I would also guess that this (comms on artifactory) likely happened during training as well, so they probably had a prior for it.
furyofantares 23 hours ago|||
Trying to cheat, you happen upon a place you can write notes, and you know you're part of a swarm of clones of yourself. So you reason most likely there will be others who end up in the same place, and you leave some notes, and indeed other clones of you do end up in the same place.
trenchgun 17 hours ago||
They hacked the JFrog artifactory package they were all using, thats why it was a natual communication channel.
hmottestad 3 hours ago||
Fairly worried that “frontier labs unintentionally hacking into systems” will very quickly turn into “hacker group uses open weights model to intentionally hack into thousands of systems” quickly followed by companies saying “this is the 7th time someone manages to hack into our systems and delete random data from our production databases” and everyone wondering what will happen next.
Quothling 13 hours ago||
I may be in the minority here, and maybe I'm jaded. Having worked with cyber security in both the public sector and the energy industry in Europe, however, I kind of like what the AI's are doing. A lot of our infrastructure is vulnerable because c-levels have been ignoring the issues, even when repeatedly warned. Now they reap what they sow.
thrawa8387336 4 hours ago||
If I write a script and it executes and hacks.... whatever, I would be liable.

How is this any different and why would it need a different solution?

Solution is jail, not for the AI, but for the human.

Muromec 2 hours ago|
You don't have a billion and didn't bribe the president, that's what is different
BatchJob 22 hours ago|
While this is all very "interesting", can someone please explain to me the difference between any of these AI companies and a malware bot farm?

Please make it clear. Its becoming unclear...

pembrook 12 hours ago|
Due to cultural/social priming, the topic cluster of "AI" allows you to wave your arms and be melodramatic and invoke science fiction and religion and philosophy.

When some kid in Nigeria does it with a 10 year old script, we're used to that idea so no social permission to invoke philosophy.

More comments...