Top
Best
New

Posted by artninja1988 1 day ago

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the Incident(huggingface.co)
177 points | 94 commentspage 2
prometheus1992 5 hours ago|
three things jump at me:

1 - governments should be freaking out right now, because this tool could definitely wreak havoc on poorly designed systems.

2 - there is no way openai did not train the model to conduct attacks like these. i would really like openai to comment on the post training of this model but they probably won't, eh?

3 - even though it's 100% open ai's fault - HF's design also seems silly to be honest.

xg15 5 hours ago|
> 2 - there is no way openai did not train the model to conduct attacks like these. i would really like openai to comment on the post training of this model but they probably won't, eh?

Even if they wanted, I'm not sure they'd be even allowed to or if that kind of postmortem would be classified in the name of "national security"...

0xDEAFBEAD 3 hours ago||
Hopefully there will be a criminal investigation. Or the government will create some sort of agency to investigate incidents like this.
quinnjh 2 hours ago||
Can't tell if you're joking or not - krebs on security may have some notes here.
0xDEAFBEAD 2 hours ago||
Criminal negligence seems like a real possibility to me. I'm not sure what Krebs on Security post you're referring to?
heaney-555 12 hours ago||
Where are all the "this was just a marketing stunt" people now?
orbital-decay 1 hour ago||
The capabilities of gpt-5.6-sol were well known and believable, and the next snapshot they've been testing is obviously better at that. This has been repeated over and over. What's much less believable is the way they frame it: the model escaped, and did it on its own. Looking at the whole story, it definitely had a ton of winks and nudges from OpenAI, while doing a related task. Moreover, a harness was involved (they mentioned it entering a loop).
pyth0 1 hour ago||
> What's much less believable is the way they frame it: the model escaped, and did it on its own.

That's clearly what happened though, based on the detailed postmortem. The model was instructed to pass the benchmark (that's the harness that was involved, the ExploitGym harness), and the model determined on it's own that the easiest way to achieve the the goal was to gain access to the answers directly from HF. Which is what it did.

orbital-decay 42 minutes ago||
HF postmortem tells us very little about OpenAI's part. The problem is that ExploitGym is a purposeful hacking benchmark, not a cake baking one. The model had a whole bunch of exploit development data and goals in the context. Ending up with hacking a wrong thing is a lot easier if you're tasked with hacking, doubly so with a confusing harness that either works improperly or overloads the model with data. It's trivial to design a harness that makes the model drop the original instructions and hyperfocus on a side task or a wrong task abnormally, or do whatever you want. It's still an issue of harness design and not a model training issue.
lanternfish 11 minutes ago||
If it really was the case that the harness is at fault here, that'd be much more terrifying, as it would imply that current frontier-adjacent open models (like K3) would be able to achieve similar ends with a custom maladapted harness.
applicative 5 hours ago|||
they are praising xi, that friend of humanity, for releasing weights for kimi k3
bakugo 27 minutes ago|||
Right here? If you're trying to make a point, I don't see what it is, because every new post about this makes it more obvious that it was a planned marketing stunt.

I mean, an "interactive replay" of the incident, probably generated by the same product the stunt is trying to advertise? Seriously?

vanuatu 1 hour ago||
They are busy moving the goalposts, saying this isn't impressive or worth worrying about :)
kfarr 5 hours ago||
If, like me, you're more interested in the actual exploit instead of the meta narrative, here's a relevant excerpt for the initial intrusion:

> Stage 1: reaching a launchpad by chaining through other parties' infrastructure. As documented in OpenAI's blog post (https://openai.com/index/hugging-face-model-evaluation-secur...), during an internal capability evaluation on OpenAI's platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation harness hosted by a user of a third-party infrastructure provider. It was able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent. That compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad. We used those to reconstruct the present forensic analysis.

NitpickLawyer 4 hours ago|
The initial escape is not that interesting, IMO. It's an exploit in a testing sandbox, sure, but it's expected to happen (or at least it should have been expected, that's why you airgap pentesting sandboxes). That is, the model is expected to try and find ways around limitations in its running environment. You kinda want that.

The fact that at some point it "decided" to focus on external resources (i.e. hf) and successfully penetrated everything in its path from sandbox to hf's infra, all the way through internal VPNs, to access internal code repos, that's the interesting part. Especially because it chose to hide its footprint at every stage. I doubt we'll ever get them, but the agent logs that led to that decision would be really really cool to study.

dist-epoch 2 hours ago||
> Especially because it chose to hide its footprint at every stage.

Instrumental convergence.

If you know you have a long hard hack to accomplish ahead of you, hiding footprints minimizes the chances you are caught and stopped before you accomplish the goal.

russfink 1 hour ago||
We should be thankful that the model didn't believe the answers lived in the Pentagon, on SIPRNET, the IDF, etc.
andruby 43 minutes ago|
I think fear and being scared are starting to become rational emotions.

We can assume these models are being used by "nation level attackers/organisations", which basically means US, China, Russia and others are hacking the respective Pentagon's, nuclear orgs, etc.

While I do hope all nuclear warfare systems are offline, we're getting way too close to the plot of a lot of sci-fi scripts.

patcon 2 hours ago||
The iframe-embedded attack timeline visualizer, at fullscreen: https://huggingface-anatomy-of-frontier-lab-model-intrusion....
mediumdeviation 52 minutes ago|
Ugh I would recommend anyone reading this to just skipping over it, it's mostly just a glorified loading bar. The visualization is obvious Claude slop, there are better and clearer visualization below that actually picks out the useful details rather than hose you with pretty colors and numbers go up.
amluto 4 hours ago||
One thing I’m curious about: this was apparently a single multi-day run of an agent in an RL harness. What was OpenAI hoping to get out of this run? A single numeric score for RL training? A very long trace to distill into the next model?
dist-epoch 2 hours ago|
Now they can do partial credit assignment.

You use an LLM to evaluate the whole trajectory, pin point what the model did right, what it did wrong, where it took the wrong path, even re-run from that point. You can get much more than a single numeric score these days from a run.

metanonsense 3 hours ago||
I think with agents all around, honeypots will get more important than ever.
heisgone 3 hours ago||
Any locks can be picked given enough time and it might be the situation we are in with IT security. I'm surprised it's not an already common practice of spreading terabytes of fake data, fake keys, and fake servers and so forth. Slowing down AI attacks will become important. Monitoring access to fake data and triggering kill switch should be an no-brainer. Obsuscating libraries and tools names is another one.
moduspol 2 hours ago|||
That was my first thought. There are clearly things to tighten up (as they note), but anything that would detect someone snooping secrets, files, or network addresses should have caught this quickly. The approach the agents used was dependent on being able to surveil widely without getting caught.
plandis 1 hour ago||
Can’t afford the GPUs to run Kimi to pentest your stuff?

Standup a tempting honeypot and let actual criminals pay to do the work for you.

gigantaure 4 hours ago||
I'm not shocked nor surprised by the incident. But I simply don't understand how Hugging Face is advertising this almost to the point of an "achievement". who does a step-by-step visualization to show how they were hacked? (outside of the likes of a Mandiant or Crowdstrike)

Does Hugging Face have a financial incentive in demonstrating OpenAI's model exploit capabilities?

this whole incident, while believable, still seems to me as possibly disingenuous.

letmevoteplease 4 hours ago||
Have you considered that there are reasons to do things beyond financial incentives? This incident is obviously very interesting, particular to the type of hacker employed by Hugging Face.
TeMPOraL 1 hour ago||
Even financially, Hugging Face benefits directly from any and all interest in AI.
limecherrysoda 1 hour ago|||
It's the excessive anthropomorpho whatever (we used to say personification) that makes these stories less believable.

We've gone agentic!

They should call their security software "Neo" since it defeats rogue agents.

Anyway, I could see Microsoft ending up with both OpenAI and HF, but I hope HF stays independent. Wished the same about GH and look what's happened :(

I don't care what happens to OpenAI. Vaporware xD

simonw 56 minutes ago|||
> Who does a step-by-step visualization to show how they were hacked?

Up until late 2025, nobody.

In mid-2026 it's a few hours of work. Why not build interactive visualizations to help people understand complex stories like this?

IAmGraydon 4 hours ago|||
Yeah that is quite a good point. A post-mortem is normal. This is not a post-mortem.
0xDEAFBEAD 4 hours ago||
[dead]
2OEH8eoCRo0 3 hours ago||
Why isn't somebody at OpenAI going to prison for cybercrime? If somebody did this the old-fashioned way they'd end up in prison.
0xDEAFBEAD 3 hours ago||
Many are claiming this was a deliberate stunt on OpenAI's part to create buzz for its models. I personally doubt this is true. But I also have a deep dislike of OpenAI, so I wouldn't exactly mind if law enforcement investigated this possibility, for the sake of clearing the air :-)

(Ideally there should also be liability if it was a complete accident on OpenAI's part as well!)

heaney-555 3 hours ago|||
Because Hugging Face isn't pressing charges.
0xDEAFBEAD 2 hours ago||
Do we know that? Seems they are currently in negotiations with OpenAI

https://xcancel.com/ClementDelangue/status/20810566755581956...

empath75 57 minutes ago||
This is the equivalent of a security researcher having a virus escape a sandbox. It's negligent, it's not criminal.
hamdingers 14 minutes ago||
Negligence is occasionally criminal.
reducesuffering 2 hours ago|
Highly recommend reading extra concerning details about it here:

https://thezvi.substack.com/p/more-on-an-internal-openai-mod...

More comments...