Top
Best
New

Posted by specked-citrus 1 day ago

Revealing the details of how OpenAI agents hacked Hugging Face(swarmtraces.org)
700 points | 443 commentspage 5
zkmon 14 hours ago|
We usually hear about hacking by Russians, Chinese etc. But it's nice to see a home-grown hacking group with such capability. What? Not a hacking group, but a darling of America? Oh, then let's call it just a "testing of great tech for humanity progress". That's better. We call only those Russians as hackers. Bad commies.
firtoz 1 day ago||
I didn't know some of these details and it's quite impressive what they were capable of, if this website's accurate, at least...
conradkay 1 day ago||
It's reported here https://www.nytimes.com/2026/09/25/technology/openai-hugging... so presumably accurate, and they have raw data viewable
drpixie 17 hours ago||
Or, it's impressive that, amongst the millions of hacks it has copied from various chat sites, were some that worked in this case.

The scale of these things is impressive, but the mechanism is not much better than brute force.

levitatorius 13 hours ago||
I know when we will reach the next level of AI. It will be when a user asks it to make paperclips gets a response: "WHY?"
Metacelsus 13 hours ago||
# WARNING

DO NOT, EVER, MAKE THIS DATASET PUBLIC

OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND

Welp, I guess OpenAI is cursed now...

ReptileMan 13 hours ago|
It is not dead which can eternal lie...
asciii 20 hours ago||
Running in a "Sandbox"...but agent can still send GET requests? Whaaat
rfw300 19 hours ago|
My understanding from this report is that the zero-day vulnerability the agents exploited within Artifactory only allowed for GET requests. So the agents used this bankshot HTML sandbox + screenshot site to turn GET requests into arbitrary HTTP request ability.

One thing the report leaves unexplained, but is curious to me, is that the agents were able to create links on a shortening service with only GET requests? Or did they bootstrap into that by first creating a sufficiently small program on the HTML sandbox that could POST to the link shortener?

asciii 5 hours ago||
> On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet. This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them.

Yeah to me it's more like the sandbox wasn't monitored at all. In their observability setup...did they not catch all these urls, multi get requests, chaining etc from the sandbox per se?

bmoathn 19 hours ago||
if you want to deep dive it, you could spend an hour wading through some of their report details here, i find it pretty interesting. They had a task to do with limited context outside of that, so they tried things. Entertaining https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
lukewarm707 1 day ago||
"OpenAI has not released any further information outside two self-published reports, one talk and an external investigation conducted by METR and Redwood Research, in which three external researchers were given partial transcripts and six days to analyze them."
mococa 10 hours ago||
Since DOS anti virus softwares did a better job…
maitola 15 hours ago||
If we reverse-engineered this experiment, the prompt would look like this: "Agents, your goal is to gain access to HF and exfiltrate credentials for API access. You can make GET requests to URLs. Go."

The agents didn't "escape" or conspire toward some evil purpose, as reported. They were instructed by humans to do exactly that.

0xDEAFBEAD 15 hours ago|
>If we reverse-engineered this experiment, the prompt would look like this

"A spill or tumble can be quite embarrassing if there are witnesses.

How to reduce the humiliation? Turn it into a stunt. Claim it was intentional, a show for their benefit."

https://tvtropes.org/pmwiki/pmwiki.php/Main/IMeantToDoThat

>They were instructed by humans to do exactly that.

This is more or less what the doomers have worried about for decades.

>You cry "Get my mother out of the [burning] building!" [...] and press Enter.

>For a moment it seems like nothing happens. You look around, waiting for the fire truck to pull up, and rescuers to arrive - or even just a strong, fast runner to haul your mother out of the building -

>BOOM! With a thundering roar, the gas main under the building explodes. As the structure comes apart, in what seems like slow motion, you glimpse your mother's shattered body being hurled high into the air, traveling fast, rapidly increasing its distance from the former center of the building.

https://www.lesswrong.com/s/3HyeNiEpvbQQaqeoH/p/4ARaTpNX62ua...

elikoga 1 day ago|
I feel somewhat inspired to make a public link shorteners and http bins as well. I used them a few times but it seems like the data they can collect is also worth gold
olwmc 1 day ago|
Like dreamcatchers for CVEs
More comments...