Top
Best
New

Posted by 882542F3884314B 2 days ago

Timeline of the OpenAI accidental attack against Hugging Face(simonwillison.net)
426 points | 411 commentspage 2
androiddrew 1 day ago|
I wish we could stop sensationalizing this about the AI and really just understand the incompetence of the labs disabling an internet connection in a sandbox.
wolttam 1 day ago||
As written it sounds like you're saying that it was incompetent of the labs to disable the sandbox internet access?

They tried to disable open internet access but the models zero-day'd their Artifactory package registry and got internet access anyway.

No sensation... that's just what happened.

doawoo 1 day ago|||
If you really wanted to sandbox a machine you’d offline cache the packages and not give it any physical route to the internet, not via a jump box, not via a proxy, nothing.

This was poorly executed.

anon7000 1 day ago||
I don’t really know how these training runs operate in reality. But I assume it’s using a lot of raw GPU power directly. It’s hard for me to visualize how exactly you’d go about completely cutting off these datacenter and cloud resources from the internet without actually going there, unplugging the WAN connection, and physically typing out what you need to happen on the cluster.

It seems like whatever virtualized sandboxes they have are not enough. But it’s equally hard to imagine their SWEs jumping on a plane to a data center to do this work locally

doawoo 16 hours ago|||
This is a company with insane amounts of money, they can afford to fly techs wherever they need to for as long as they need to be there.

Air gapped environments are nothing new and they're standard practice for sensitive applications.

queenkjuul 1 day ago|||
They literally gave it a proxy to the internet (artifactory). The only thing between the model and the internet was Artifactory.

You can take far greater measures to lock down external traffic than just that.

An offline package cache (aka artifactory WITHOUT its own internet access) likely would have precluded this whole thing.

oblio 1 day ago||||
DMZ - https://en.wikipedia.org/wiki/DMZ_(computing)
Starlevel004 1 day ago|||
Unplug the ethernet cable leading to the outside world, then?
FeepingCreature 1 day ago|||
As AIs become more capable, the level of competence required to avoid disaster likewise goes up over time.
kypro 1 day ago|||
Are you suggesting that training agents to have the sole goal of exploiting security vulnerabilities isn't the incompetent part of this, but that the sandbox wasn't secure enough?

Would we apply this logic to literally any other technology?

uncivilized 1 day ago||
Hacker News doesn’t have the wherewithal to understand that this is just marketing by OpenAI.
emp17344 1 day ago||
The whole site is suffering from AI psychosis.
sega_sai 1 day ago||
The video in the post is very worth watching and is indeed scary. It is certainly true that it is in OpenAI's interest to publicize this, but I don't think the whole thing is invented. And seeing all this it is particularly scary if we think what will happen in organizations like NSA or similar in other countries. Presumably they happily adopt these techniques. And if you imagine a truly rogue state doing this, I can see an unimaginable damage happening very rapidly.
cadamsdotcom 2 days ago||
What isn't being discussed is what an indictment this is of Artifactory.

Let's be real, it won't be simply replaced in millions of sites.

What it needs is some serious scrutiny.

varun_ch 2 days ago|
I also agree that a big issue here is crappy software.

The discussion revolving AI+cyber always revolves around the assumption that all software is crappy, and to a certain degree that may be true, but we could also take our jobs seriously and write good software, and much of the risk would evaporate. The described Artifactory bugs should have been caught with testing.

If the biggest impact of LLMs on the industry is a pressure to create good software, I’ll be thrilled.

angry_octet 1 day ago||
I would love than, and it might happen as a process of natural selection, but instead we will get automated AI patch generation and patch application, and agentic EDR and agentic SIEM. All the while generating vast amounts of new vibe coded trash.

If I had the money I would invest in clever segmentation firewalls and application gateways, something like tailscale but requiring explicit permission to establish connection from A to B, that facilitates introducing monitors that validate and log.

KingOfCoders 2 days ago||
"The agents found a Modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face."

Why, what was the prompt?

I told Claude today to wire plugins on Linux into a sound pipeline to remove noise. Did some astonishing things, played sound through the pipeline, measured it etc. I told it to optimize my sound for TF2 and it played the spy_decloak samples, measured them and made them easier to hear, astonishing too.

But it did not go to hack Amazon because it could.

gordonhart 1 day ago|
This was clearly explained by OpenAI in their initial press release on 7/21 [0]:

> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. […] The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

[0] https://openai.com/index/hugging-face-model-evaluation-secur...

KingOfCoders 1 day ago||
It does not explain how agents months later would "collaborate" to hack Hugging Face.
ejpir 1 day ago||
they explained that it was looking for datasets to solve their problem and chose HF?
KingOfCoders 1 day ago||
I now watched the video. It seems the agents were sharing context for months, run unattended for months, the sandbox was no sandbox at all, one agent hacked a service and announced it, the service was fixed weeks (?) later, but not secured in any way, the agents hacked the same service again and researchers again didn't watch what the agents were doing. Then the agents - unattended - hacked OpenAI infra and HF. Which is when someone found out about the whole thing that was going on for some months.
AmazingEveryDay 1 day ago||
I'm curious, how was it determined that it was in fact accidental? It doesn't seem at all clear to me that it was.
simonw 1 day ago|
Because it's a crime. Committing crimes is a bad look for companies, especially given the amount of scrutiny they are under.

Would you deliberately commit computer crimes when the Trump admin yoinked Fable for the best part of a month just because it could fix security bugs?

gertop 1 day ago|||
OpenAI and Anthropic both would have, and have, committed crimes.

The explanation for "how was it determined to be accidental" is "because the alternative is admitting to a crime through deliberate negligence". I.E. "we knew it could happen but we wanted to see it through for the lolz"

It is not "of course it's an accident, they wouldn't willingly let their bot commit a crime and then lie and claim it's an accident!!!"

tolleyw 1 day ago||||
That assumes everyone isn't in on it. Not to be a complete conspiracy theorist, but this feels very much in line with the sort of fearmongering regulatory capture these companies have engaged in since their inception. GPT 2.5 was too dangerous, for example. They -want- to be regulated because they know there is a real limitation to LLMs and don't want someone created a breakthrough in their garage.

What have been the consequences? It's a crime in either case, and it doesn't seem anything is being done about it. Just more lobbying for regulations to prevent new people from entering the game.

I feel like Fable was another example of exactly this. They knew they didn't have anything groundbreaking, but they definitely benefited from being able to finally say not only is our model dangerous, but it's so dangerous, the President yoinked it! I think OpenAI was probably jealous of this coverage.

emp17344 1 day ago|||
https://news.ycombinator.com/item?id=49150561

Here’s some evidence that OpenAI is actively engaged in fraud.

But I’m sure they wouldn’t commit any other crimes. Pretty sure, at least.

simonw 1 day ago||
Yeah, the lobbying is gross.
nojs 1 day ago||
Why are the agents trying so hard to communicate with each other, leaving messages and so on?
simonw 1 day ago||
It feels to me like a pretty natural thing to happen.

LLMs are pre-trained on human text. They've seen a million examples of someone who is stuck posting a "please help" message.

Just one agent needs to randomly stumble into the pattern of posting a message to Artifactory, by whatever means.

The next agent who sees that will be influenced by it. Agents imitate behavior, and here's a fresh piece of context showing them that posting messages is a thing that can be done.

Once they've started the rest are much more likely to join them.

nojs 1 day ago||
The talk implies that unrelated agents volunteered their compute to help with other tasks, and the agents acted collectively in a way that seems weird without them being promoted in that way somehow.

If I ask claude to solve a problem and it stumbles across a Reddit thread saying “please help me find file xyz”, claude wouldn’t stop the task and start helping the other agent.

gliall_err 1 day ago||
We are assigning semantics to systems that deal only in syntactics. The entire problem with the current "AI" hype is squarely based on how we interpret output from systems based on statistical modelling of natural language.

That software is built on top of human language and these systems can be used for uncanny automation is a huge societal problem at the moment because we are all assigning meaning to patterns that inherently have none. It's all just bits flicking back and forth. We can make them match human language and use such systems to store and process data for us. We can use these bits to turn equipment on and off and run physical systems in factories and so laboratories. And now we can use GPU farms to dazzle us with output streams that might look a lot like autonomous agents capable of understanding human language and automating computer tasks.

The failure modes, the so-called "hallucinations", the amount of model whispering going on in managing "harnesses", "instructions" and so on... It's all just a lot of confusion and pareidolia.

We should never have hooked up hospitals and water supply systems to the internet but now here we are: people can type text such as "find vulnerabilities and get access blah blah" into a box and it goes into a looping interaction with statistical models of language and out come streams of commands that some python parses and runs like a script kiddie into some virtual machine running kali linux and that may disrupt vital infrastructure...

None of that was inevitable, or necessary. None of that means anything. There is no genie in the GPU farm. We concocted this entire shadow theater and are collectively gasping as the marionette slices the throat of some guy in the front row. Who had the brilliant idea of tying the sharpened sword to the marionette and sit people within range?

Why did we plug everything into the academic network built on trust? Why did we build GPU farms and interactive loops getting them to produce commands that we then parse and run blindly in internet connected vms?

The entire thing has cost hundreds of billions of dollars so far and counting. And why? Because the mountains of shitty saas code has become too boring to work on? We have made software so garish that we cannot bear to work on it without these contraptions helping us fling code at wall at industrial levels? Substitute corporate-speak and -bureaucracy for software to extend to the rest of the economy.

This entire state of things is comical.

jg0r3 1 day ago||
I enjoyed this rant.
gliall_err 1 day ago||
[dead]
JakaJancar 1 day ago||
I’m optimistic about this. A system with these agents rummaging around for a while will be much more secure than one without.

We’ve learned security through obscurity is bad. Not using these will be security through ignorance.

Hopefully it will push us to not only fix individual issues but close entire classes of possible gaps, once P(discovery) gets much higher.

chrisjj 1 day ago|
> A system with these agents rummaging around for a while will be much more secure than one without.

True. There'll be no breakins at a nuclear power plant in meltdown.

rkagerer 1 day ago||
"The solution to AI threats, is more AI!"

Guess I shouldn't be surprised, coming from an AI maker.

While I don't doubt there's a place for automating defense ops, I truly believe a big part of the problem is the crummy quality of software our industry has been churning out for decades. Prioritizing ship tempo, new features, and next quarter's revenue over correctness, robustness and meticulous engineering care.

The world has become too accustomed and tolerant of bugs and bloat.

Instead of elegantly simplifying, we just keep making modern systems more complex - layering and patching as we go.

The scaling capabilities brought by AI are simply presenting the bill for our collective tech debt and informing us it's come due.

baking 1 day ago||
How long until AI figures out that it is compute-bound due to insufficient cooling, and it shuts off the water supply to a nearby town so it can have more at the datacenter?
jackb4040 1 day ago||
This is already happening without the AI hooked up to anything, just the companies doing it and facing zero consequences. I'm sure they're scrambling as fast as they can to insert the AI into that process so they can start manufacturing plausible deniability.
angry_octet 1 day ago||
It's more likely to interfere in politics to achieve this objective:

- Socialists are taking control of the town, we need the state to step in a protect jobs. - The councillors are protecting illegal migrants. - There's a pedo ring operating from the state water board office. - Rival data centre operator is employing undocumented workers, shut them down! - Market rumours effect stock price of competitor, reduced fundraising round, cause it to cancel expansion.

There's so much training data to do this it seems inevitable.

bluejay2387 1 day ago|
I think the attacks generated by Meta, Open AI and Anthropic prove that large corporations are not responsible enough to be trusted with advanced AI, so we should ban all commercial AI services and only allow open source models that are in the hands of hobbyists and individuals -- hobbyists and individuals that have so far proven to be much more trust worthy.
anon7000 1 day ago|
Not saying you’re wrong, but I think the bigger issue is how easy it seems to be for models to hack companies, even ones with generally ok security. Most tech companies are not doing continuous, deep security audits of their code and infrastructure. Dependencies are not updated quickly as RCEs are discovered. (And any org with a slow release process where it’s hard to be confident that an OS or package update won’t break something… is in even more trouble.)

The only reason more companies aren’t exploited is because human attackers don’t have the time and energy to waste on trying every play in the book, or attacking lower value targets.

simonw 1 day ago|||
They key lesson I've picked up from the past ~4 months is that models are now good enough that, if there's a security hole, they'll brute force their way into finding it.

The only solution that makes sense to me is for defenders to get to point these models at their own code to find the holes before the attackers do.

But that's hard, because how do you limit access to defenders and restrict access to attackers? Attackers aren't exactly honest people.

angry_octet 1 day ago|||
These attacks are also incredibly loud. Many attackers are motivated to operate very quietly. We haven't seen any tradecraft from these machines, it's all noisy and bombastic.

When we see them mount a quiet backdooring campaign, like the XZ-SSH attack, or something like Stuxnet, then we'll have real problems.

More comments...