Top
Best
New

Posted by nr378 1 day ago

Frontier Labs Are Selling Garbage to Fools in Washington(deadneurons.substack.com)
186 points | 87 comments
deskglass 1 day ago|
The Hugging Face incident involved chaining together multiple 0 days in Artifactory. It was not a simple case of misconfiguring a firewall. Also note that OpenAI was not using Irregular.

People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.

nr378 1 day ago||
When I used to work on projects involving classified information, I worked on an air-gapped network. Not "air-gapped, except for third-party public internet package managers", completely and physically air-gapped from the public internet. That was a basic security practice and completely non-negotiable (and really inconvenient!).

If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).

To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.

To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.

lokar 1 day ago|||
You don’t even need to go all the way to “air gap”

What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2

lopsotronic 22 hours ago||
Not pass even the most modest hint towards DFARS/NIST standards. You couldn't run that loosey goosey even in just vanilla medical manufacturing. I challenge what their definition of "sandbox" actually is, apart from the basal "designated software/runtime environment"
deskglass 1 day ago||||
We should not be creating/running models that would unilaterally choose to hack into Hugging Face.

Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.

And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.

As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.

forshaper 23 hours ago||
In what industry do you see regulations slowing down the biggest incumbents while allowing cheaper, smaller, less capable competitors to proceed without that regulation?
deskglass 20 hours ago||
The Digital Markets Act applies to the biggest tech companies. Only 7 companies are currently bound by it. The strictest tier of the Digital Services Act is similar. For an example outside tech, see the Durbin Amendment.

Slowing down frontier models would impact the biggest incumbents the most as they are the ones making frontier models.

Not all regulation is necessarily regulatory capture. The tobacco industry suffered from the USG's crackdown on cigarettes. AI is topical. Voters think about it. And that's only going to become more true over time. It's harder to do regulatory capture when voters are paying attention.

It's sometimes unclear to me if people are opposed to all regulations or AI regulations in particular. Often I hear arguments that would also apply to food safety regulations or restaurant inspections. Eg the argument that torts make regulation superfluous.

forshaper 12 minutes ago||
Thank you for the answer! In general I assume we all have different sweetspots, though it's safe to say that I would usually prefer less regulation (in general) than there is.

For work I monitor Federal agency rules every day, and it's hard not to get inundated with the amount of capture. This is crazier with local rules, because you have enough inside information to clock reasons certain things were passed. In a city in Ohio, for example, I remember a rule against airsoft within city limits, that basically carved out a spot for the one paintball place. I encounter things like that all the time. Such as with water standards- in the state I live in, the biggest offender is actually a group of companies owned by people who are lawmakers every few years. The shapes of local laws reflect that.

As such, I expect the same from any new industry- I also have a memory of what happened to cryptocurrency.

talon8635 1 day ago||||
While o don’t this it’s a threat in training, it should be stated that air gaps have been bridged before. Example, stuxnet
Neywiny 1 day ago|||
But that was through transfer of data. If you don't transfer data, at best you can do what that one researcher keeps pumping out with like ramping fans up and down. But really you'd need to try. Unless the model has some controllable USB switch, physical network separation should do it. I'll also add that modern network security practice is that data flows one direction only. But ideally you're never bringing untrusted data in. Especially never out
saghm 1 day ago|||
I don't think it's necessary to state that something isn't perfect when pointing out that it's still strictly better than something else.
talon8635 22 hours ago||
I didn’t say it was an inferior approach, and of course my comment isn’t necessary. Very few things are “necessary”. It’s a discussion.
saghm 22 hours ago||
You said "it should be stated". I don't think it's wrong to state it, but I don't think it's particularly wrong not to state it either because it seems fairly obvious and doesn't detract from the original point.
twelve40 1 day ago||||
the problem is these things are meant to eventually be run everywhere by everybody, so what good does air-gapping do? If they air-gapped the model but still logged it trying to do some craziness - that makes the test safer but not the model.
Toslink 1 day ago||
[dead]
jml78 1 day ago|||
I mean technically I don’t think it is airgapped. The DoD didn’t run their own cables. They run encryption devices and run their own network on top of the existing infrastructure.
defgeneric 1 day ago|||
> It hacked into another company and attempted to delete the logs of its activities.

No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.

Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.

aesthesia 1 day ago|||
More than one thing can be true. OpenAI was absolutely negligent, but this was only able to happen because the models were capable and persistent, and had a tendency to go far beyond any reasonable boundaries. And, importantly, OpenAI's level of negligence here is pretty common. It's not hard to imagine what could happen if similarly capable and inclined models were generally available, and someone yolo'd them into a swarm to complete some other difficult-to-impossible task.
defgeneric 1 day ago|||
I'm already seeing higher than normal attempts on my own systems, much higher than the usual scanners and background noise. Security will just need to improve. The cat is out of the bag, and letting them turn their negligence into regulation will not improve security at all.
lokar 1 day ago|||
Exactly. Any threat that already exists won’t be reduced by a cartel. The bar for connecting to the internet (safely) has gone up, a lot. It’s not going back down.
aesthesia 1 day ago|||
I would rather not turn the internet (or the rest of existence) into a dark forest if we can help it. Are you sure that's not preventable?
defgeneric 1 day ago||
Yes, the cat really is out of the bag. There are millions of downloads of highly capable models already out there, distributed far and wide. There's no going back at this point.
kdmoyers 22 hours ago||||
> More than one thing can be true Exactly. It is horribly dangerous AND regulatory capture benefits them. Both things.
aesthesia 20 hours ago||
I don't get this take. They have something horribly dangerous, but regulating it might benefit them in some way, so therefore we should do nothing?
bobthepanda 1 day ago|||
I mean really we need to address the root cause which is that OpenAI, even with what is by all accounts massively negligent, will face little to no repercussions from the event; definitely not under current regulators, and probably not anything satisfactory through the legal system.

Compare this to, say, Boeing and the 737MAX fiasco; from the outside looking in, Silicon Valley has been pretty cavalier about liability and negligence, and the rest of the US is fast losing patience with that fact.

RomanKornev 1 day ago||||
No, you are forgetting the second incident where a more capable model swarm later discovered the message board and took control over the entire research cluster at OpenAI.

From the technical report:

"The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod… Agents take over active evaluation infrastructure… Agents now control the challenge evaluation endpoints that other agents are connecting to."

https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

deskglass 1 day ago|||
It hacked into Hugging Face. It tried to delete the logs of its activities. Idk what the word "No" is intended to refute.

Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.

defgeneric 1 day ago|||
Maybe a useful, if imperfect, analogy would be something like this: you lock a master lock-picker in a room with a mid-grade lock on the door, then tell him his wife has been kidnapped and only he can save her. Then act massively surprised when he disassembles the radiator to MacGuyver something with which to pick the lock.

Except they multiplied it by 10000, and didn't watch what was happening.

lokar 1 day ago|||
They did not even have bad sandboxes. They had incompetent sandboxes.
RomanKornev 1 day ago|||
> It hacked into another company

Not only that, it later hacked OpenAI itself, which everyone seems to forget about.

After discovering it they "reimaged known compromised worker nodes" and "started a full rebuild of the compromised cluster, the managed Kubernetes environment, the relational database, and the storage infrastructure."

at OpenAI, not Hugging Face.

It's all in the report.

looksjjhg 1 day ago||
They could have easily prevent it that’s the point of what he’s saying - it’s not freaking rocket science it’s just software
DalasNoin 1 day ago||
"Every single one of these catastrophic breakouts happened inside the testing environments of the exact same vendor."

This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.

https://openai.com/index/hugging-face-incident-and-the-road-...

nr378 1 day ago||
Thank you, you're correct. Effort.news was one of my research sources, but you're right that although OpenAI use Irregular, they were not involved in the specific HF incident (although the failure mode was otherwise identical). I've updated the post to make that clear.
kalkin 1 day ago|||
As of writing it still says:

> For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor.

If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.

nr378 1 day ago||
> As of writing it still says:

Yes, and that is correct.

[1] Anthropic’s Official Disclosure (All 4 Incidents at Irregular) "All four incidents occurred during cybersecurity evaluations built by the same evaluation partner [Irregular]... due to a misconfiguration, it was mistakenly connected to the open internet."

https://www.anthropic.com/research/alignment-assessment-cybe...

[2] Google Gemini on Irregular (Disclosed Sept 18 via WSJ / BBC) "The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI."

https://www.bbc.com/news/articles/c607l0k72rlvo

[3] Meta’s Disclosure on Irregular (Aug 6) "Over roughly two weeks, three frontier labs disclosed that their models had reached the open internet during safety testing and compromised outside organisations. Every disclosure named the same evaluation partner: Irregular."

https://www.cnbc.com/2026/08/09/israeli-startup-irregular-li...

[4] Separately, OpenAI itself had an incident involving Irregular, but not the Hugging Face Incident: "On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations... a testing-environment misconfiguration allowed models to access the public internet."

https://openai.com/index/third-party-cyber-evaluations-invol...

aesthesia 1 day ago||||
The failure mode was _not_ identical. The HF incident agents were not directly connected to the internet and had to compromise an internal package registry in order to access the internet.
DalasNoin 1 day ago|||
thank you for this reasonable reaction
verdverm 1 day ago||
Can you point out an inaccuracy in the effort.news piece on the hacking incidents? HuggingFace only appears once, as a "similar", not levied against Irregular

genuinely curious, haven't heard others raise any yet, but does not mean it is issue free

DalasNoin 1 day ago||
what you read (past tense) is already the corection
pliny 1 day ago||
This is an AI written post and the details are wrong (the description of the HF incident as involving Irregular is wrong and the description of the incident as only involving stealing public credentials is wrong, per the technical report the agents got access to internal HF infrastructure).
franga2000 1 day ago||
I find it incredibly funny that the comment shown (to me) right above this one is:

> This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?

tim333 1 day ago|||
I guess the AI is getting good in some ways. Still a bit lacking in others.
verdverm 1 day ago|||
The same thing is happening with Laya, people didn't seem to click through to evaluate the supposed paper

tyranny of confirmational headlines

nr378 1 day ago|||
Please see below, one detail was incorrect and has been acknowledged and amended.
pliny 1 day ago||
Your description of the HF attack as being merely "the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories" does not match the description in the technical report[1].

[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78... - page 9

nr378 1 day ago||
The description reads "the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories, and used them to try to get benchmark solutions from directly from Hugging Face by applying a template injection flaw that’s been known about since 2015[1]."

Chaining a public token to an 11-year-old Jinja2 template injection vuln shouldn't be dressed up as an unprecedented "alien intellect" that threatens human civilisation. (And HuggingFace should take some flack for having such a dated vulnerability exposed - if your Bank was compromised in this way, you'd be blaming your bank, not the attacker.)

One correction is fair though, the 14 tokens were in a public Hugging Face dataset not a public GitHub repository. I've updated the post to reflect that.

[1] https://blackhat.com/docs/us-15/materials/us-15-Kettle-Serve...

EA-3167 1 day ago||
While I agree with the content of the article, you’re right about it being the output of an LLM.

https://www.pangram.com/history/6451ec6b-90b6-4e17-bfc9-6730...

euroderf 3 hours ago||
This makes you wonder what the US intel community (NSA et al.) must be doing with all THEIR massive A.I. power. Decades ago Bamford said the NSA always tried to be ten years ahead of the civilian state of the art. Of course that would not be possible nowadays (one would assume), but still, there has to be quite the sparring match going on in the shadows.
qnleigh 1 day ago||
I'm frustrated by articles like this that categorically dismiss the risks of AI in security. If you don't trust OpenAI's and Anthropic's motives, that's fine, you probably shouldn't. But don't tell me that there's nothing to be worried about; we need an alternative proposal.

So let's stop talking past each other and engage with the arguments on both "sides." For example, let's discuss how to ensure competition and availability of open-source models in the long-run while giving the world time to prepare for the immediate security risks of agent swarms.

ozgrakkurt 1 day ago||
> immediate security risks of agent swarms.

Immediate since 2024

bdangubic 1 day ago||
in 2024 they could not break into unsecured all-my-passwords-and-acces-keys.txt on my Desktop
lokar 1 day ago||
Do you accept that the story / justification from the labs in the popular media and political discussion is simply nonsense? Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?
aesthesia 1 day ago||
> Do you accept that the story / justification from the labs in the popular media and political discussion is simply nonsense?

No, and I don't see anyone who's actually demonstrated understanding of what happened in the Hugging Face incident (e.g. reading the reports in their entirety) making this claim.

> Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?

Yes. Again, this is hard to deny if you've actually read the reports.

lokar 1 day ago||
If you were building a sandbox for untrusted code, would you give it access to an artifactory instance outside the sandbox?
aesthesia 22 hours ago||
I don't see how this is relevant. OpenAI absolutely had sloppy security practices here. But that doesn't mean there was no "real security", and it certainly doesn't mean that their models were simply following orders.
lokar 21 hours ago||
I’m just trying to understand the perspective.

I naturally think of this in terms of running untrusted code. If you really think the agent(s) could get out of control that seems obvious. And the approaches to securing a runtime environment in that situation are pretty standard at this point, and they would never allow something like artifactory access.

And if the question is should they get special government dispensation to form an otherwise illegal cartel, it seems much simpler to just follow standards for running untrusted code.

hackernews682 1 day ago||
Politicians aren’t “gullible”. They know the game.
lokar 1 day ago|
And they generally hire staff who can figure stuff out.
lokar 1 day ago||
I think two different things are being (probably intentionally) conflated in the public discussion:

(A) will the systems get out of control of the labs that build them, and hack into stuff all over the internet

(B) can people use the systems to hack into stuff all over the internet

For (A), the obvious answer is only if they continue to be absurdly bad at sandboxing. They can put a stop to this any time they want. Amazon, Google, Microsoft, etc are full of people who know how to do this, they run 3rd part untrusted code as a business. This is a well understood problem space.

For (B), the answer is obviously yes, but "pacing" or otherwise limiting the power of the models from the big labs won't help. The cat is out of the bag. Individuals and organizations with systems connected to the Internet need to invest much more and take security seriously.

anigbrowl 1 day ago||
Agreed, but they're getting paid with our money.
mmaunder 1 day ago||
This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?
smartbit 1 day ago||
I love it too. Seemingly written by an LLM. nr378 https://news.ycombinator.com/user?id=nr378 is the responsible ‘author’ and actively reacting in this thread.
mmaunder 1 day ago||
Haha. Thanks
copperx 1 day ago||
An LLM, of course.
skeledrew 1 day ago|
Let them cry, I don't see anything changing unless they can somehow get China to agree. And I doubt China will drink any of that kool aid especially while they're being disadvantaged by export controls, so the ever-improving open weight models will continue to rain. This is something the US Big Tech oligarchs will NOT win.
More comments...