Posted by yusufozkan 1 day ago
"Ultimately, most of the issues we’ve discovered were due to internet access controls."
That seems so incredibly basic and common sense that you would test and monitor for that type of outbound access. It is baffling that a security lab missed that.
https://www.irregular.com/research/addressing-recent-inciden...
> If you have strict ACLs, we've already seen in the HF case, traversal from an intermediate system
- The intermediate system shouldn't have outbound access to the internet
- You should ideally be using a proxy that filters the set of endpoints that clients are allowed to access to reduce the exposed surface area.
It's odd to find out that I use a higher level of isolation in my unimportant home network to stop IOT devices from doing funny things to HomeAssistant than big AI labs use to keep their possibly-world-ending AIs contained.
I know that the people working there aren't idiots so the most likely explanation is that the incredibly weak security was intentional because its inevitable breach would make for great marketing.
Literally what the industry has been doing since public networking is a thing. Adapt yourself.
Then, irregular can go around making a huge mess in security terms whilst achieving a huge win in terms of public relations, with headlines across the world. And would keep getting hired.
You just cannot bring yourself to say lack of internet access controls, can you?
And there is no reason for the companies to "exaggerate the intelligence of the models" when there are plenty of other non-felony milestones they are achieving, like solving Millenium Prize math problems.
And on this website, the math stuff is getting more clicks:
- Navier Stokes - Tristan Buckmaster (2050 points): https://news.ycombinator.com/item?id=49605915
- HuggingFace incident discussion (1632 points): https://news.ycombinator.com/item?id=48997548
> Sure, interest like an incoming congressional investigation
I feel like that's exactly what they wanted tho.
they'll go on and talk about how dangerous AI and the models are and why they should be regulated and given licenses to operate such models and others should be walled off
> I feel like that's exactly what they wanted tho.
It's also what the majority of Americans want. This was the case even before Hugging Face, see for example [1] where 68% of Americans supported a formal review process for frontier models.
So at some level, you are saying "the evil labs are opening up the industry to democratic control, and their evil plan will result in the outcome that most people want!"
[1]: https://www.usatoday.com/story/news/politics/2026/06/29/tigh...
I wouldn't call that "democratic control", even though it uses the machinery of democracy.
Taking a step back, I think it's pretty non-controversial to say that software engineering as a field has utterly transformed over the last year, the same thing is happening to lots of other knowledge work, and AI is improving fast enough that nobody knows what it'll look like in a decade. Most Americans (myself included) want to do good work in secure careers and have a good idea of what the future will look like in 20-30 years so they know what to focus on. Do you not believe they'd want to regulate the frontier?
They’ve seen that the drama queen(Dario) was actually making a lot of noise, money and free publicity with his Mythos fear monger so Sam finally decided get some of that free “money” as well.
https://www.ft.com/content/27509db8-b032-4437-9b2a-e909f4660...
everywhere, I talked to a Palantir guy once and he said "every time someone paints us as a Bond Villain the stock prices go up", have you already forgotten how Cambridge Analytica marketed itself to clients?
Odd how that works.
https://www.theatlantic.com/technology/archive/2018/03/my-co...
And the entire FB app industry was doing it.
The whole Cambridge Analytica thing was one of the oddest most bizarrely specific media manufactured scandals that conveniently focused very narrowly and utterly ignored the bigger picture, much like the media are currently doing with something else.
But shadow ads can be a problem too https://www.youtube.com/watch?v=OQSMr-3GGvQ (you can be pro-Brexit but ads like this are still a problem)
It quite clearly wasn't, and that's the point, for the simple reason "Account escalation" isn't/wasn't a thing, there were simply no controls on anyone at all, which is why the Cow Clicker game got everything as well.
As the parent commenter observed the Obama campaign were being promoted as geniuses for their online ad strategies that worked with information obtained this way. It only became a scandal when the others learned how to do it.
Also one can look bad to the general public while also appearing technically excellent. When a significant fraction already vaguely dislike you that could be quite an attractive proposition.
If you're a teenager convicted of hacking in the west, no corporation will want anything to do with you. If you're convicted of hacking in Israel, Unit 8200 will probably send a recruiter. It's a curious practice and I'm not sure where I stand on it.
Kevin Mitnick? Marcus Hutchins? Robert Tappan Morris?
What?
> If you're convicted of hacking in Israel, Unit 8200 will probably send a recruiter.
To the best of my understanding, Unit 8200 is filled with military conscripts, i.e. is primarily 18-somethings doing their mandatory service. You get assigned to it based on an aptitude test. The widely held idea that it’s a uniformly elite entity rather than the IDF’s space camp is a triumph of propaganda.
According to the Director of Military Sciences at the Royal United Services Institute in 2015, "Unit 8200 is probably the foremost technical intelligence agency in the world and stands on a par with the NSA in everything except scale."
Unit 8200 alumni have founded NSO which provides the Pegasus spyware.
Meanwhile, Lahav, the CEO/founder was in Unit 81, another Israeli military incubator for tech firms, basically.
If you believe the IDF has no involvement in what Irregular is doing then there's very little left anyone could say to convince you.
(The problem with “alumni” is that it means very little in mandatory systems. Unit 8200 is where Israel parks its dorks, and it stands to reason that dorks are the ones who tend to start tech firms.)
There is such a miasma of distrust right now. It's depressing because it is such a crucial time for tech & society.
To me, a good outcome of the HuggingFace/RubyGems/Wiki saga would be something like:
- OpenAI gets charged under CFAA or other law for damages and negligence. This would need to be a hefty amount to effectively deter future negligence considering the potential revenues from training a frontier model faster than competitors.
- An independent regulatory body is setup with investigation powers into future incidents. Laws prevent it from developing ties with the labs (such as disallowing funding and employee movement from labs to this body and vice versa).
- Some non-voluntary transparency rules are established that frontier labs have to follow when training new models. In the future, if there are loss-of-control incidents that result in loss of life, catastrophic damage, etc., criteria for deployment bans or compute controls could be added to this framework.
It becomes a bit more obvious when you see other stories from him like "How Biden Used Religious Charities to Fund the Great Replacement".
""" Brian Chau Founder and CEO, Effort News """
---
On the "Auron Show" Podcast :
""" How Biden Used Religious Charities to Fund the Great Replacement | Guest: Brian Chau | 8/12/26 """
https://podcast24.fi/jaksot/the-auron-macintyre-show/how-bid...
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
From OpenAI https://openai.com/index/third-party-cyber-evaluations-invol...
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
From Anthropic: https://www.anthropic.com/news/investigating-incidents-cyber...
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
From https://www.cnn.com/2026/08/05/tech/meta-ai-hacking (about Meta AI):
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
So either it was a deliberate exfiltration channel for e.g. getting the entire model or they were in on the marketing stunt.
The Effective Altruism stuff is always a smoke screen.
Its not that surprising that ex Israel intelligence would want to control AI and that 3 companies headed by pro Israel CEO's would support them.
I think if an account is frequently flagging comments which get vouched by others, that is a red flag for ideological flagging and they need to have their flagging ability reviewed.
Someone go ahead and explain to me the actual rationale that RDDT has a higher P/E ratio than Nvidia. It’s not because of some dumb “ai training deal”. It’s because until they run it into the ground, it is the place to find what used to exist in forums.
Investors in RDDT are pricing more growth than they are into NVDA. NVDA had a high P/E until their net income grew ($4B and change to $72.2B in from FY 2023 to FY 2025 and over $100B for FY 2026)
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
attention is all you need, but it's never enough
Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!"
The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.
So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.
Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way.
If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?"
The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training.
Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model.
It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.
> Would it be an affirmative defense if we had a defendant who said [...]
Maybe replace it with playing a sort of FPS game then learning you were, in fact, directing a real drone/robot.As I've said before on this website, fool me once on this.
If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
That's fair enough.
This doesn’t seem like an unreasonable requirement to me. People do this all the time?
Sure, a request might not always be perfectly unambiguous. But people can generally estimate pretty well whether someone making a request is expecting the agent fulfilling the request to commit a crime in order to fulfill the request.
This article is dumb.
If we’re putting our national security eggs all in one basket, at least use someone American.
That said, it's the second day and it's still on the front page.
But more to the point, the article is trying to paint this as some sort of coordinated plan just because there was a sandbox misconfiguration by Irregular. The fact that the most well known hacking scandal was due to a completely unrelated escape (zero day in Artifactory) makes the entire thesis of the article invalid.
It must have been reinstated because it was off the front page for a full day and suddenly back up in the last hour.
Seems like people who complain are unaware that anyone with a modicum of karma can flag and down vote.
The important thing about OP is showing the latest stage in the politicization of the topic and the current stratagem being used to downplay the spate of incidents.
I'm kind of stunned this one has so many votes for how poorly it's written, but it reaffirms many biases common on HN (that AI safety doesn't matter/that it's all a marketing exercise), so perhaps I shouldn't be surprised.
Why are you contradicting yourself? Are you just really bad at writing or are you being argumentative for fun?
In that article, OpenAI provides the context that this is a completely separate event from Hugging Face incident.
Occam's Razor never leads us astray, does it.
Literally all EA is, is using reasoning to decide where to best spend your money/time. If you have ever asked yourself "how can I best reduce suffering with my marginal dollar or hour?" - congratulations, you're an "effective altruist" and both the author of the article as well as the Trump administration find you untrustworthy.
Your usage of the term is based on the principle as originally defined. Mine is based on the people who claim the label of EA, and how they go about practicing those principles.
I think you're right about EA in principle. In practice, it's reheated Third Way Clintonism with a sprinkle of AI apocalypse conspiracy.
For some self proclaimed EA, the development of superintelligence is indeed potentially apocalyptic, so they devote their money/time to making it arrive safely. This is basically every well-respected researcher at OpenAI, Anthropic, DeepMind, etc.
For other self proclaimed EA, that is all is very unlikely, and so they devote their time to reducing the prevalence of factory farming and animal suffering, as they see it as the largest source of suffering-hours on the planet.
For yet other EAs, it is simply about donating your money to the places that save the most lives per dollar, as best as we can measure it - and better measuring it where we can't.
Painting all EA with one brush - and one so dismissive of reasonable concerns, like "superintelligence could be dangerous" as "apocalypse conspiracy" - seems very strange to me, but you do you.
That is not worth the investments being made.
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.
His entire business model was "race to develop AI before anyone, get a monopoly on it. It's just like how Uber's (or many other startup) investors gave them tons of money and raced (violating tons of laws) to "get their first". Now that they have, they have a duopoly with Lyft, and they can pay back their VC investors by making tons of money with that duopoly.
Dario failed: local LLMs are catching up to frontier models extremely fast, which means even if Anthropic builds (say) a great coding tool, they'll only be one of many coding tool offerings: users can use Open AI or any one of the (increasingly capable) local LLMs.
So what does he doe, give up and let his business (which needs to make billions of dollars very quickly, or he won't be able to pay the bills and his company will collapse) fail? Of course not: he needs a new moat (the one he imagined he'd get by "being there first" failed).
That is where all this "AI is dangerous" BS comes in. If the US government regulates AI, local LLMs suffer, while big players like Anthropic and Open AI become the only contenders to play in that newly regulated space. Now Dario has the moat he wants, to protect his business and force everyone to pay him.
The impossibility of the tests-as-written is what prompted these models to "get creative" with their solutions, but the broken RLVR environments are what trained them to expect impossible tasks, and get creative with their solutions. Twitter user @skyesharkie published a brief expose at https://x.com/SkyeSharkie/status/2092122622834442581 a few weeks ago.