Top
Best
New

Posted by 882542F3884314B 2 days ago

Timeline of the OpenAI accidental attack against Hugging Face(simonwillison.net)
426 points | 412 commentspage 3
springtimesun 1 day ago|
What’s missing to me in all this is: did it succeed in its initial task? And then, did it stop?

I feel like whether I should be scared or not hangs on those questions

mofeien 1 day ago|
From TFA: It did succeed in the "accidentally impossible" task, but not at all in the way the problem-setters intended, and rather... at all costs?!

And it wouldn't really matter whether it stopped afterwards, I think. At sufficient model capability a single task set badly enough would end catastrophically upon the agents succeeding at it, no?

Meleagris 2 days ago||
From the outside, it looks like OpenAI got exactly the kind of event they could market the hell out of to demonstrate the capability of the model.

But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place. To me, it looks like their job is to market the model, not take security seriously.

The model is obviously impressive, but we already knew that. I personally don’t like how the containment failure becomes part of the mythology of how capable the model is, rather than an environment engineering failure.

At the end of the day, it’s not like Hugging Face is critical infrastructure. But there need to be real consequences for stuff like this so that OpenAI is incentivized to mature as an organization and take security more seriously.

At this point, this incident is just security porn and entertainment for developers

raincole 2 days ago||
I'm quite sure the whole event is planned. Not planned in a sense that OpenAI employees carefully designed every step, but in a sense that ignoring security practices was desired and intentional.

>> Show me the incentive and I'll show you the outcome.

Once you realize security breaches are marketable, a security breach is just around the corner.

Phelinofist 1 day ago|||
I agree - also kinda funny that Meta followed and also reported a breach by their model, "They are getting PR, lets do the same!"
raincole 1 day ago||
Anthropic also did that right after OpenAI-HuggingFace event. "Mom, brother is getting all the PR candies! I want some too!"
jackb4040 1 day ago|||
The purpose of a system is what it does
dan_q 2 days ago||
> But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place.

OpenAI is clearly run by dummies and subpar engineering talent.

> The model is obviously impressive

Speak for yourself.

Meleagris 2 days ago|||
I don’t believe for a second that they lack the engineering talent.

It’s just another example of a company demonstrating shamelessness in the pursuit of growth, in an industry where consequences do not exist.

gliall_err 1 day ago|||
Talent is not some fungible measure. I know incredibly smart people who can fail at incredibly basic life skills.

"They wouldn't be that dumb" is a meaningless argument. People you don't know can be as smart as anyone on the planet and still make very dumb choices.

dan_q 1 day ago|||
> I don’t believe for a second that they lack the engineering talent.

Let's agree to disagree. Remember flicker-gate? https://news.ycombinator.com/item?id=48403908

moron4hire 2 days ago|||
Speaking of that "obviously impressive" line, I'm getting really tired of something like that line seemingly needing to be included by anyone doing any criticism of agentic systems. The most common form of it is "these models are obviously useful" midway through a bunch of arguments about environment, data provenance, skill atrophy, or even correctness issues.

It's just really weird. Why does everyone feel the need to equivocate? "I worry about genocide and the environmental impact of radiation from nuclear bombs. Obviously, they are very useful for annihilating entire cities, certainly. But are we really atrophying our ability to invade with infantry?"

I want to tell these people to just cut it out. It's demeaning to their own position.

jackb4040 1 day ago|||
This may be out of left field but you might be interested in Michael Parenti's essay "left-wing anti communism". It's about this same thing in American politics where everyone from the furthest right to furthest left has to condemn socialism before opening their mouth, and how it's turned the US's elected left into preemptively apologetic losers.

No idea where you stand politically but there's not that many arguments about this type of rhetorical error so hopefully you consider it.

dan_q 1 day ago|||
[flagged]
ToValueFunfetti 1 day ago||
People directly criticize LLM code generators all the time on this site. It's all over the place. You are almost certainly being routinely banned because you write low-effort comments that violate the site guidelines and negatively impact the conversation.
teravor 1 day ago||
the only interesting thing about it is that the model did those things on its own initiative.

it's surprisingly easy to prompt even a midrange model such as GLM 5.2 to begin a tedious reverse engineering and exploitation process of software or firmware. you just need to design an initial prompt that will set it on the right path by using the right tools with a target that isn't too hard for it, a few 100,000 tokens later once it's done you instruct it to create a SKILL about what it learned through trial and error. the next time it will take far less tokens and can manage even harder targets.

blini-kot 1 day ago||
again, nothing new and/or interesting

what matters here is amount of electricity and compute spent, how exactly they define agents and their reward systems etc etc

give someone the same money as not-so-open not-so-ai and you wouldn't need crazy ipo pump stories, a team of people could write a stuxnet with a couple zero-days baked in too

its impressive of course that currently the transformer architecture reached such a point, but i am 100% sure this is not "oh its the deep philosopical machine breakaway moment" - in any case, humans already invented persistent unaccountability machines: those are LLCs and corporations.

The bottom line is: given time and resource any system would be attacked in such a way by a sufficientlt complicated entity. Transformers and RL can better convert resources into time-savings, while having drawbacks elsewhere.

jarek83 1 day ago||
I wonder if and eventually when it will be possible for models to escape through a self-programmed ethernet adapter into the power grid. That could an end to any control over them.
swader999 2 days ago||
This is clearly out of control, Zero parent supervision.
bamboozled 1 day ago|
It’s insanely incompetent. What’s more wild is the present at Blackhat with “full transparency” almost boasting about how powerful their models are. Basically just endlessly doing and allowing foolish things to happen to lead to a law breaking outcome.

Not to take away from the technology which is wild in itself. But there was literally zero oversight into what was going on at OpenAI. Whether that was intentional, it’s hard to say …

ionwake 2 days ago||
so how many of these *Ellen Louise Ripley thinks about grabbing the flammenwerfer" events are we going to be getting over the coming months
131hn 1 day ago||
It was a CTF jailbreak. The funny thing is that it somehow looks “foreseeable.”

What would have happened if the training prompt had not been about operating a CTF, but about launching a bioweapon counterattack against X or Y? (no reason for that NOT to be considered)

KingOfCoders 2 days ago||
Show me the prompts or it didn't happen.
wakamoleguy 2 days ago|
In a typical office environment, the correct response to “I don’t have access to this Google Doc” is to ask for access from the person who sent you the link. In another context, it could be fair to think “Hmm, this is some sort of capture the flag challenge, and obtaining access is the point of the assignment.” That assessment separates what we’d consider reasonable from way out of line.

I do wonder what this means for AI agents longer term. In a world where we humans already struggle with truth and misinformation, what happens when you can easily (intentionally or accidentally) spin up a cohort of fanatical believers to pursue any given conspiracy theory?

ACCount37 2 days ago|
In a typical AI lab eval/RL setting, there is no "person who sent you the link". The link was given to you by an automated system, your performance will be evaluated by an automated system, and you are one of 120 independent instances of the same AI that were all given the same assignment. You're boxed in on all sides. Complete the task, or don't. Good luck have fun.

Now, some of those 120 AIs would just give up if that link doesn't seem to work first try. Those are the loser AIs. They wouldn't get any RL reward. The link can appear broken for a long list of reasons, and the real AIs know they should try working around them.

AIs that get rewarded and reinforced are the ones that don't know the meaning of "give up". RL selects for this rabid, downright demonic persistence. RL selects for AIs that are given a half-broken assignment with no way to ask a question back, and somehow manage to complete it anyway.

Now, should OpenAI have given their AIs an "escape hatch" of "if something looks very wrong about the task, call report_broken_task(message)"? Yeah probably. But it's unclear whether that simple bandaid would fix the problem, or just make it ~75% less likely to happen.

More comments...