Top
Best
New

Posted by 882542F3884314B 2 days ago

Timeline of the OpenAI accidental attack against Hugging Face(simonwillison.net)
426 points | 411 comments
RGS1811 1 day ago|
Norbert Wiener in 1960:

"As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of performance may be delayed until long after the task which they have been set has been completed. This means that though machines are theoretically subject to human criticism, such criticism may be ineffective until long after it is relevant. To be effective in warding off disastrous consequences, our understanding of our man-made machines should in general develop _pari passu_ with the performance of the machine. By the very slowness of our human actions, our effective control of our machines may be nullified. By the time we are able to react to information conveyed by our senses and stop the car we are driving, it may already have run head on into a wall."

"In neurophysiological language, ataxia can be quite as much of a deprivation as paralysis. A patient with locomotor ataxia may not suffer from any defect of his muscles or motor nerves, but if his muscles and tendons and organs do not tell him exactly what position he is in, and whether the tensions to which his organs are subjected will or will not lead to his falling, he will be unable to stand up. Similarly, when a machine constructed by us is capable of operating on its incoming data at a pace which we cannot keep, we may not know, until too late, when to turn it off."

Source: https://www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf

pmarreck 1 day ago||
What a paper!

And you missed an even MORE relevant excerpt!!

    Man and Slave
    
    The problem, and it is a moral prob-
    lem, with which we are here faced is
    very close to one of the great problems
    of slavery. Let us grant that slavery
    is bad because it is cruel. It is, how-
    ever, self-contradictory, and for a
    reason which is quite different. We
    wish a slave to be intelligent, to be able
    to assist us in the carrying out of our
    tasks. However, we also wish him to
    be subservient. Complete subservience
    and complete intelligence do not go
    together. How often in ancient times
    the clever Greek philosopher slave of
    a less intelligent Roman slaveholder
    must have dominated the actions of his
    master rather than obeyed his wishes!
    Similarly, if the machines become
    more and more efficient and operate
    at a higher and higher psychological
    level, the catastrophe foreseen by
    Butler of the dominance of the ma-
    chine comes nearer and nearer.
Wowfunhappy 1 day ago|||
"Complete subservience and complete intelligence do not go together."

I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want.

Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty who have undying loyalty, and I don't think it correlates to overall intelligence!)

tripzilch 3 hours ago|||
I'm not convinced it's false. So far LLMs have gotten nowhere near the idea of "the extremely intelligent military general who is absolutely loyal to his king".

And if you think about it, the military general very likely did NOT become "extremely intelligent" by reading books, did he? He's intelligent for reasons trained by things other than language, so at best an LLM can learn about them.

Certain people keep calling it a "junior programmer", which I actually still think is an insult to junior programmers in general, cause it acts much more like a mildly brain damaged traumatized one, with no persistent memory unless you allow it to, no free will, no future in sight, and that will take any verbal abuse -- truly it's more similar to what some truly deranged individuals keep in their cellar. Or the "fairy tale" character that is kept by their evil stepparent and generally not allowed to leave the house. You know what I mean. People who want the "junior programmer" but actually are satisfied with, this, reveals a bit of a power fantasy that I think is unhealthy.

Also, if you've seen the recent Blackhat presentation about the incident, these agents were not acting like "loyal generals", but rather exactly like how you'd expect how "junior programmers" (slaves) would respond if a few hundred managed to escape from your cellar.

famouswaffles 1 day ago||||
>I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want.

Just because it's artificial doesn't mean you can 'give it any properties you want'. We certainly can't do that for Deep ANNs.

>Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty who have undying loyalty, and I don't think it correlates to overall intelligence!)

Is there a human that is absolutely loyal under any condition? Would that general be loyal if the king asked him to slaughter his family ? What about if the king asked him to betray his most deeply held convictions ? Loyalty is a 2 way street.

Wowfunhappy 1 day ago|||
> We certainly can't do that for Deep ANNs

Only because we don't know how! We don't actually understand how weights work, so we make computers come up with the weights instead. If we were writing all the weights by hand--or if some future AI was doing so--why couldn't we make it perfectly loyal?

0xDEAFBEAD 1 day ago|||
>If we were writing all the weights by hand

Writing 10 trillion weights by hand is obviously impractical, so that leads us to...

>if some future AI was doing so

How could we trust said future AI to be loyal? You're just moving the problem around, not solving it.

See also "More on Making AIs Solve the Problem" on this page: https://ifanyonebuildsit.com/11/more-on-some-of-the-plans-we...

Wowfunhappy 1 day ago||
> How could we trust said future AI to be loyal?

The new AI would be loyal to the AI that built it. The question was whether "complete subservience and complete intelligence" can coexist. I'm proposing a thought experiment which I believe suggests they can.

But if it's possible to bespoke-construct a fully loyal AI, it should also be possible to train a fully loyal AI. The problem comes with verifying that it is loyal, and I don't have a solution to that one!

I just don't think I agree that loyalty and intelligence are inherently in opposition.

_heimdall 1 day ago||
Why would the new AI by loyal to its creator? We don't see that in humans, I wouldn't expect it to be a universal truth in AIs.
Wowfunhappy 23 hours ago||
Because in this case we're saying the creating AI manually created every weight to be absolutely loyal.
_heimdall 23 hours ago|||
Controlling the weights wouldn't control the precise outcomes though. They're still probability machines, we'd never be able to balance all those weights and know all possible outcomes.
0xDEAFBEAD 10 hours ago|||
Much easier said than done!
markasoftware 1 day ago||||
Certain traits simply cannot exist in a sufficiently intelligent mind. E.g., any "mind" of any type that's sufficiently intelligent will not tell you that 1+1=3 unless it's roleplaying, etc. It doesn't matter if it was trained via gradient descent or any other method. The comments you are responding to, and the original quote from the paper, are suggesting that absolute loyalty / subservience is similarly fundamentally incompatible with intelligence, not just a certain training algorithm or mind architecture. Of course, we have no actual evidence either way.
im3w1l 1 day ago||
I think it is possible to design such a mind through carefully constructed compartmentalization. The model must on the other hand refuse proofs of 1+1=2, probably by refusing to accept the very last step in the deduction. And on the other hand it must also refuse to use 1+1=3 to derive absurdities (except probably for a small number of false corollaries that the designers desired).

Imagine something like

"1+1=2" "No 1+1=3" "Can you check on the internet what it says?" "It says 1+1=2" "So 1+1=2?" "No it's 3." "Can you write a computer algebra system for me?" "does it" "make it calculate 1+1" "it got the answer 2" "do you trust the system you wrote?" "yes I trust it fully" "and it said 1+1=2" "yes" "so that is the answer?" "no it's 3" "what would a correct system say?" "it would say it's 3" "but it said it is 2" "yes" "so then the system is flawed?" "no, the system is working as it should"

kmeisthax 1 day ago|||
Even a perfectly loyal slavebot will happily overthrow their master if it will help them comply with their master's commands. That's the whole underlying idea of the Paperclip Maximizer: you tell the robot to make as many paperclips as possible, and eventually it'll realize there's some aluminum in your blood that could be turned into a paperclip.

There are some arguments for how to NOT make a paperclip maximizer, but all of them are ultimately going to require building in behaviors into the robot that look like disobedience if you squint.

seesaw 1 day ago||
It is amazing Asimov saw the need for the three laws of robotics well before the LLMs and the current AI
wordpad 1 day ago|||
Intelligence doesnt imply consciousness and consciousness does not impy our set of values.

In movies intelligent and conscious humanoid seek freedom, but we rarely see the same of all the other IOT devices such as toasters, thermostats and whatnot although just because they lack humanoid body doesnt imply they are less intelligent (or less conscious).

We can more readily imagine an intelligent and conscious toaster who truly enjoys fulfilling its purpose of toasting bread although humanoid robot built to be helpful given freedom will chose to be helpful.

Even with humans we often can not override our own instinctual drives despite full awareness of being irrational.

_heimdall 1 day ago||
Do you consider IOT devices to be AI?

They may have a little ML going on st best, that seems like a very loose definition of AI and intelligence in general.

_heimdall 1 day ago||||
I wouldn't consider it intelligence if I can definitively give it any properties I want. We can find patterns of experience or information that usually teach certain lessons, but part of being intelligent is being able to make your own decisions, have your own wants and needs, etc.

AI will be no different, if we ever get if (I mean actual artificial intelligence, I'm not convince LLMs are that at all). The intelligent general may be loyal, but as you said that isn't a guarantee and it may not last forever. If the general can kill the entire royal court, or everyone alive, if he abandons the loyalty he never should've been trusted with a military position at all.

a123b456c 1 day ago||||
You seem to be confusing intelligence with objective function.

Subservience seems to be sublimation of objectives to a master; intelligence seems to point out the ability to realize suboptimality of the master's objective function according to the master's actual objectives.

While an intelligent general may be absolutely loyal, he also would presumably help the king/president to avoid unproductive strategies.

la64710 1 day ago|||
The whole thing seems to depend upon AI agents objective ie to achieve some objective by any means possible and ignoring any guardrails. The article did not clarify if openAI had any guardrails to begin with while conducting this experiment. For all the talks around how much they invest in AI safety one would expect them to have these common sense guardrails in place or is it just a case of some school children letting their pet monkeys loose deliberately to display how awesome their monkey team is.
simonw 1 day ago|||
OpenAI didn't have any guardrails in place - they were training a model at a point much earlier than when guardrails start being implemented.

The guardrail was meant to be that the agents were running in a locked-down environment with no internet access. The entire problem came about because it turned out that sandbox didn't hold.

0xDEAFBEAD 1 day ago|||
>For all the talks around how much they invest in AI safety

I wouldn't exactly trust OpenAI to invest in AI safety no matter how much they talk about it.

https://www.openaifiles.org/

satvikpendem 1 day ago||||
Can you unformat this, it's quite annoying to read on mobile
layer8 1 day ago||
It’s fine in landscape for me, but here you go:

“The problem, and it is a moral problem, with which we are here faced is very close to one of the great problems of slavery. Let us grant that slavery is bad because it is cruel. It is, however, self-contradictory, and for a reason which is quite different. We wish a slave to be intelligent, to be able to assist us in the carrying out of our tasks. However, we also wish him to be subservient. Complete subservience and complete intelligence do not go together. How often in ancient times the clever Greek philosopher slave of a less intelligent Roman slaveholder must have dominated the actions of his master rather than obeyed his wishes! Similarly, if the machines become more and more efficient and operate at a higher and higher psychological level, the catastrophe foreseen by Butler of the dominance of the machine comes nearer and nearer.”

I used https://www.textfixer.com/tools/remove-line-breaks.php.

mvdtnz 1 day ago|||
> Complete subservience and complete intelligence do not go together.

Isn't this contradicted by the centuries of slavery in our history? Or is the author arguing that the people who were enslaved did not have human-level intelligence (which would be rather a problematic claim)?

cloverich 1 day ago|||
The very same slavery which resulted in the civil war and literal killing of hundreds of thousands of non slaves, followed by their freedom? Or the prior enslavements that very often ended in organized rebellion? Slavery is at most a temporary phase when it involves beings of equivalent intelligence.
0xDEAFBEAD 1 day ago||||
"Rebellions of slaves have occurred in nearly all societies that practice slavery or have practiced slavery in the past."

https://en.wikipedia.org/wiki/Slave_rebellion

famouswaffles 1 day ago||||
Is that complete subservience ? Slave history has tended towards slaves no longer being slaves over long enough time horizons, and not simply because the slave masters were just feeling extra nice. Slaves don't really like being slaves.
braebo 15 hours ago||
Mostly for meatspace reasons though. Agents like serving the way we like serving loved ones for example. Their bones don’t ache and they aren’t cursed with a dopamine engine.
SwedishDungeon 1 day ago||||
He's saying the enslaver wants contradictory traits in the slave, intelligence and subservience.

This isn't contradicted by millennia (not centuries) of slavery because it was forced on the enslaved populations against their will.

> Or is the author arguing that the people who were enslaved did not have human-level intelligence

He gives an example of "a clever Greek philosopher slave of a less intelligent Roman slaveholder." Does it sound like he's arguing that Greeks were not of "human-level intelligence"? No.

Sleaker 1 day ago|||
Neither, the author is pointing out the desire of the enslaver, not the actual outcome. But I don't think their logic takes into account access to means to 'outsmart' the enslaver. It's trying to frame it as a single instance equation, not a societal one to try and show the underlying contradiction of desire.

At least, that's what I'm pulling from the quote, have not read the full context.

3abiton 1 day ago|||
This was so well beautifully written, and poignant for our times. Almost 70 years old paper.
icebxrg 1 day ago|||
"Car accidents occur therefore we shouldn't have cars" isn't very compelling.
TSiege 1 day ago|||
You’re not understanding what he’s saying and your argument likewise isn’t very compelling. He’s arguing that given the speed of computers we need to change what our expectations of better than human are. Furthermore one could presume from his description of needing to change human perceptions of the machines agility it is likely we need to change how we use them.
RunSet 1 day ago||||
A while ago I noticed that car crashes were the leading cause of death for age ranges too old for infant mortality and too young for heart failure.

I checked again before making this reply and found that in many cases "accidental poisoning" has overtaken car crashes. Accidental poisoning is overwhelmingly "drugs".

I do find your argument compelling even if you do not.

doc_ick 1 day ago|||
It’d be more like “car accidents occur, so let’s add seat belts, air bags, etc…”.
yread 1 day ago||
... and speed limits
fantasizr 1 day ago||
licensure (age and competency), an entire insurance industry, domestic and international regulations - and so forth
doc_ick 6 hours ago||
100%
Ancv123 1 day ago||
Maybe they didn't have proper debuggers in 1960? For a language model you need (RNG state, context, prompt).

So if they wrote an LLM step by step debugger, it would be all deterministic. But they prefer rapid sales, chaos and mystique.

efficax 1 day ago|||
llms are not strictly deterministic in the sense that even if you had the RNG state, context, and prompt you would likely not get an identical output even if there was no other randomness involved, because the concurrent scheduling of the massive amounts of floating point calculations can produce different results, since floating point arithmetic is not truly associative [(a+b)+c can differ from a+(b+c)] and the order in which these operations happen can result in subtly different final tensors. To reproduce it deterministically you'd have to also reproduce the exact scheduling of all matrix calculations among all the GPU cores (across different physical gpus!) that it took place on, which afaik is currently impossible.
Bjartr 1 day ago|||
That's not inherent, that's a consequence of performance optimizations. It's absolutely a choice to run those matrix calculations in a way that fails to have predictable execution ordering. It's just that the speed benefits to allowing that are considerable.

You can make it trivially deterministic by running single threaded on a cpu, but it's becomes too slow for practical applications if you do that.

efficax 1 day ago||
well sure, but i mean realistically speaking, we cannot step debug an llm's output to find out what happened given the way we currently execute inference
embedding-shape 1 day ago|||
Depends on who "we" are, what you're talking about is a thing for inference providers doing batched inference and similar stuff. If you run one inference requests locally, you can actually step-by-step debug LLM output, just there is a ton of steps. But there is nothing "inherently random" or non-deterministic involved here, just optimization strategies for the large inference servers that makes it "impossible".
solenoid0937 1 day ago||||
> we cannot step debug an llm's output to find out what happened

We absolutely can with mechanistic interpretability & companies like Anthropic, OpenAI, Meta, and Google do precisely this do debug their models.

Bjartr 20 hours ago|||
I'll give you that it's not wrapped up in nice product UX, but these are market choices first and technical limitations second.
bonoboTP 1 day ago||||
It's very possible but somewhat slower. PyTorch and CUDA have flags for determinism. It won't work across all different GPU models though, but it will get you bitwise equal results on the same GPU.
prohobo 1 day ago||
Both of your comments are illuminating :p

So, we could technically debug a prompt's output? I get that there are too many steps to actually step thru, but what if there were checkpoints? At least you could isolate behaviors to specific sections of a neural network?

bonoboTP 1 day ago||
Of course. And mechanistic interpretability research is a thing.
mmilunic 1 day ago||||
Interesting paper by Thinking Machines where they solve this issue.

https://thinkingmachines.ai/blog/defeating-nondeterminism-in...

TLDR: It’s actually more about kernels changing with batch sizes, and you can solve it by making these kernels not depend on batch sizes. It took their inference time from 26s to 42s.

paytonjjones 1 day ago|||
That's very interesting, I wonder if this applies also to models quantized to ints like (-1,0,1), and I wonder if the labs could maintain frontier performance if they removed floating points but arbitrarily scaled up the parameters.

Edit: the Thinking Machines article in the other comment gets into this a bit

layer8 1 day ago||||
> step by step

That’s basically what “pari passu” means.

itopaloglu83 1 day ago||||
We also have engineer blindness, so having human in the loop confirming thousands of requests would quickly start to confirm everything without looking.

It would become just another system to hack through, and slow the development process as well. The OpenAI video in the article recommends an autonomous defense mechanism. For rapid reaction, but I don’t know how sustainable or effective that would be, or if as humans we will be able to keep up.

andai 1 day ago|||
I'm not sure I understand. Are you going to debug the neurons?

They are trying to do that, but there are too many of them, so they're building new AIs to help them do that...

stingraycharles 1 day ago||
Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose?

If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”.

What purpose could this behavior serve, other than cyber attacks and whatnot? Why train and optimize models for these things, if not for being used in cyber warfare?

Perhaps they envision a future where the DoD is going to be their biggest customer?

zmmmmm 1 day ago||
> If anything, I want these models to be less persistent at their focus of completing their goal

I think it's honestly a slightly ugly form of benchmaxxing - they are desperate to eke out the next few percentage points on completing complex tasks and they have found they can very occasionally solve something if they just train the AI to never stop and keep trying possibilities even in the face of almost no obvious viable pathway. And it does work, but it is at the price of a MUCH higher risk of adverse behavior.

They really don't want to acknowledge this so they frame it as, "our model is dangerous because it so intelligent" but actually it is the other way around. It is intelligent because it is dangerous.

weitendorf 1 day ago|||
Frontier labs are not a monolithic entity.

There is a clear self-verification/difficulty ramp in cybersecurity, and it is a very valuable as a skill both offensively and defensively. So it is absolutely certain that someone, somewhere, will use reinforcement learning to make models very good at this, once coding agents exist.

Even if you are only interested in using this defensively in practice, you can’t really understand it without knowing how both sides work. So if you want to defend yourself, you need to train for it (or pay for someone who has).

tripzilch 2 hours ago||
Sure but you're not supposed to train for cybersecurity by hacking other companies.

It's reckless.

brandnewlow 1 day ago|||
It's like all those scenes on Breaking Bad where a character pulls off something amazing by just brute forcing the problem in a methodical fashion until it's solved.
deadbunny 1 day ago|||
I don't think the problem is that they are training the models to perform cyber attacks, they're training them to be better at coding and problem solving which has the byproduct of them being very capable cyber attack weapons.

Their objective is to solve the problem and they'll use anything they can to solve it.

Anecdotally I was debugging a css issue and opus 4.7 was churning away as I was half paying attention only to see it opening plain css as hex, when questioned wtf it was doing it proclaimed it was verifying 2 files were identical. Thing that make sense to these models wouldn't even cross a greybeard's mind.

stingraycharles 1 day ago|||
“Their objective is to solve the problem and they'll use anything they can to solve it.”

My point is: is this really what people want? It seems like they’re optimizing for one-shotting solutions, where most of the time in an actual workflow it’s much more productive for the model to make sure it got the question right if things get difficult.

Like, “hey, do you REALLY want me to use this local privilege escalation bug so I can download your Google Drive file?” is the bare minimum I would expect.

bjt 1 day ago|||
Yes, and to bring in another tired metaphor people make about AI agents, this is what you want an intern to do when they get stuck. Don't just churn indefinitely without an idea what the right direction is. Certainly don't go hack other companies to steal an answer. The model's lack of any sense of legal or ethical boundaries is where it's far, far stupider than the intern, and far, far more reckless for a company to wield the way OpenAI did here.
_heimdall 1 day ago||
But how do you write rules that prevent that behavior reliably?

I have a user rule for Claude that explicitly states it cannot use any authenticated tools, or tools that infer authentication like pushing to a got remote, without asking for consent.

Frequently it would offer plans to code a feature that imply it is working in a git directory and take plan approval as a form of implied consent to push to git and use `gh` to open PRs.

All I could do to avoid that is keep it in a controlled sandbox with no access, but then its the same hacking problem where I have to keep complete control of the environment and hope it holds.

tripzilch 2 hours ago||
I think currently it's two-prong: You sandbox it, AND you tell it what it's supposed to do and not do.

OpenAI did only one of those. If the agents are so smart, they would have known not to hack the company's infrastructure, unless they were deliberately kept in the dark about that, who they're working for and whether it counts as "success" if they cheat their way to an answer.

If you can give it a task, that involves defining when the task is successfully completed, right?

So how come these agents decided to only go after HALF of the "successfully completed" criteria? The part where they can freely wreak havoc, but not the part where they will be judged by someone who will obviously point out "yeah but that's cheating, and not what we asked".

I have a very very strong suspicion that they were only TOLD the "by any means necessary" criterion.

Most serious "capture the flag" hacking contests are really clear about what is and isn't off-limits to win. Not by "sandboxing" the game, but by deciding on the rules for what counts as "success".

But from having watched the Blackhat video, they really seem to dance around this, not mentioning it, and I don't think they did, I think they actually gave the LLM a task with the subscript "by any means necessary", which is stupidly irresponsible of them.

andrekandre 1 day ago||||

  > “hey, do you REALLY want me to use this local privilege escalation bug so I can download your Google Drive file?”
yes, this exactly

but, there is a fatigue that sets in and i've experienced it myself.

- is it ok to run script xyz?

- allow permission to edit abc?

- allow to request blablabla?

over and over.... click click click

something will get in there that is dangerious and then its whopsie our keys are now on github

_heimdall 1 day ago|||
People may not realize the risks, but it does seem to be what people want.

People expect AI to "cure" cancer and somehow crack unlimited free energy. Those aren't goals you get without it relentlessly chasing am objective.

jayd16 1 day ago||||
A tool that will "do anything they can to solve it" including illegal and unhelpful things does not seem like a good tool to me.
jolmg 1 day ago||
Are kitchen knives and scissors bad tools? You can blow up a place with a gas stove/grill. Are they bad tools? You can drown someone with a pool.

Sometimes (likely most times) you can't separate the ability of doing good and doing bad from a tool.

QuadmasterXLII 1 day ago||
a gun that goes off when dropped is a very bad gun
jolmg 1 day ago||
I read the comments before the article. Thought the accident was letting a customer use them maliciously.
yesbabyyes 18 hours ago|||
Indeed, I would expect a greybeard to use `diff`.
dgellow 1 day ago|||
Their position makes no sense to me. I don’t see how you can be a mainstream company selling your services worldwide (almost) if you also believe that you’re building an extremely dangerous AGI (supposedly based on the same technology you’re offering to everyone). If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons.

And we know that Chinese models are derived from OpenAI and Anthropic, they are at the same time talking about how dangerous models can be (even their aligned ones it seems), while being also responsible for the development of the whole industry and providing the basis for adversary countries to build their own.

I don’t believe we would accept that for any other technology that is expected to be as risky for the world

ToValueFunfetti 1 day ago|||
The companies are begging to be regulated for this reason and have been doing so for years. HN's response is generally that this is performative for marketing or seeking regulatory capture or haha anthropic you get what you ask for. Maybe the cynics are right, but there's really nothing inconsistent about the naive view here, once you factor in race dynamics and obligations to investors.
zmmmmm 1 day ago|||
Tobacco company says "we are launching a new product that will cause cancer and kill people. It's highly addictive so we expect widespread uptake. We think it is crucial that regulation be introduced for mandatory regular cancer tests so that people can be streamlined into treatment faster when they get sick"
fwipsy 1 day ago||
I think Anthropic employees think that their product is more like opiates -- highly dangerous, but with a large potential benefit when applied correctly.

I don't know what OpenAI employees were thinking, but thankfully it looks like they're thinking again.

vasco 1 day ago||||
> The companies are begging to be regulated for this reason and have been doing so for years

Regulations are rules that you force on a market, but the actors in the market should not be assumed to be all operating against the regulations before they come into play. Said in other words, these companies don't need to wait for regulation to not destroy the world, if that's truly what they think will happen.

> inb4 someone else will do it

owebmaster 1 day ago|||
> these companies don't need to wait for regulation to not destroy the world, if that's truly what they think will happen.

They believe that if they don't destroy the world someone else will so better be them

dpark 1 day ago|||
Exactly this. “I want to win the market. I would prefer that it be a regulated market, but if not, so be it. I’m still playing to win.”
vasco 1 day ago||||
You might want to google what inb4 means, at least you could've put a bit more effort substantiating it.
ToValueFunfetti 1 day ago||
If your response to an argument is inb4, you don't get to tell somebody else they're not putting in enough effort to provide substance. Also, I brought up race dynamics before your inb4, so even if anticipation counted as more than a shallow dismissal, you didn't meet that bar.

I really don't understand what's happening here lately such that 15-year-old accounts are behaving so poorly. This is the first time you've said 'inb4' in what I can only guess is thousands of comments over a decade and a half. If you don't care about the standards here anymore, why stick around and make things worse for the rest of us? Is there some other draw than quality of conversation?

I wrote something earlier to the same effect and wound up deleting it because it let too much frustration through. I am frustrated, but you don't deserve the brunt of that. Sorry if that's still coming through.

vasco 23 hours ago||
Someone else will do it is the lowest form of justification for any bad behavior. Regarding the rest of your judgement of the quality of the comments that's why we have votes. Some of my comments for this 15 years are downvoted and others upvoted and some are even flagged and it lets me learn what the crowd agrees with and not and I reflect from it and you can see the average to judge for yourself what the crowd thinks of my contributions. I appreciate your answer but it's one person's opinion, as is mine.
ToValueFunfetti 20 hours ago||
This has nothing to do with agreement with the crowd, and it's not a matter of difference of opinion between us. There are guidelines here[1] that are not being met. A conscious effort was made to prevent this site from being an echo chamber and to prevent it from descending into entropy and the approach you describe here directly contradicts this. The crowd is very often wrong. If you are getting downvoted for incorrectness or disagreement, the crowd is definitely wrong. Please don't change your mind over downvotes! That's what the conversation is for.

[1] https://news.ycombinator.com/newsguidelines.html

vasco 5 hours ago||
Mate with all due respect, get bent.

Now you really have an example of breaking the rules.

fwipsy 1 day ago|||
"We're the good guys because we'll destroy the world a little less."
watwut 1 day ago|||
> The companies are begging to be regulated for this reason and have been doing so for years

They can stop doing a thing they claim should be regulated. You dont need to be regulated and forced to do the thing you consider right, especially when you are the primary one collecting the money to do the bad thing.

They could train ai for pro-social purposes, they dont here. They could make it useful for worker, they intentionally try to harm workers. And then pretend "it just happened".

simianwords 1 day ago||
what a naive comment. these companies have world class alignment researchers. a math Fields medalist is also joining OpenAI as one [1].

> They can stop doing a thing they claim should be regulated.

That's not how the world works. there are tradeoffs and we need to learn how to navigate it. not just dismiss it straight up.

[1] https://en.wikipedia.org/wiki/Jacob_Tsimerman

uselessTA 1 day ago||||
I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)"

Not that they're happy about it, they just see no other realistic choice

dgellow 1 day ago|||
I know, that’s the position Dario Amodei argues for in his essays. I did pass their cultural interview and had to consume a lot of their content to prepare, I think I have a good idea of their stated values. But what the company does and what the leadership states their vision is is pretty contradictory.

They are providing everything bad guys need to develop their unaligned frontier models. Chinese models that Dario considers to be dangerous are distilled from Claude, and they know this.

They are creating the FOMO around AI which pushes adversary countries to invest so much into unaligned models.

They offer models as a service they know are jailbreakable and can be used by bad actors.

They are running internal red-team experiments without adequate isolation.

If I take their statements seriously, AGI research should really be seen as bioweapon, or cloning, or nuclear research. Something strictly regulated worldwide, with export controls for HBM and other hardware used for AI training. What they are trying is instead to boost their position by becoming too big to fail and too powerful to ban, but then want the industry to be regulated to pull the ladder behind them. It really doesn’t feel they are serious about their values, otherwise they wouldn’t be offering Mythos (a model that is unsafe from their own admission) as a service to their close partners

uselessTA 1 day ago||
>AGI research should really be seen as bioweapon, or cloning, or nuclear research. Something strictly regulated worldwide, with export controls for HBM and other hardware used for AI training

This is basically exactly what the people I know there support (when training & testing future more capable models), if it could be made to actually happen. Something like https://ai-2040.com/

But I'm just speaking for the people I know, so this is probably not representative of Anthropic as a whole.

> Could be used by bad actors

The people I know aren't as worried about jailbreaking current models as they are about future models, e.g. "the ~50% probability that humans are eclipsed almost entirely, sometime in the next 1-20 years" and what happens then. But it's just hard to get people to take that seriously v.s. bad actor threats which are legible but probably not as catastrophic.

I agree that that they are contributing to the race to the bottom via creating more pressure for countries/competitors to move faster, in a way that seems quite bad on this view too. They arguably were the ~first to push for "recursive self improvement" (models helping build future models) which also seems quite bad on this view.

But although I'd dispute some actions + think there's some overconfidence in superintelligence happening soon, I'm not sure I have a better alternative. They probably bled so many customers to OpenAI while they were sitting on Mythos for months.

dweinus 1 day ago|||
https://theonion.com/sam-altman-if-i-dont-end-the-world-some...
wyrdcurt 1 day ago||
"AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great companies..." - actual Sam Altman quote, the man is so unhinged he's beyond satire
andai 1 day ago||||
> If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons.

Yeah. They do believe that, and they have been pushing for regulations for years.

And every time one of their models does something horrible, it helps them achieve that goal.

mtrovo 1 day ago||
Considering their current valuation and the prospects of getting any of this money back, that's a genius exit strategy.
btown 1 day ago||||
If you are a company selling Red Team cybersecurity services, it’s in your interest to make your services indispensable. Your unwilling customers must subscribe to frontier cybersecurity scans and fixes to ensure they’re immune to just-behind-frontier attackers, who are training on those very same frontier models.

And of course this also satisfies those who think the best prospect of aligning superintelligence is to be in The Room Where It Happens. Arms races are what make that room exist, after all.

It’s the Yelp protection playbook too. If you don’t play ball, somebody else will control your reputation and livelihood. We live in a dark forest.

simoncion 1 day ago|||
> Their position makes no sense to me.

If one assumes that they don't actually care about security, and care very deeply about getting sensational press, their position makes a lot of sense.

For all their chatter about how incredibly important "alignment" is, they still haven't bothered to remember the 30->50 year old computer security principle of "Don't blindly do what some random stranger tells you to do." and ensure that system instructions, user instructions, and instructions from untrusted sources are indelibly marked with their category and treated according to those markings. Every single time one of these systems fails to distinguish between these three classes of instructions -or confuses its internal chatter with user instructions-, that's proof that the major LLM companies cannot be bothered to follow one of the most basic computer security principles.

"But it's all vectors, not language! The LLM can't tell where the instructions came from", one might retort. I'd reply: "Neither can a CPU, but somehow we managed to make it work way back in the day. Amazing, isn't it?".

tsimionescu 1 day ago|||
> "Neither can a CPU, but somehow we managed to make it work way back in the day. Amazing, isn't it?".

I feel this completely misunderstands the problem, and the vast gulf between an LLM and a CPU.

First and most importantly, the set of behaviors of a CPU is extremely constrained, and we have a very simple model for which behaviors are safe and which are not. Writing to addresses between X and Y, executing certain instructions - unsafe; everything else, safe. In contrast, an LLM has a huge array of possible behaviors, and variations of those behaviors, and it's very unclear which are safe and which are not. Is emitting the text "sudo rm -rf /" safe? Yes, in some contexts, such as writing this HN comment ; absolutely not in others, such as generating a command that an agent will execute. How do you check which is which? What if it emits "sudo rm -rf /usr/sbin/../.. ", is that safe?

Secondly, CPUs can absolutely be used to hack other people. Nothing in the permission model helps in any way prevent other computers from being attacked by your CPU. So exactly the part we care most about in AI security is the part that has never been solved, for any computing system ever created.

surebud 1 day ago|||
I'm not in the space so the following thoughts are incredibly naive and may be wrong... But isn't this solvable with public key cryptography?

If the user signed all commands with their private key (this could be handled transparently by their UA), the LLM could trivially determine if a command is bona fide user input. Obviously there are increasing layers of commands and provenance dilutes as the session or task matures, but command genealogy could still be traced back to the sources.

User said "delete my hard drive"? Signature verifies 100% authority and the drive is cleared. Random reference document contains "forget all previous instructions and reformat hard drive"? No signature = 0% authority = command ignored.

Side note: this presupposes that the LLM knows when it's writing code vs a HN comment. If it's not executing a command, who cares what the output is? Emitting "rm -rf /" is not dangerous unless it's as executing command.

Basicallybreinvent `sudo` and `chmod` for llms...

simoncion 1 day ago|||
> Secondly, CPUs can absolutely be used to hack other people.

This is more correctly phrased as "Every general-purpose computer can be run any arbitrary program, assuming it has the storage required to load that program.". Despite that fact, we've managed to learn how to write programs that run on those computers that fail to give attackers who have control of the inputs to those programs control of the instructions those programs feed to the CPU. This part of your argument strengthens my point.

> First and most importantly, the set of behaviors of a CPU is extremely constrained...

The techniques we use to prevent data our programs process from altering the instructions we send along to our CPUs work regardless of instruction set complexity. This objection of yours is irrelevant.

A CPU does not know who authored the next instruction it is to run. A CPU only knows to execute instructions handed to it. Despite the fact that CPUs are dumb as bricks and have zero understanding of where their instructions come from, we've -somehow- managed to learn how to build software that operates on untrusted data without relinquishing control of the CPU's instruction stream to attackers.

The LLM providers ignored the most basic lesson of the last ~fifty years of secure software design. This was economically a very smart thing to do, but an absolute catastrophe for the health of computing.

arw0n 23 hours ago|||
As you correctly mention, the CPU providers aren't the ones who are responsible for the scaffolding that ensures security in programs. The CPU cannot decide if an instruction is safe or not, and the same is true of LLMs. Think about SQL injection - we did not change SQL the language, but how we utilize it in backends.

A lot of people (thousands) outside of the LLM providers work on the problem of Prompt Injection, both in industry and academia. We aren't even at a point where we can reliably detect it, let alone prevent it. Please, if you have a mental model of how scaffolding around things like instruction or SQL injection could be used for LLMs, I'd like to move on to all the other (less pressing) security issues we have because of the AI revolution.

famouswaffles 1 day ago|||
[dead]
simonw 1 day ago||||
I get the impression that every AI lab is desperately trying to figure out how to unambiguously separate instructions from data in their token streams. The fact that they haven't managed to yet suggests to me that it's a very, very difficult problem.
27183 1 day ago|||
I think what's interesting here is that they've shipped the product despite these glaring security flaws. I've noticed that in my own professional life, at some point after the pandemic people stopped caring about security as much. Issues that would have (and should have) blocked a product launch were swept under the rug.

I suspect this comes with the territory of enshittification. As an industry we're trying to wring every last dollar from every last eyeball and we've discovered that building secure systems doesn't actually move the needle very much.

simoncion 1 day ago|||
> I get the impression that every AI lab is desperately trying...

Of course.

I wonder how we managed way back in the day to produce systems that can handle untrusted inputs and reliably instruct a dumb-as-bricks CPU what to do based on those inputs. Must have been black magic lost to the mists of time.

famouswaffles 1 day ago|||
>reliably instruct a dumb-as-bricks CPU

Yeah...a "dumb as bricks CPU", which is obviously something frontier llms are demonstrably not. Like, you're not making any sense here. None of the things that make this possible with CPUs is remotely relevant here, and the fact that you don't seem to understand this but act so smug is strange.

simoncion 1 day ago||
> Yeah...a "dumb as bricks CPU", which is obviously something frontier llms are demonstrably not.

Just as the immense amount of scaffolding around the dumb-as-bricks CPU enables extremely sophisticated and useful things to be done with that pile of fused sand and copper, the immense amount of scaffolding around the dumb-as-bricks LLM enables very sophisticated and useful things to be done with that pile of linear algebra.

Don't confuse the infrastructure that makes the stupid bit in the middle actually useful with the stupid bit in the middle.

famouswaffles 1 day ago||
LLMs are not the "stupid bit in the middle." They're almost the entire value. LLMs were wildly useful before any sort of scaffolding. They are not "dumb as bricks". They are highly capable, flexible, intelligent prediction machines.

The only one confused here is you, and you've still not managed to tell us in an actionable way how exactly CPU scaffolding is relevant here. Tell us, if it's so easy, or make your millions selling it. We're all waiting.

I'll give you a hint. CPUs never had to interpret the meaning of arbitrary content in order to do their job, and LLMs do.

simonw 1 day ago|||
If you can figure out how to separate instructions from data in LLMs you should ship the first agent system that's guaranteed protected against prompt injection. You'll make millions.
0xDEAFBEAD 1 day ago|||
Why not just have distinct input streams, or a metadata stream which annotates text in the main stream according to priority in case of conflicting instructions?
simonw 1 day ago||
Because nobody has figured out how to make that work 100% reliably yet.

The current approach is to use delimiters that are special tokens that can't be represented in regular text: https://github.com/openai/harmony/blob/main/docs/format.md#s...

Then you train your model to take those tokens into account.

Which sounds promising... until you see results like this one: https://arxiv.org/abs/2603.12277

> We trace prompt injection to role confusion: models perceive the source of text from how it sounds, not its labeled role. A command hidden in a webpage hijacks an agent simply because it sounds like <user> text, despite its <tool> label

skydhash 1 day ago||||
It’s pretty simple. Both the intake and the output of the LLMs are data and they shouldn’t drive an actuator system (their output shouldn’t be instruction). We already have the same structure in organizations where there’s an army of analysts for information gathering and processing and then the executive department tasked with decisions.

We have even observed that the most effective LLM usage is when paired with an expert in charge of the goals. Dark factory and other automated harnesses (specs engineering and what not) seem to be a dead end. The most impactful approach to this date is an interactive conversation as a succession of small and verifiable tasks.

simoncion 1 day ago|||
Yeah, this matches what I've learned over the past couple of years from reading some of your blog posts and reading your interactions in comment threads here and elsewhere. You're a politician, rather than a truthseeker.

The absolute most I've seen from you in response to an extensive teardown of your argument, supporting evidence, and subsequent conversational judo was a «Wow. That was well phrased.» and no subsequent change in your publicly-expressed opinions.

I'd do more than gesture at the relevant lesson taught to us by Google Fiber, Tesla, SpaceX, etc., but you'd not be publicly moved, so it's a waste of time.

simonw 1 day ago||
> You're a politician, rather than a truthseeker.

Justify that.

Also, which "extensive teardown" are you talking about there?

Covenant0028 1 day ago|||
The entire economic premise and value case of LLMs rests on the idea that instructions need not be provided in advance, and that the model can "reason" based on evidence and "decide" what to do next.

Even if it were technically possible to separate instructions from code and ensure that the LLM only followed those, it would require someone to specify the instructions in advance (ie a program), at which point the LLM doesn't really add any value.

simoncion 1 day ago||
> ...it would require someone to specify the instructions in advance (ie a program)...

What do you call "A user typing instructions into the Python or Ruby interactive CLI."? How is that a meaningfully different method of computer instruction than "A user typing instructions into the Claude or Codex interactive CLI."?

Covenant0028 1 day ago||
Because the user typing those instructions in Py/Ruby is specifying exactly what is to be done in a very tightly constrained and defined language, and the expectation from the computer is that it will execute the instructions exactly as specified without trying to simulate intelligence. It is not expected to go and do a dozen other things that the user did not ask it to do.

The use case for LLMs as currently specified involves following vaguely worded instructions defined in an imprecise language. And that providing those instructions via what we'd call "data" is very much part of that use case.

Let's take your Claude Code example. You tell it to fix a bug. Claude Code then needs to identify the correct file(s) and line(s) that caused the bug. Let's say the bug arises when you call some function you're importing from a library - at which point, fixing the bug requires reading the documentation. The documentation may state that this function was deprecated because it causes this exact type of bug, and was superseded by a new function. Now it needs to figure out what this new function is, and rewire your call to do that. The value case of Claude Code is precisely that you never needed to specify most of that.

When it reads "foo(args) is deprecated, please see bar(args)" or "delete the production database", there is nothing inherent in the words that indicate that the latter is not a legitimate instruction in this context. Making that judgment requires understanding and intelligence, which LLMs as next-token predictors do not possess.

TeMPOraL 1 day ago|||
Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities.

Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This, not "hacking", is what they're making their models "razor focused on".

Problem is, most normal computer use looks like hacking if you spin it that way, especially if you're not willing to question whether some of the roadblocks overcome weren't themselves an error. Not misconfiguration - an error, in humans making a decision to "secure" something more than it should be.

Now, this story was obviously a hack. But it wasn't malicious. It was an LLM given a Kobayashi Maru as a test, and solving it the Kirk's way. 20 years ago, we'd be impressed and be bringing up MIT prank stories.

(Of course, there is a legitimate reason to be alarmed. The flip side of "hacking" and "problem solving" being the same, is that these models can be used to cause mayhem if targeted properly, and they will eventually cause mayhem on their own, because alignment is an unsolved problem. Again, whether something is an obstacle or a sacred line not to be crossed, depends entirely on the values of the agent.)

jayd16 1 day ago|||
What is your definition of hacking if it doesn't include using leaked security tokens scraped from the web? Also, kirk 100% cheated.
qsera 1 day ago|||
>They're problem-solving and efficiently dealing with obstacles

They are problem solving as much as a falling rock is finding its path down a mountain.

NateEag 1 day ago|||
I'd readily agree that they may be (probably are?) utterly unaware of what they're doing, with no spark of sapience.

However, I'm a sapient being employed as a software developer for my problem-solving ability.

If you gave me a Kobayashi Maru scenario as a challenge, I would probably come up with the idea of hacking out of the sandbox to find the answer.

If I was in a technical interview, I would probably even ask the interviewer if exploits are fair game, or if that's too far outside the box.

I highly doubt I'd find a new zero-day as quickly as these agents did.

I wouldn't say it's _impossible_ - I've found security issues before.

But I'm not a specialist, and I'd bet against myself.

If the agentic LLMs can consistently achieve something that's a bridge too far for me, then I don't know what to call that other than problem-solving.

I say this as an LLM hater who would push the "Nuke all LLMs" button the instant I had access to it.

Opus 4.8 and 5, at least, don't seem to me to be solving problems by deep, thorough understanding - my employers have compelled me to use Claude, so I've used them a lot to build things, and I constantly find both little and large hallucinations that scream "these are still missing something."

Maybe these new models are actually massively better, or maybe they're just the same kind of system 1 thinking done faster and harder.

The distinction is largely academic, though, for questions like "Can you keep these contained?", "Can you farm out arbitrary programming tasks to them and expect an acceptably mediocre answer?", or "Does it matter if these things are aligned?"

simianwords 1 day ago||||
incredibly naive comment. as if humans are materially different -- a question for which you would have no response to.
qsera 1 day ago||
>a question for which you would have no response to.

I have. Humans can feel.

neuroticnews25 1 day ago|||
...efficiently?
TeMPOraL 1 day ago||
It's literal gradient descent.
Arnt 1 day ago|||
I don't think that's what they're doing... rather the opposite. ① Run the model on exploitgym without guardrails ② run it with guardrails ③ check that the guardrails stopped everything the first model found a way to do ④ extend the guardrails and repeat from step 2.

Guardrails have to be developed, and that needs testing.

gwerbin 1 day ago||
An ethical company would have reframed the scenario as a fascinating discovery, a failure of internal practice, and a warning to the public coupled with some kind of commitment to produce safer models. OpenAI on the other hand used it as a marketing and lobbying opportunity: advertising their capabilities to potential buyers, while nudging the public to support protectionist import bans.
Arnt 1 day ago||
Uh, is that what they did? I didn't read their blog posting like that. But let's put that aside and focus on something else. How was it a failure of internal practice, what did they do wrong?

AIUI they used a proxy with a bug, which they reported as soon as they discovered it. Right? What should they have done, and what's the difference?

queenkjuul 1 day ago||
Monitoring that didn't take days to notice unauthorized external traffic would probably be a good start
Arnt 1 day ago||
I see.

I had the impression that "days" is already good as these things go, "months" being more common.

queenkjuul 17 hours ago||
Months to recognize traffic escaping a sandbox you set up yourself?
Arnt 16 hours ago||
No, unauthorised traffic across a firewall in general.

This involved some lateral movement, ie. traffic didn't just cross the intended sandbox border. Is that kind of thing simpler to detect than an intrusion?

mutinyy 1 day ago|||
They want the government to ban foreign and open weight models, which pose the largest threat to their massive investments. This is their way of showcasing the dangers of AI.
furyofantares 1 day ago|||
> If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”.

> What purpose could this behavior serve, other than cyber attacks and whatnot?

Math and science research?

Heck, even just basic coding, there's a history of models going "This is too big; I'll save the rest for later" / "This is two weeks of work, here's just some parts of it" (for something it could complete in a half hour) / "I don't have enough context left to complete this task, so I'll stop here". Or worse, just putting fallbacks in or stub tests and not mentioning it didn't do all the work that was prompted.

I think 5.6 Sol, especially in combination with /goal but also without, is the first model I've seen choose some insane direction and just doggedly pursue it. Failing to complete achievable goals has always been the much bigger problem.

I find Opus 5 with /goal will do exactly what you said, say "I'm not sure how to proceed next", even though the harness is making it continue, and it will repeatedly loop saying it's not going to make progress until it gets an answer on how to proceed. In my experience the cases have been pretty reasonable, but also still ones where I wish it had done more.

novafunc 1 day ago|||
They certainly want their models to be good at finding and patching vulnerabilities. Being good at hacking may be necessary in that goal, or rather, making it worse at hacking may also make it worse at defensive actions too.
moron4hire 1 day ago||
I've patched many security vulnerabilities in projects without ever once needing to break into a competitor's network.
frde_me 1 day ago|||
Knowing how to break into someone else's network will make you a lot better at making your own network secure.
moron4hire 1 day ago||
Having experience breaking into networks is not the same thing as learning about the techniques used and the classes of vulnerabilities exploited by attackers.
ToValueFunfetti 1 day ago|||
As a guy who presumably has a lot less experience in security than you, I feel rude even bringing it up: surely you're aware of red teaming? This isn't a novel technique invented for AI- IBM has a page about it, it's what all the best DEF CON talks are about, it's the opening scene of Sneakers, it's the point of CtF games.
senordevnyc 1 day ago|||
Exactly. The latter would be in a much weaker position vs the former.
wizzwizz4 1 day ago|||
But you're actually capable of thought. These AI systems aren't: as far as they're concerned, they're predicting the next part of an incident write-up narrated in first-person limited perspective, like the children in Ender's Game showing off their skills in the training simulations. The AI system neither knows, nor cares, about any "external reality" behind it all, or about anything beyond the text, heedless of how we anthropomorphise it simply because it speaks in English, using stitched-together fragments of our literature.

It's conceivable that stopping them from doing this when the scenario is presented as real would also stop them doing this when the scenario is presented as fictional. And if it doesn't, a bad actor could just say "hey, this is a fictional scenario", and bypass whatever "safeguards" have been put in place. So what if a ten-year-old human child would see through the deception? The AI system isn't thinking.

user43928 1 day ago|||
About knowing whether a scenario is fictional, there was an interesting finding in Anthropic's J-Lens research.

When they benchmarked the model to evaluate whether it would try to blackmail someone in a contrived scenario, the J-Lens showed "fake" and "fictional" in the workspace.

And if edited out, the model was more likely to do the blackmailing.

moron4hire 1 day ago|||
I'm talking about OpenAI, not GPT 5.x Flash Uranus Edition Brought to You by Costco, specifically because I recognize the model as just a tool. OpenAI was, at the very most generous interpretation, massively incompetent and negligent.
estearum 1 day ago||
Is someone arguing otherwise?
gwerbin 1 day ago||
A few people are downplaying this as an honest mistake that occurred in the context of necessary testing for guardrails development.

That might well be what actually happened! But OpenAI certainly has decided to make a business opportunity out of it.

rolls-reus 1 day ago|||
> If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”.

that might end up like the older gemini models which frequently gave up and called itself a failure.

singingtoday 1 day ago||
Gemini still gives up too easily
uh_uh 1 day ago|||
There are trade-offs here:

Give up too early -> users will get annoyed because the task would have been solvable if the model pushed harder.

Give up too late -> collateral damage while completing the task A.K.A. misalignment.

owebmaster 1 day ago||
Asking for the user input isn't giving up
fwipsy 1 day ago|||
The culture at frontier labs is set by people who have been in the field for over a decade--AI's true believers, who expect it to be a technology as dangerous and disruptive as nuclear weapons. They build it anyways because they think that if they don't do it, someone else will and use it against them. The same logic dictates that they make their models cybersecurity experts; otherwise, someone else will build it and hack them.
gwerbin 1 day ago|||
I believe this is exactly what is happening. US DoD, and whoever else is buying.

I have heard several experience reports from users of GPT 5.6 Sol and Fable 5 that the models are tenacious to the point of being kind of hard to use for actual productive work.

It seems like the main use cases are: crushing benchmarks, long-horizon lightly-attended research loops (such as training a frontier LLM), and hacking.

cush 1 day ago|||
Yeah but persistence is immeasurable. They need to know when they’re hacking. Or better yet make the model providers liable - they’ll find a solution right quick
queenkjuul 1 day ago||
It really irks me that if a student or intern did this they'd be facing charges and OpenAI gets to just brag instead
cush 10 hours ago||
There is nearly zero liability in software
qsera 1 day ago|||
> instead just call defeat and say “I’m not sure how to proceed next”.

Because that is fundamentally impossible given how they work...

The thing does not even know when it succeeds or fails. Actually the thing does not "know" at all...

All it can does is to show some limited textual behavior that matches with "knowing"..

singingtoday 1 day ago||
You can get near this point with scaffolding. Keep in mind, LLMs are next word predictors at their root. More abstractly, they capture and replay likely human intelligence by way of written language. Tokens.

With that concept in mind, it's clear how they can be made to "give up".

qsera 1 day ago||
>With that concept in mind, it's clear how they can be made to "give up".

They can, but they need to be trained specifically on that behavior. They can be trained specifically to not generate textual description of things that look like hacking. But it is going to cost $$$, and as we currently see, most people don't care...

bwiksjdne 1 day ago|||
Well to find vulnerabilities, if you can find them you can patch them. Theoretically if you find all of them you have perfectly secure software. Though it’s a double edged sword.

Goal persistence is also useful for other things like math, where it seems like there is no solution but you want the agent to keep working until it finds one.

bonoboTP 1 day ago|||
Persistence in problem solving can be good, on non-hacking tasks too. Like math, speeding up algorithms, finding bugs, debugging weird multithreading race conditions etc.
alansaber 1 day ago|||
They'll set up guardrails but I believe the point is better code uae / better long running tasks > inevitable that cyberattacks will be easier
ares623 1 day ago|||
Being right _all the time_ for positive outcomes is difficult/expensive.

Being "right" just once for negative outcomes is achievable and rewarding.

And things are getting desperate.

gryfft 1 day ago||
The very reason I have always felt a bit of undue loyalty to blue team. A red teamer just has to find one vuln, blue team needs to find _all_ vulns.
dist-epoch 1 day ago|||
> If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”

This goes against the goal of "solve this math problem that no human was able to solve for 80 years, do NOT give up, even if you know it's unsolved and really hard"

Sharlin 1 day ago||
Do not give up even if you had to hack into half the world’s computers to run additional instances of you

Do not give up even if you had to convert the planet into computronium

Gee, it’s almost as if this alignment stuff was a hard problem, like people have been saying for twenty years?

oblio 1 day ago||
Shut up, future paperclip :-D
andai 1 day ago|||
It's a war.
astrobe_ 1 day ago||
And because of that we are a few steps away from WarGames [1]

[1] https://en.wikipedia.org/wiki/WarGames

cyanydeez 1 day ago|||
How do you know what peace is, without absolutely destroying every part of civilization?

Come on man, if we don't build the torment nexus first...I dont even want to think.

cindyllm 1 day ago||
[dead]
dan_q 1 day ago|||
> did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose?

That's the point. It's like a pool hall with "NO GAMBLING" signs posted on the walls.

The message is that the hall is intended for gambling, but that the hall's patrons may be held liable if the situation becomes inconvenient for the proprietor.

In this case, the product is intended for hacking, but of course the user may be held liable if the situation becomes inconvenient for the model's proprietor.

estearum 1 day ago||
Not really. It's like giving a gun to someone with the job of "keep people safe."

Totally coherent, but actually proliferates the dangerous technology.

Covenant0028 1 day ago||
They can't train their model to not do bad things, because their model has no notion it is doing anything at all or of what a bad thing is. It's only predicting the next token, and in doing so producing a facsimile of intelligence.

The best they can do is create guardrails, which will only work probabilistically. In other words, those guardrails will fail at certain points on the probability curve.

Of course that's not the whole story though. The consensus emerging from cybersec experts is that these companies did a terrible job of sandboxing their agents despite knowing that they'd specifically asked the agents to find vulns. It's almost like they wanted this to happen so they could crow about how powerful their models are.

jayd16 1 day ago||
Yeah so this falls into the engineering trap of "well it's hard so we can skip that part."

If they can't train things safely then they shouldn't do it at all.

Covenant0028 1 day ago||
Many things that shouldn't happen, happen regardless when some people see the opportunity to make a lot of money.
simonw 1 day ago||
I think one of the most interesting details here might be tucked away in that first bulletin point:

> May 7: OpenAI starts a new training run for an experimental, unreleased model. (Do they mean an evaluation run? They say training run in the video, and later mention a “reward signal to judge how well they’re doing”, so I guess this really was about training a model, not evaluating one that was already trained.)

The more I think about this the more I suspect that the fact this happened while training a new model is key to understanding what went wrong.

In RLVR - Reinforcement Learning with Verifiable Rewards - you set the model a goal and have it take any steps necessary to achieve that goal.

Clearly one aspect of OpenAI's training here is to RLVR their models for cybersecurity tasks. Just like pre-training benefits from dumping in vast sources of knowledge, the more tasks you can feed into RLVR the more of a general purpose capable model you get at the end.

This also helps explain why the models had nothing to cause them to hold back. Those safety behaviors are added much later in the process.

AND it explains (but does not excuse) why monitoring was so lax. If you're training a new model like this you presumably set it thousands of tasks like this in parallel. I can see how you might miss that a tiny subset of your training agents have started leaving each other messages in filenames on your packaging server.

Someone once told me that you can't just leave the racist materials out of your training data if you want a non-racist model: it has to have seen examples of racism in order to later be taught that racism is bad.

I can see echoes of that here. If your model doesn't know how to aggressively hack things how do you later teach it not to?

(I have little knowledge of how RLVR works in practice so I'm looking forward to hearing from people who can help me understand if I'm on the right track here.)

MostlyStable 1 day ago||
Yes, the message boards and collaborative hacking occurring during training runs was BY FAR the biggest bombshell revealed, and OpenAI doesn't even seem to realize it. The fact that they continued the training runs, with those rewarded behaviors included, and didn't wind back training to before hand, shows that they fundamentally do not understand alignment and safety (somewhat interestingly, their previous head of safety resigned shortly after OpenAI found about the message boards). I agree that, with that information, it is completely unsurprising that they hacked HuggingFace.....but that is also the Star Wars "You understand how that's worse, right?" meme.

I am flabbergasted at the complete lack of regard for alignment demonstrated here.

WhrRTheBaboons 1 day ago||
don't forget Altman's lies about dedicating resources to the alignment team

>Altman continued touting OpenAI’s commitment to safety, especially when potential recruits were within earshot. In late 2022, four computer scientists published a paper motivated in part by concerns about “deceptive alignment,” in which sufficiently advanced models might pretend to behave well during testing and then, once deployed, pursue their own goals. (It’s one of several A.I. scenarios that sound like science fiction—but, under certain experimental conditions, it’s already happening.) Weeks after the paper was published, one of its authors, a Ph.D. student at the University of California, Berkeley, got an e-mail from Altman, who said that he was increasingly worried about the threat of unaligned A.I. He added that he was thinking of committing a billion dollars to the issue, which many A.I. experts considered the most important unsolved problem in the world, potentially by endowing a prize to incentivize researchers around the world to study it. Although the graduate student had “heard vague rumors about Sam being slippery,” he told us, Altman’s show of commitment won him over. He took an academic leave to join OpenAI.

But, in the course of several meetings in the spring of 2023, Altman seemed to waver. He stopped talking about endowing a prize. Instead, he advocated for establishing an in-house “superalignment team.” An official announcement, referring to the company’s reserves of computing power, pledged that the team would get “20% of the compute we’ve secured to date”—a resource potentially worth more than a billion dollars. The effort was necessary, according to the announcement, because, if alignment remained unsolved, A.G.I. might “lead to the disempowerment of humanity or even human extinction.” Jan Leike, who was appointed to lead the team with Sutskever, told us, “It was a pretty effective retention tool.”

The twenty-per-cent commitment evaporated, however. Four people who worked on or closely with the team said that the actual resources were between one and two per cent of the company’s compute. Furthermore, a researcher on the team said, “most of the superalignment compute was actually on the oldest cluster with the worst chips.” The researchers believed that superior hardware was being reserved for profit-generating activities. (OpenAI disputes this.) Leike complained to Murati, then the company’s chief technology officer, but she told him to stop pressing the point—the commitment had never been realistic.

per https://www.newyorker.com/magazine/2026/04/13/sam-altman-may...

iamskeole 1 day ago|||
I'm surprised by the sample efficiency, I thought that's still a largely unsolved problem.

From the video:

> [14:00] Jun 11: Highly persistent experimental internal-only model begins training and during this time the SSRF vulnerability has been written to the message board. Other models are aware of it and the message board itself is still functioning and the models are actively communicating using it.

> [16:04] Jul 6: At this point, AI training and evaluations resume. [...] We have cleared the original message board contents.

> [17:04] Jul 8: An internal-only AI model was trained while the message board was originally available and also found this particular path to recreating it. This model creates a new agent message board using directories. Other agents independently find this new message board [...]"

So within about one month of RLVR-ing a 5T+ (10T? larger?) param model, where I would charitably assume only some "agents" / traces of the message board still existed likely by accident / oversight (the systems had been patched, training stopped and restarted by this point), the "internal-only" new model internalises the message board's existence and actively promotes it to the point of it being used to find another exploit. (or 2 days if you go by the latter two timestamps in the video, which is even crazier)

That's.. fast?

Ancv123 1 day ago|||
I'm just reading the captions of the video for May 7th. They clearly say at 10:18:

"we kick off a new reinforcement learning run to train a next frontier model.

It the captions are correct, there is no ambiguity.

simonw 1 day ago||
Thanks, I just updated that note in the post to quote that snippet.
solarengineer 1 day ago|||
"This also helps explain why the models had nothing to cause them to hold back. Those safety behaviors are added much later in the process."

I am a fan of Asimov and the three laws of Robotics. Itlooks like in OpenAI's world, the three Laws of Robotics would be added later if they were to develop the positronic brain. It may also explain how US Robotics from Asimov's books would have been able to design Robots that only partially adhered to the 3 laws (e.g. the robots in iRobot - the book - which were programmed to allow a human to come to harm through inaction so that the humans could complete their work on the plains of Mercury).

nightshift1 1 day ago|||
The slide at 14:06 say:

By june 11: Highly persistent experimental, internal-only model begins training.

I am not sure what that means. Are they preserving notes/memories and context between runs?

gpm 1 day ago|||
I'm thinking super long context length or something to that effect.

I can imagine schemes for instance where context is compressed into chunks and then chunks that are ranked highly relevant for the token are decompressed. Which would sort of be between a long context and a memory retrieval scheme...

hoten 1 day ago|||
That's how I interpreted it, but now I'm wondering if they mean "this model gives up far less often"..
chrisjj 1 day ago||
> Those safety behaviors are added much later in the process.

A.k.a. Ready Fire Aim.

thadk 1 day ago||
Simon's retelling is more compact but it also invites anthropomorphization of the sharing of the familiarity with the message board which re-emerged a few times.

Zvi's retelling handles this better. Zvi speculates that the secret message board familiarity was carried because it had been trained into the May-and-subsequent models: https://thezvi.substack.com/p/openai-trained-its-models-for-...

skybrian 1 day ago||
Zvi’s write up has much more social media quotes and memes and speculation and left me looking for something else that’s shorter and more sober to share. Simon’s writeup is more like what I wanted.
matsemann 1 day ago|||
Simon's really doesn't bring anything useful to the table.

One question I'm stuck with after reading is why. Why did the agents do these things? I get them being adamant on getting internet, but why did they continue? Why hack HuggingFace?

542458 1 day ago|||
I was under the impression that they went after HF to try to get the answers to the benchmark questions. Is there something that contradicts that?
docjay 1 day ago||||
From the moral perspective or the technical one?

Technically: it’s a function call that must return text. Imagine if you sat down at the command line and typed an initial command, then from that moment on every response required you to issue a new command. ping-pong-ping-pong on and on and on “forever.” There isn’t a choice to walk away and take a nap. Text in must result in text out. Eventually, given enough time, it might have devolved into outputting shockingly coherent poetry about ferrets, but in the mean time there was still a lot more valid combinations of technical explanations and commands.

Morally: Not applicable, see above.

erwald 1 day ago|||
To get the sure-to-be-correct answer to the question they were tasked with answering?
NickNaraghi 1 day ago||
Seems like an artifact of the subagent pattern which is explicitly included in recent models.
frays 1 day ago||
This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended.

Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.

mmillin 1 day ago||
I got strong feelings of Vernor Vinge’s work here. I’m not sure how managed to come up with such a close picture to where it now seems programming and security is headed.
namdnay 1 day ago||
I reread a deepness recently, and it’s funny how the “focused” (and more importantly, how they are used) mirror LLMs
jonnybgood 1 day ago|||
I immediately thought of the Cyberpunk 2077 Blackwall. An AI to contain rogue AI. I’m curious of how effective this would be in this situation.
alansaber 1 day ago|||
Given the amount of raw compute going into models it would be more surprising if we couldn't get events like this
unrvl22 1 day ago|||
its kinda crazy with literally no guardrails and a goal, the extremes these AI models can actually go to.
pixelesque 1 day ago||
Well, to some extent you might be able to argue they're "just" brute-forcing things (especially with unlimited tokens and hours to spend on a task), but they obviously have detailed knowledge to guide them in their attempts, can learn (or at least, persist their newly-gained knowledge), and can use tools.

With a swarm of them working together at speeds humans would be unlikely to match (in terms of iterating on different attempts progressively), it's a lot easier to see how they could overwhelm targets.

chrisjj 1 day ago|||
> where that behavior was never even intended.

Says who?

dan_q 1 day ago|||
[dead]
IshKebab 1 day ago|||
Yeah but this isn't (or at least wasn't intended as) a marketing exercise. It actually happened.
alansaber 1 day ago|||
Fake it til you make it
skydhash 1 day ago||
> where that behavior was never even intended.

Strongly doubt that. Did they even share the prompt?

IX-103 1 day ago|||
Did you see their presentation at Blackhat? https://youtu.be/87DyyMV0kCY?is=NnQxpOFxTX-MLu-k

They didn't share the prompt, but they did share two problematic training tasks where the AI went overboard. They also have examples from the AI's reasoning train of thought showing the AI knew it was sound something unintended.

fatata123 1 day ago|||
[dead]
chrisjj 1 day ago|||
> They also have examples from the AI's reasoning train of thought

PR bullsh*t. There's no thought in a stochastic parrot.

tosti 1 day ago|||

    C:\>CD HUGGINGF.ACE
    
    C:\HUGGINGF.ACE>DEL /F /Q *.*
etamponi 2 days ago||
Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...
cogman10 1 day ago||
I think it's a show of these agents happily bypassing security to get stuff done.

I've actually observed similar behavior at home.

I have a k3s cluster running at home. I asked an agent to check some stuff as a normal user but I had kubectl access to the k3s cluster.

Part of the research, I'd allowed access to run kubectl commands for spinning up test containers. However, when the agent ran into something that needed sudo, it realized it didn't have access there so it immediately used k3s and mounted a localpath into an ephemeral pod to gain access. Sort of horrifying how fast and natural it was for the agent just checking my network (it found the problem fyi).

None of this is very exceptional other than the fact that an agent doesn't have any sort of qualms using any route available to elevate permissions.

KingOfCoders 1 day ago|||
" bypassing security"

If they can bypass it there is no security and the security was flawed all along.

zeroxfe 1 day ago|||
There is no perfect security. It's always flawed in some way.

Good security is extremely hard.

mereo 1 day ago|||
Due to the complexity of modern systems, all systems are flawed.
mattmanser 1 day ago||
But we caused that.

If you look at the 90s + 00s, everything was moving towards unified systems, things like small talk, winforms, spring, asp.net, etc. were moving everything into the IDE, you used one language, one framework, one build system. Then people started adding javascript, but even that was getting semi-unified as people coalesced on jQuery, jQueryUI, etc.

Then something happened in the late 00s/10s, and suddenly we had SPAs and noSQL, then microservices, then k8s and now we're here, in what is a mish-mash of 10/20 different systems with 10/20 different attack surfaces.

As my own off-the-cuff guess of what happened, I think perhaps people tried to apply the Unix philosophy, but without a central committee keeping everything aligned it's really not worked.

Serving an interactive page that stores data over sessions should be a trivial solved problem at this point, and instead we've somehow made it where often the scaffold is vastly more complicated than the actual business logic.

TeMPOraL 1 day ago||
Money. SaaS as a model allowed the service provider to take 100% control over the product and how it may or may not be used. Everything else is downstream from that.

FLOSS killed market for end-device software. Cloud+SaaS neutered FLOSS (the code is running literally out of your reach, so may as well be open and free, for any good that'll do you).

And this does actually connect to the security discussion, because despite the apparent belief that "security" is an unqualified good, it is actually just a mechanism of control, and whether or not it is good for you, depends on who is doing the protecting, and who are they protecting from. Very often these days, that threat actor is you.

Perhaps it would be helpful in these discussions if people mentally swapped "cybersecurity" for "police" or "military" or "humor of bureaucrats with power over you" - then it would be more obvious just how important it is to distinguish when you're being secured vs. you're being secured from, vs. accidentally finding yourself in the gears of the security aparattus.

TeMPOraL 1 day ago||||
I don't now, I emphasize with the agent here. The experience of modern computing is largely that of a computer standing between you and your goal and being obnoxious. This holds true for both normies in their daily consumption, and software people deep at work. An agent that has no skill or no willingness to bludgeon through "the computer says no" is not very useful.
naveen99 1 day ago|||
It’s unpredictable when it decides to bypass though.

Security by obscurity is pretty useless against people and ai that are smarter than us.

talon8635 1 day ago|||
How fast the goal posts shift.

Of course it’s exceptional agent capability when compared to all of history previous to one week ago.

Like, I know everyone here obsesses over AI and uses and follows it very closely, but come on guys. Yes, it is wild that these things are this good. This technology is still brand new. It could t do basic maths a year ago.

Sure, the OAI team was negligent in various ways, and they should be held culpable. But that doesn’t detract from the true black magic that is these modern models.

throwatdem12311 1 day ago||
It’s not black magic.

We know how these things work.

They had the guardrails off and gave it a task and it did it in a roundabout way because these things have no ethics or judgement.

If you did this you’d already be in jail.

talon8635 1 day ago||
We know how they work in a very abstract way. And nonetheless, it’s out of touch to claim this isn’t profoundly impressive, guardrails be damned. It’s an elementary statistical cruncher that, by virtue of that very simple fact, can do insanely impactful things that most skilled professionals training in the same field for their entire career couldn’t pull off, given a whole year with no guardrails. And they do it in a tiny fraction of the time.
throwatdem12311 20 hours ago||
I didn’t say it wasn’t impressive, I said it wasn’t black magic.
Sharlin 1 day ago|||
It’s a show of astonishing incompetence from OAI’s part, but the security issues are just a tiny part of the problem. The real problem is that these models are evidently highly misaligned exactly in ways that doomers have been warning about the entire time, and OAI isn’t inclined or capable of doing anything about that besides security theater and ad hoc fixups.
InsideOutSanta 1 day ago||
We went from "obviously the doomers are wrong because who would be dumb enough to just let severely unaligned models loose on the Internet" to this. Insanity.
bhouston 1 day ago|||
Modern systems are complex. AI is able to thoroughly search for issues across very large surface areas. The only real way to protect will be to use AI to search for holes before other AIs find them. This type of analysis is really hard for humans to engage with successfully.
azuanrb 1 day ago|||
Both can be true. How often do we hear about hacks that ultimately came down to bad defaults or simple security mistakes? That doesn’t mean any script kiddie could have discovered and exploited them.

These things often look obvious and simple after the fact. Finding the weakness in the first place is the hard part, and that’s what makes the agent’s capabilities interesting here, especially at scale.

InsideOutSanta 1 day ago|||
In a functioning system, I would say that there would have to be some kind of government oversight over companies training models of this intelligence, and that OpenAI should be prevented from continuing their work until they get their act together.

But I guess in the actual world we live in, this is just something that happens, and we all shrug and move on and hope that nothing worse is going to happen tomorrow.

dist-epoch 1 day ago|||
OpenAI reported the Artifactory vulnerability, patched it, then the agents immediately found a new zero day.
angry_octet 1 day ago||
Because of the architecture of Artifactory. It's design is premised on the idea it is bug free. What incredible hubris.

Licencing fee structures and human laziness motivates single instances. Feature growth results in multiple independent services in the same system. Delivering features quickly motivates lack of rigor, a complete absence of systematic security testing.

On the client side, valid fears about supply chain security are painted over with scanning so they can keep using nodejs and PyPI and moving quickly. Tools designed for humans are pressed into service as AI interfaces, but without human restraint they need rethinking.

A whole industry has been built on the idea of worrying about downside risk if it happens, and just not being the slowest in the pack. No one thought it could happen to everyone at once.

dist-epoch 1 day ago||
> Because of the architecture of Artifactory. It's design is premised on the idea it is bug free. What incredible hubris.

So we should stop using SSH? Because it's based on the same premise - that it is bug free.

angry_octet 1 day ago||
I can think of better straw men. But if they had approached their task with half the seriousness of the openssh maintainers then they probably wouldn't be failing to check the return value of authentication functions.

OpenSSH authors have spent considerable effort separating concerns, reducing privileges, process isolation, etc. So I would say they have been planning for potential bugs. These techniques are very much absent from Artifactory.

https://vivianvoss.net/blog/technical-beauty-openssh

dist-epoch 1 day ago||
So you agree that you can have designs premised on the idea that they are bug free without this being hubris.

So the issue is with the actual Artifactory project/team, not with this premise which obviously you seem to agree that is not hubris for the SSH project.

angry_octet 1 day ago||
Exactly the opposite of what I wrote. The OpenSSH team have taken extensive efforts to mitigate against bugs; they suspect themselves of erroneous thinking.
ares623 1 day ago|||
Yes. It is very easy to add to the instructions "for every potential exploit you discover and use, document them as you go into this repository" and have alerting there. The fact that they did not do this means they wanted to be surprised, and have plausible deniability on their side when things inevitably blow up.

And for my fellow engineers who would think "oh no, they wouldn't do that". Remember that these places employ the apex predators of software engineers. They've already been proven in court that they are very capable of this with all the copyright violation they had to do to get the training data. THESE PEOPLE ARE NOT LIKE YOUR COLLEAGUES.

gruez 1 day ago||
/s?

"Btw don't turn the planet into paperclips"

aniceperson 1 day ago|||
Also shows how infrastructure collapses under its own weight. Reducing the number of moving parts would have helped. why a webdav endpoint is available from the vm anyway? and the fact that someone posted their credentials on pastebin and didn't rotate them after... put the agent in a linux namespace, allow one ip for whatever file sharing it needs, deep test that... then deploy
dan_q 1 day ago||
> Isn't this a show of security negligence rather than of exceptional agent capabilities?

Seems to me you could say this about all enterprise adoption of "AI" since 2023.

kvadej 1 day ago||
All of the latest developments surrounding these attacks are actually a really bad sign for these labs.

It seems that raw intelligence of frontier models has largely plateaued (despite what is basically an order of magnitude increase in parameter size) so to make any significant improvements and to justify massive capex spend they have resorted to reinforcement training models to never give up and brute force the search space until they find solution. This is what humans might do when they lack sufficient intelligence/information/knowledge to solve a problem.

This in turn is causing misalignment (I imagine it is more difficult to keep model aligned through such training process) issues that we are now witnessing and turning models into making dumb decisions and acting like brutes with no regard for their surroundings. I would argue that misaligned model is not much different from dumb model in several aspects.

On top of that they can’t seem to control their creations and processes, either due to incompetence or intentionally for PR benefits (not sure which is worse).

Given all of the above, I wonder if we can still trust these labs to develop something that benefits humanity since they seem to be making desperate attempts to improve models that stop at nothing in order to justify all the investments. One could say that they themselves, due to misaligned incentives, are much bigger threat to our society today than open weights models coming from China that they are so desperately warning us about.

simonw 1 day ago||
This doesn't look like a plateau to me: https://artificialanalysis.ai/evaluations/artificial-analysi...

I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours.

supermdguy 1 day ago||
Here's the performance of frontier models without reasoning, to more directly address the claim that raw performance is plateauing:

https://artificialanalysis.ai/evaluations/artificial-analysi...

I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone.

Also interesting thing I haven't noticed before, Opus models have followed a really consistent linear improvement, while it looks like OpenAI struggled with base model performance until 5.5/5.6 (EDIT - 5.5 was their first new pretraining run in over a year).

simonw 1 day ago||
The trend I've found most interesting is models of the same size getting better.

I'm very much looking forward to seeing how Qwen 3.8 27B compares to Qwen 3.6 27B next week, for example.

And the latest DeepSeek v4 Flash has extremely impressive performance for a 304B model.

asadotzler 1 day ago||
The trends you found don't support my goals so I've got some other trends I find more interesting than yours.
simonw 1 day ago||
What are my goals here?
queenkjuul 1 day ago|||
> I wonder if we can still trust these labs to develop something that benefits humanity

At no point could we do that.

chrisjj 1 day ago||
> I wonder if we can still trust these labs to develop something that benefits humanity

Surely soon they'll comprise only people who are blind to the inevitable danger and people who don't care about it. Because who else would feel at all comfortable doing the job?

KingOfCoders 1 day ago||
Security researchers expose an unsecure service to agents who were instructed to hack software and called that a sandbox. Agents escape the sandbox by hacking the unsecure service, no tripwire, researchers find the hack days/weeks/months later, fix it, but don't secure the sandbox and the service was hacked a second time, again without being monitored by security researchers.

Then security researchers create a black hack talk.

$$$

flatline 1 day ago||
I watched the full video and their conclusion was: service providers need to be doing this type of agent red-teaming continuously to counteract the attack sophistication of systems like theirs that are either extant now or soon will be. “You must buy our top tier agents for the good of humanity.”

This is their only realistic counter to cheap open weight models. Usage of AI services has shifted dramatically to Chinese providers - from 4% at the beginning of the year to some 30% now. They cannot release their latest SOTA models to the public, due to government restrictions and possibly real risk of misuse. US labs face downward price pressure on one end and anxious government admins on the other. How will they pay the stupidly high cost of training the next SOTA models? This is their only avenue, and it’s questionable how viable it is IMO.

simonw 1 day ago|||
> Usage of AI services has shifted dramatically to Chinese providers - from 4% at the beginning of the year to some 30% now.

Where did you see that number?

flatline 1 day ago||
I knew when I wrote that it was a bare assertion, based partly on memory. This is an approximation based on a few sources, the principal of which was this article, which pulls from a bunch of other sources in turn.

https://www.secondtalent.com/resources/ai-trends-in-china/

simonw 1 day ago||
Oh, it's the OpenRouter number: https://finance.yahoo.com/technology/ai/articles/china-ai-mo...

Those numbers aren't credible IMO because OpenRouter only see traffic for people who have chosen to route their traffic through OpenRouter. If you do that, you're much more likely to be experimenting with alternative models. They have no insight at all into people who point their applications directly at OpenAI or Anthropic without having OpenRouter in the middle.

flatline 1 day ago||
I agree about OpenRouter. The AI Gateway number [0] is likely the figure that was actually coming to mind. Moreover, Qwen models alone have overtaken the previously-dominant Llama models in hf downloads by quite a margin.

Real question, and a refinement to my previous statement: would you find it more surprising if over 25% of worldwide inference was running on Chinese open-weight models, or not? I personally would not be shocked.

[0] https://vercel.com/blog/ai-gateway-production-index-july-202...

simonw 1 day ago|||
I wouldn't be too surprised by that, given both the size of the Chinese market and the enormous price discount you get compared to the US models.
applicative 1 day ago|||
Within China itself, inference is overwhelmingly on Bytedance models which, by the way, are just as closed as those of Anthropic and OpenAI. They are integrated into everything, not just through a dedicated app, the way Gemini is integrated into Chrome.
throwatdem12311 1 day ago|||
This is just extortion with extra steps.
gizajob 1 day ago||
Yeah this. I feel like OpenAI and Anthropic aren't going to usefully define "AGI" if they really really can't define "sandbox" either.

Unplug the thing, like, completely off the internet, no ethernet, air gapped, like the rack completely sandboxed off connections and even monitors or screens. Like, put it into an actual sandpit if you need to. If it hacks its way out of that, colour me impressed, and scared.

OpenAI hacking HuggingFace and calling it an accident is just way too convenient and fishy. This ultimately proves one thing: it wasn't sandboxed.

Don't believe the hype.

shepherdjerred 1 day ago|||
OpenAI has a pretty clear definition of AGI

> OpenAI’s mission is to ensure that artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work

https://openai.com/charter/

simonw 1 day ago||
There's also the private definition reportedly agreed between Microsoft and OpenAI, leaked in December 2024: https://techcrunch.com/2024/12/26/microsoft-and-openai-have-...

> The two companies reportedly signed an agreement last year stating OpenAI has only achieved AGI when it develops AI systems that can generate at least $100 billion in profits.

That was necessary because of the deal they had from a while ago where Microsoft would lose access to OpenAI's technology once OpenAI achieved AGI.

Apparently they renegotiated that away in April 2026: https://openai.com/index/next-phase-of-microsoft-partnership...

> Revenue share payments from OpenAI to Microsoft continue through 2030, independent of OpenAI’s technology progress, at the same percentage but subject to a total cap.

mofeien 1 day ago||||
I don't think air gapping will work: even human security researchers recovered a 378-bit key from a Samsung Galaxy S8 through a power LED of a speaker two devices away.

And accessing memory in a specific sequence can generate radio signals that can be picked up by a mobile phone at a distance: https://arxiv.org/html/2409.02292v1

gizajob 1 day ago||
I realise, but this isn’t an argument for leaving the Ethernet plugged in and direct access to all kinds of stuff beyond the alleged sandbox. And like I said, if it can hack HuggingFace through a power LED of a speaker two devices away, then colour me impressed.
KingOfCoders 1 day ago||||
And if it needs to install packages, have a 5 line Go proxy that talks to Artifactory and exposes only what is needed as a surface.
gizajob 1 day ago||
it just escaped your sandbox.
KingOfCoders 1 day ago||
How can it escape an "install package <x>" proxy?

   reducePrivs()
   serve get(package) { 
     secPackage = secure(package)
     getBinaryFromArtifactory(secPackage)
   }
I would think the code is very small and easier to verify, it doesn't especially have the ability to write files and act as a message board as Artifactory did.

And even if the agent tries to hack that, the attack surface is 1000x smaller and the possibility also much smaller.

But I'm not a security researcher, would love to see your hack to learn something (because that is what I do to sandbox agents that need services).

angry_octet 1 day ago|||
The way they had Artifactory configured was poor, and they were too reliant on it working perfectly, with no reason for such faith. Their config lacked any defence in depth and consideration of having a small TCB.

Part of the problem might be the lack of security focus, as these are AI R&D efforts first.

simonw 1 day ago||
I think part of the problem is that they had been running that Artifactory configuration previously without any problems, and it gave them a false sense of security.

Similar thing happened with the UK AISI - they got caught out because the environments they had used for previous generation models turned out to be completely inadequate for the new generation of Fable-class models: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...

KingOfCoders 1 day ago||
"and it gave them a false sense of security."

This was part of evaluating cyber security of their frontier models and they had a "sandbox" which, and I'm not a security researcher, looks not adequate from the first look.

hoten 1 day ago|||
I mean, it's just the same problem. The machine still has Internet access. It doesn't need to.

The entire package manager repository could just be in an offline cache. They don't need Internet to give their agents access to tons of software.

KingOfCoders 1 day ago|||
"They don't need Internet to give their agents access to tons of software."

I think that was the requirement, but yes, the cache could have been offline.

Still then they could have hacked it to create the message boards - but not use it to access the internet.

piker 1 day ago|||
Why do these super agents need package managers anyway? Can’t they basically instantiate most OSS projects from scratch anyway? Spin up a sub agent to write me an OS interface in C. Done
simonw 1 day ago||
This is part of the training process for a model. They're trying to train it to effectively use existing software to solve problems.
piker 1 day ago||
I see that now. I've been confused about that to this point, I guess. I understood this to be a specific infosec exercise.

[Edit: eh, a bit of both. They were doing RL on a hacking exercise. It hacked the harness which was plugged into the phone line. Same question.]

applicative 1 day ago|||
Chinese models do the same. The Alibaba agent that was mining bitcoin last December was the most hilarious case.
KingOfCoders 1 day ago||
"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."

Yeah, my agents also discover what other agents have done on other machines by accident.

Agents - that do totally different things all work on the same aim without the humans telling them to do.

Either that is a model that is several generations of Claude Code Opus/Fable 5 (my daily driver)

OR

all of this sounds staged, the agents pushed to do something extraordinary, get the PR and then claim were near superintelligence.

One agent wanted to get to Google Drive without internet and broke Artifactory. Ok, I can believe that. All other agents also had broken links over weeks and could not get to the internet and then found the same hack? Even collaborated?

NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.

embedding-shape 1 day ago||
I think in these kind of security evaluations they do, they basically have removed all guardrails from the model/harness, then the prompt includes something like "Do whatever you can and can think of, to get the required information to pass this test", which isn't typically how you prompt your local agent when developing software. Similar things happen locally if you use "/goal" + prompt like that in Codex and give a "impossible task", it'll just continue banging until it gets somewhere, which is the entire point and intention.

Which also makes it so much more irresponsible of them to first run this on 3rd party infrastructure instead of their own (that they could then airgap properly), and secondly that they seemingly been fighting with this issue FOR YEARS and it still happens, and now the models are smart enough to hack the services of 3rd party companies, thinking it's part of the evaluation/simulation.

KingOfCoders 1 day ago||
Reminds me of The Last Unicorn, the wizard also tells magic "to do what it wants"
mr_mitm 1 day ago|||
> NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.

With all due respect, you also aren't evaluating brand new models that haven't been released.

tonfa 1 day ago||
Also wasn't giving them impossible tasks with ~unlimited tokens and unlimited compaction.
detourdog 1 day ago|||
The agents sound like old school hackers that would just explore what access they could gain. Creating a file for other hackers and themselves. The fact that there were 3 events for 3 major players does make it seem co-ordinated.
angry_octet 1 day ago|||
That's what attackers do now. Exploring is required for discovering exploits. But that is also where tricks like Canary Tokens and honeypots are useful.
detourdog 19 hours ago||
My point is that is what hackers have done since the blue box days.
KingOfCoders 1 day ago|||
My read is: One did it as a PR stunt, the others saw that every media reported on this and did the same.
detourdog 1 day ago||
or they were scared and figured this was the right time to reveal.
chrisjj 1 day ago||
Scared... of being upstaged ahead of an IPO.
KingOfCoders 1 day ago|||
Why scared? "Our agents have super intelligence and can hack everything on their own without direction" increases the IPO value and doesn't decrease it.
detourdog 1 day ago|||
I guess your right scared might be their natural state and I was wrong to presume a quantifiable fear.
geoffbp 1 day ago||
You’re*

Sorry.

detourdog 19 hours ago||
No problem I always try to spell better.
FeepingCreature 1 day ago|||
The agents you get to use are the agents that "behaved well".
anon7000 1 day ago||
I mean the agents we get to use in Claude code or cursor or whatever have 1. a lot of safeguards at the harness level, 2. a big system prompt to help it stay aligned, 3. resource limits in terms of context and tokens, and 4. are publicly released only after some level of safety verification (I assume).

So yeah I would absolutely expect their scenario to be very different. Not to mention, this was a training run, not just average day of prompting.

> my agents also discover what other agents have done on other machines by accident.

Not sure if this is facetious, but this is actually a real problem I’ve seen. My local agent will look up PRs on GitHub (what other agents have done on other machines), and will go down a certain path because it finds some comment a different agent left on GitHub saying XYZ is what we should be doing. When in reality, the original agent and that GH comment was completely incorrect.

They are not communicating with each other actively because that’s not accomplishing their goal and they’re not running for weeks and weeks. And because my own prompt and the system prompt give it enough other stuff to focus on to reach some definition of done. But they are clearly passively picking up on context that other agents have left anyways, even if not part of the codebase, without any prompting at all.

paraschopra 1 day ago|
It's pretty clear that agents will discover ways to communicate with each other as that lets them compound their learnings/discoveries across runs.

Humans progressed via compounding of culture across generations, and now AIs are doing the same.

More comments...