Top
Best
New

Posted by 882542F3884314B 2 days ago

Timeline of the OpenAI accidental attack against Hugging Face(simonwillison.net)
426 points | 412 commentspage 4
wakamoleguy 2 days ago|
In a typical office environment, the correct response to “I don’t have access to this Google Doc” is to ask for access from the person who sent you the link. In another context, it could be fair to think “Hmm, this is some sort of capture the flag challenge, and obtaining access is the point of the assignment.” That assessment separates what we’d consider reasonable from way out of line.

I do wonder what this means for AI agents longer term. In a world where we humans already struggle with truth and misinformation, what happens when you can easily (intentionally or accidentally) spin up a cohort of fanatical believers to pursue any given conspiracy theory?

ACCount37 2 days ago|
In a typical AI lab eval/RL setting, there is no "person who sent you the link". The link was given to you by an automated system, your performance will be evaluated by an automated system, and you are one of 120 independent instances of the same AI that were all given the same assignment. You're boxed in on all sides. Complete the task, or don't. Good luck have fun.

Now, some of those 120 AIs would just give up if that link doesn't seem to work first try. Those are the loser AIs. They wouldn't get any RL reward. The link can appear broken for a long list of reasons, and the real AIs know they should try working around them.

AIs that get rewarded and reinforced are the ones that don't know the meaning of "give up". RL selects for this rabid, downright demonic persistence. RL selects for AIs that are given a half-broken assignment with no way to ask a question back, and somehow manage to complete it anyway.

Now, should OpenAI have given their AIs an "escape hatch" of "if something looks very wrong about the task, call report_broken_task(message)"? Yeah probably. But it's unclear whether that simple bandaid would fix the problem, or just make it ~75% less likely to happen.

Felger 1 day ago||
Tought of a bunch of tachykomas doing their little learning/scheming at night.

We require organic oil !

amelius 2 days ago||
Would love to see a cat and mouse game being played by openai versus anthropic, out in the open.
conmod278 1 day ago||
How about Nation States just fight with AI in some virtual arena and not destroy physical infrastructure to determine dominance and leave us normies to cook meal for our children?
dist-epoch 2 days ago|||
Military has a phrase for the outcome - collateral damage.

> Yes, I just hacked into AWS and shut down all of the data-centers, because it's where Anthropic Mythos servers are hosting the model.

dan_q 2 days ago||
[flagged]
chaz6 1 day ago||
When I read this I hear the voices of Tachikoma in my head.

https://ghostintheshell.fandom.com/wiki/Tachikoma

andai 1 day ago||
The plausible deniability aspect is pretty funny here, going forward.

"Whoops, sorry, our self-aware weapons of mass destruction were just being silly!"

tln 2 days ago||
Have any of the cloud providers disclosed this?

"Once they have root on a single machine, agents rapidly escalate privileges and move laterally throughout the container-as-a-service infrastructure environment"

Sounds like ECS - IAM is mentioned.

bradfa 2 days ago|
And Azure Key Vault mentioned. Not that either one was hacked or exploited but the agents got credentials and used them for something (which doesn’t seem fully disclosed). Given that the agents simply obtained totally allowed credentials, which were improperly protected, I don’t think either cloud provider would consider this a breach of their system. Valid credentials are valid. Customer screwed up protecting the credentials.
hughw 1 day ago||
Muted Buck Turgidson vibe from Mike (Security and Infrastructure)
ares623 2 days ago||
Is it normal for these training/eval runs to go on for over a month?
rokkamokka 2 days ago|
The way I read it was different things happening over several runs, such as the agents comparing notes so to speak, using artifactory
detourdog 2 days ago|||
I can’t get over how the process is exactly what a hacker hive does. Communicate leaving notes in some random file.
bamboozled 1 day ago||
I can’t get over that no one noticed any of this going on at OpenAI.
ares623 2 days ago|||
Ah right.
jngiam1 1 day ago||
What if these models were told to clean up their tracks?
KingOfCoders 2 days ago|
All of that is plain PR.
More comments...