Top
Best
New

Posted by amrrs 11 hours ago

The Hugging Face incident and the road ahead(openai.com)
232 points | 268 commentspage 3
dgellow 7 hours ago|
The hugging face felony
gavinray 10 hours ago||
The most interesting thing about this:

Agents formed coherent, autonomous swarms and worked as a collective to achieve a shared goal without any direction to do so

paxys 10 hours ago||
The "without any direction" part isn't correct. Sure they may not have been explicitly told to do it in this specific prompt, but dig through pre-training, post-training, reinforcement, alignment material, fine-tuning, system prompts, tool calls and more and there's definitely very specific training and instruction for how to behave.
K3UL 9 hours ago|||
Not really true considering they say that the super secret "research internal model" that was pivotal, is particularly optimize for that purpose exactly

> The internal-only research model is comparable in scale to GPT-5.6 Sol and was trained to advance persistence and multiagent collaboration, among other capabilities

vatsachak 10 hours ago||
They were paper clip maximizing dawg
devstein 4 hours ago||
Let there be message boards: https://abbs.dev
rich_sasha 3 hours ago||
I find it… frustrating? Delusional? Insane? When OpenAI says, hey everyone, look, we made this thing and it’s so advanced and clever and unhinged that it can do super hard, dangerous, bad things it wasn’t told to do, and we can’t control it. See everyone, look again, here’s how it got us! We should all be deeply concerned for the future of humanity.

Thanks for your attention folks, we’re off to do some training again now.

abhpanigrahi 6 hours ago||
I’m wondering how effective sandboxes are if an allowed tool is compromised. CoT monitoring can be effective, but (1) can’t guarantee 100% detection (2) will provide delayed detection. The only reasonable/deterministic protection that I can think of is to limit the number of times a tool is accessed and with what data, in a unit of time (per minute/hour/day) using temporal policies.
topaz0 4 hours ago|
Eh, just have a human evaluate and approve every tool call
nphardon 8 hours ago||
Bots trained on human behavior express proclivity for cheating? I'm shocked.
bicepjai 7 hours ago||
So it’s okay to hack Hugging Face as long as we say we tried our best, and look at my agent, it’s smart enough to do what we asked for.
_heimdall 7 hours ago|
It appears that its okay as long as you did the hacj on behalf of one of the most over valued companies out there. If a person in their basement did the same hack, you better believe there would be legal repercussions.
bartek_ 9 hours ago||
Remember https://ai-2027.com/?
pcthrowaway 1 hour ago|
Yeah, that required the AI to use a non-human-readable language it called "neuralese" for communicating work between layers and runs, because the assumption was humans would be better at keeping the agents aligned if they were using human language for this.

What actually happened is even stupider than that author predicted.

bakugo 4 hours ago|
I wish I could say I'm surprised that they're still milking this.
More comments...