Posted by specked-citrus 1 day ago
How is this any different and why would it need a different solution?
Solution is jail, not for the AI, but for the human.
How long till we get some fun trusting-trust attacks on internal OpenAI infra?
At a minimum I would expect an FBI investigation, but given that the US government is right now a failed state I’m assuming no such investigation will happen.
Cybercrime legislation in other countries might not require intent, and then I would hope to see some prosecutions. OpenAI has clearly been negligent, and this negligence is causing harm in the world. Someone should be fined or jailed for this.
Until then, I do wish that both the sorcerer's apprentice LLMs and the orgs failing at securing their data (remember, data is a liability) would face damning consequences.
One is allowed to dream on a Saturday morning.
If LLMs can do this with everything stacked against them, imagine what the NSA has Astra doing right now.
The logic is good, the execution is disgustingly noisy.
honest question, but almost literally everyone doing anything with web technology knows this is simply not true, right?
there's no such thing as "read only Internet" and restricting an agent to GET-requests only to accomplish that, is akin to using base64 for "encrypting" your password
The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?
Give me a break. What a bunch of amateurs.
They "monitor" this by having classifiers watching the model output that'd stop the session/punt you to a weaker model/raise an alarm if they see anything suspicious. They can't do that in a cybersec eval because the normal safeguards would just be going off at all times.
Why didn't they attach a special classifier, which'd allow hacking-within-the-task but not going off the rails? Good question; part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.
And maybe "they are running after glory, not safety"?