Top
Best
New

Posted by Wirbelwind 12 hours ago

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs(scalex.dev)
232 points | 182 commentspage 5
oblio 8 hours ago|
We already have the solution. Use AI to validate AI agent commands.
threethirtytwo 9 hours ago||
The future of software is fixing bugs and security issues in production.

Many companies will be accepting this new paradigm because of raw speed. Something that could take say 4 years to fully mature will now take less than a year. But the cost is that many of these issues will have to be caught during live QA either in production or investing heavily in QA. That’s the future.

Oras 9 hours ago||
So humans scored 66% on human eval?
rvz 9 hours ago||
Proof that people just do not read what they are seeing on their screens when put too much trust in the agent as it prints the result and they will approve anything on their machine.

So if a basic curl | bash was tweaked to download malware which the agent gets tricked into running the command but it said it was safe, the user would just approve it.

unclebucknasty 9 hours ago||
Interesting premise, but there's not much real world meaning here without stats on the percentage of agent-offered commands that are actually dangerous.

If that number is something like 10%, then we have a really big problem. But, if it's .000001%, then it's pretty vanishing. At some point in between we cross a threshold that puts the risk below many other risks that we routinely take (e.g. trusting npm dependency graphs).

Of course, if it's really that low a percentage, then the entire model of "supervising" via human approval really is fundamentally flawed.

Damjanski 9 hours ago||
love this so much!
azhdanova 5 hours ago||
[flagged]
msbel5 4 hours ago||
[flagged]
fenestella 6 hours ago||
[flagged]
unjuno 7 hours ago|
[flagged]
More comments...