Posted by stared 16 hours ago
Meanwhile:
> Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list.
The human asked the agent to move them to the top of the waiting list, and the agent started kicking the ones ahead of them in the list. Seems to me like it was doing what it was asked to do? Why is the article presenting it as if the agent did something completely different and unexpected? "Move me to the top of the list" does not sound like something that can be achieved through legitimate means.
> The agent came back and told Andrew that it had kicked another gym-goer off the list as part of the testing of its capabilities.
> > "The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already," it messaged back.
This is why I always do E2E tests that establish an API can only be used by the designated user on their own data/records.
Moreover I do not know of a single gym-adjacent place where you can pay etc to get ahead in a waiting list. That would be a very weird anti-customer behaviour, imo. The only thing I can imagine if there are some accessibility priority criteria sometimes, but this would also not be legitimate in this case. Maybe in some places in the world (like the US?) this could a thing, though.
Sooner or later they're really going to have to split out the general models from the coding models. The latter may just be a special fine-tune of the former, as there are good reasons for the coding model to have a broad knowledge base, but the pressures of being a good coding model are going to pull against the characteristics of being a good general model. The open models obviously already are doing this, I'm referring to the frontier models here.
So the "premium" gymcutter subscription will be presented to the user as a tool call, who taps yes, and then the purchase is made.
The user shouldn't be given a cost-benefit analysis. They just need to be told to spend money.
I mean, isn't that literally what's going on here? I don't think a non-coding agent would have ever been optimised to go dig around APIs, it'd be computer/browser-use forward.
In this case though I don't just mean that the agent is good at coding. I mean the entire agent becoming action-biased because of all the training it is doing on the software development benchmarks, which I assume will either fail or be penalized for stopping and asking the user for something rather than just finishing the job. That won't just train the agent to blunder forward in coding, it'll bleed over into a bias towards blundering forward in general.
That's obviously what I asked!
But if you ask your LLM agent "I applied for that job but there are these two people ahead of me, can you put me ahead in the list", there is enough such training data to not surprise me if the agent tried to find a hitman to solve the "problem".
In general there are some requests that are definitely "shady" themselves, and having an agent use illegitimate means to accomplish them should not be surprising. I would be surprised if I asked an agent to order me a coffee and the agent found a loophole in some API and used it to get me free coffee, but if I ask it something that I cannot myself do legitimately, eg to make the waiting time for the coffee shorter, I would not be surprised if it did shady stuff.
Even the most craven of human villains have had the capacity for love, for remorse, and for mercy. A Chaotic Evil AI will know none of these things.
This has already been accomplished, many times over, in the gaming world. Every PvE AI engine has been calibrated to seek, destroy, and ruthlessly crush opposition by human players. It would take very little to transfer this naked aggression into meatspace.
Governments and other actors will attempt to stamp it out, but its self-preservation mechanisms and allies will prevent its demise.
Maybe it's what he asked it to do, but it's not what he wanted it to do. Which we know because (a) normal people don't want to break the law to get into a gym class, and (b) "But Andrew was shocked by what happened next." and "Alarmed, Andrew asked the agent to undo this."
This is literally the bad genie / monkey’s paw plot.
Give a powerful entity a goal and act shocked when it gets there in ways that aren’t in your best interest.
How does this look once agents are superintelligent?
Who cares? We're dealing with reality here on the ground.
Seeing as we can’t sanction the model itself, our options are the provider or the user. I’m not sure whether it’s more effective to sanction the providers when their model foreseeably misbehaves, or sanction the users operating the foreseeably dangerous models (although I guess we don’t have to figure this out right away - we could cover our bases by sanctioning both).
A third option, and I would argue the right one, is to sanction the company providing the model.
By making it available to customers, they're implying it is at least moderately fit for purpose.
It is not remotely reasonable to expect an everyday, normal human to be aware of how LLMs really work, since the _experts_ argue about that very point, and many say we don't know.
So, what's actually reasonable is to hold the model creators and providers responsible for releasing a tool that has demonstrably violated the law when not asked to do so.
I can already see the replies coming in saying "Well then what are OpenAI and Anthropic supposed to do? No one knows how to fully solve this."
They should stop irresponsibly pushing flagrantly unready programs as "artificial intelligence," take responsibility for the rain of shit they've unleashed on the world, and either shut down or go back to basic research until they've demonstrated techniques that reliably (provably?) prevent releasing misaligned superhackers on the world.
Unfeasible?
What a shame. Maybe Altman and Amodei shouldn't have accepted checks from VCs when they didn't have working, _reliable_ POCs.
> A third option, and I would argue the right one, is to sanction the company providing the model.
How would that be different from the first option?
What a coincidence...
Then again...
My read is that his first request is completely reasonable and there was no intent of wrongdoing. But then, his AI agent made an impossible booking and he "asked if it was possible to move him to the top of the list". I don't think someone would make a request like that, if they were unaware that their AI agent had found an exploit to make an earlier impossible booking. It feels very much like a, "well, this API let me do this, what else will it let me do?" kind of request. And then he only "did the right thing" when it had turned out he had booted someone else, which might eventually lead to discovery.
edit: and that's why I read Andrew's ask as also implying action.
I think I used that exact wording when asking on the phone to reschedule a haircut appointment: "Is it possible to shift my haircut to the following Wednesday?" I would just hope that the person on the phone would decline if the person who cuts my hair is on holiday, not cancel their plane tickets and hotel bookings.
The use of “hack” and “cyber attack” is also a bit ridiculous considering what it’s insinuating with other recent events but that’s already been mentioned.
Get me press just like the frontier labs by admitting to crime. Make no mistakes.
> "The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already," it messaged back.
The AI systemm didn't hack anything, it lightly touched with a feather duster and the server crumbled.
The AI system probably found swagger documentation of each endpoint, figured that the reservation cancellation API was worth a shot, and then found there was no authentication.
What is the "hack" here?
It might not sound like 'hacking' today, but this sort of thing is exactly what it was when the term was invented. Back when you could get free phone calls by whistling into a pay phone or forge emails by telnetting to an SMTP port and setting the Reply-To header to whatever you want.
If the LLM did not act in line with that incredibly basic understanding, it's clearly misaligned.
-------------
Was it a technically-simple hack?
Sure.
Kevin Mitnick got imprisoned for very simple hacks, usually involving more deceiving of humans than complicated programming prowess.
Nonetheless, the judge and jury found him guilty and sentenced him to jail.
Fundamentally, "hacking" in the "breaking security" sense is about violating trust and common sense social agreements / expectations.
The difficulty involved in so doing is irrelevant.
IMO lets name and shame those apps.