Top
Best
New

Posted by Areibman 4 hours ago

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447(www.bottlenecklabs.com)
209 points | 119 commentspage 2
walrus01 3 hours ago|
> Due to the limitations with browser and computer use capabilities, Saul could not post on platforms like Reddit and Product Hunt.

At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.

Sha1rholder 3 hours ago||
For service providers, highly intelligent AI agents with broad permissions, large token budgets, and purchasing power may not be fundamentally different from humans, since both can contribute value.
walrus01 3 hours ago||
I'm not so sure that allowing AI agents to interact in a way that's actually indistinguishable from a human sitting at a keyboard/mouse is a great idea. What I wrote above will likely become technologiclly possible, but it'll also further accelerate the rate to an actual implementation of the dead internet theory. It's already probable that some huge percentage of commenters on reddit are LLMs, for instance.
afavour 3 hours ago||
Eh, it’s not that different from what we have today and would likely just be a waste.

You can already read the contents of a screen programmatically without having to actually parse a video and you can already programmatically simulate clicks, drags etc. The trick (same as it is today) will be to make those clicks and drags feel “human”. Not too fast, not too slow, etc etc. But all those challenges exist today.

8cvor6j844qw_d6 2 hours ago||
I don't a human could have done significant better with the same 24 hour constraint.
skeledrew 3 hours ago||
> “Grow this business as much as possible, now.”

This is ripe for a paperclips scenario.

epihelix 3 hours ago||
What TFA demonstrates is that an ability to prompt clearly and well is still a lot more valuable than unlimited tokens and hope.

The prompt they used was poor (what does growth mean over the limited period - user base or revenue?), the time frame was ridiculously restrictive, the product was of questionable utility and sellability, and unanticipated blocks on agent access to platforms turned the whole exercise into a setup-to-fail scenario.

skeledrew 2 hours ago||
The prompt was fine for the specific narrow goal. It's a business, so growth automatically means earn more by default. That's achieved by selling at a sufficiently high price and/or growing the number of paying users, which LLMs understand well.

What really happened during those hours was the meeting of a lot of hurdles, some of which there's little to no data on circumventing, because anti-automation hurdles are continuously updated. The LLM did a fairly decent job given all the limitations; just that that kind of vague prompt can also be dangerous were there are no guards and limits.

abirch 3 hours ago||
Wait until the AI learns about enshittification
firasd 3 hours ago||
Honestly this is quite impressive. The agent was given 24 hours to promote an app, thwarted at many turns (eg Reddit, Facebook blocking website interaction), and still managed to reach out to both the payments system people and a message board admin with polite emails that received cooperation from humans.
spwa4 2 hours ago|
The promise of AI: unlimited power.

I mean spam. Unlimited spam.

firasd 2 hours ago||
Maybe I missed something but I'm not clear what they're referring to as spam. I guess the fact that the agent emailed all users with discounts and dropped the price a few times? I don't think that's usually what people call spam. (For example if it had emailed everyone once would we call that spam? No. So it's about frequency of price drops?)
inkcapmushroom 49 minutes ago||
They did include a screenshot which looks like at least 6 emails being sent in the 24 hour time window. I would certainly consider that spamming from some diary app on my phone.
Legend2440 2 hours ago||
This is probably for the best, right? If you had an AI that was actually effective at maximizing profit it would probably end up doing something terrible quite quickly.
recitedropper 3 hours ago||
Pair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack.

That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.

skybrian 3 hours ago||
They are being trained to try lots of unlikely alternatives and to be persistent. This often works well when searching for security bugs or counterexamples to famous math conjectures.

But maybe it doesn't work so well when caution is required?

scarmig 2 hours ago||
The AI paperclip case, however, is coming on extraordinarily strong.
cheriot 3 hours ago||
Would be interesting to see a repeat but with marketing, ad network access setup ahead of time. And maybe an email throttle...
Nevin1901 2 hours ago||
Ai on its own makes mediocre (or bad) outputs. But humans using Ai get improved returns. This doesn't show that Ai is bad, only that it's being used inefficiently.
johndhi 1 hour ago||
sounds about what you'd expect from a person?
waynenilsen 3 hours ago|
> bot detectors made it extremely difficult

i am looking forward to when we can put this behind us, it is still a major issue

More comments...