Top
Best
New

Posted by Areibman 4 hours ago

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447(www.bottlenecklabs.com)
233 points | 134 commentspage 3
Nevin1901 2 hours ago|
Ai on its own makes mediocre (or bad) outputs. But humans using Ai get improved returns. This doesn't show that Ai is bad, only that it's being used inefficiently.
waynenilsen 4 hours ago||
> bot detectors made it extremely difficult

i am looking forward to when we can put this behind us, it is still a major issue

dylan604 4 hours ago||
"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?"

"It Lied, Spammed, and Lost $447."

Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.

gtowey 4 hours ago||
Right, and currently we are limited by how many teams of people can get together to run campaigns like this.

Now imagine that LLM agents make this possible for nearly anyone. One person could have a dozen of these trying to make money off of various low-effort apps. Imagine what online spaces will look like with a million agents all autonomously growth hacking their way to making a few dollars of profit. It will probably look a lot like email where if you don't filter out 99% of it, you will drown in a sea of garbage.

dylan604 3 hours ago||
> if you don't filter out 99% of it, you will drown in a sea of garbage.

Sounds like the app stores

qznc 4 hours ago|||
Maybe they should have given it a billion dollars and the strategy would have worked fine?
freeone3000 4 hours ago||
Given a billion dollars, it would have likely ended up with a million-dollar company
onraglanroad 3 hours ago||
Not $447 million? Sounds like a result!
leros 3 hours ago||
I think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, repeat.
codedokode 3 hours ago||
Turing test passed, acts indistinguishable from a human, although the scale of loss is not human-like yet.
verdverm 2 hours ago|
Turing was testing our gullability, v2 is a preference test
kritr 4 hours ago||
I’ve found that when the right cli tools are preprovided / provisioned for the LLMs to get the job done, they tend to do okay.

But when hunting for them in the wild, they get a lot more confused.

cynicalsecurity 1 hour ago||
Shitty instructions = shitty outcome. Blame yourself, not the AI model.
paxys 3 hours ago||
Sounds like it is as intelligent as the average startup founder.
Razengan 3 hours ago|
So, just like humans?
More comments...