Top
Best
New

Posted by dbreunig 2 hours ago

Fable and the End of the Free Lunch(www.dbreunig.com)
71 points | 58 comments
nchmy 1 hour ago|
The real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc...

I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis

geniium 33 minutes ago||
I was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years.

It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it

r_lee 24 minutes ago||
imo this is the problem some of these labs are gonna face, because open models will do this just fine and you as the consumer don't need to pay their training costs

especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA

ksh09 22 minutes ago|||
I'd be content if I could get the DS4 flash, luna, mimo level intelligence running on MY low-end hardware completely offline and bearable TPS, not otherwise.
matteoraso 41 minutes ago|||
>I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster.

There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.

lilbigdoot 1 hour ago|||
If they could be cheap+fast and not try to do too much, that's a good spot for me. I don't use the smarter models as much because of cost and because they're still not good enough to let loose on a lot of problems. For assistance I prefer something that can very quickly spit out a specific piece I can review on the spot and keep going. I let smarter models handle things that I treat as external dependencies and don't care how they're written, but in my core domain I'm still mostly hand coding
nchmy 28 minutes ago||
I have a similar process - its just a pair programmer most of the time. I dont understand how people can have a fleet of agents working a bunch of waterfall specs..
poincareball 25 minutes ago||
Evidence actually supports that capabilities are leveling off, and cheaper/faster is not really coming. Just log-linearly more capability at smaller parameter counts as they saturate.
bad_haircut72 10 minutes ago||
not an AI researcher - this is probably true for these "everything" LLMs but I think specialized models are gonna be the next big thing
rmast 12 minutes ago||
Most of the things I work on are at least security adjacent. At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.

Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.

nicoburns 8 minutes ago|
That's completely valid. But worth noting that most of the stuff I work on is not security adjacent (mostly UI / layout / rendering related), and I almost never run into this.
pigpop 42 minutes ago||
Reading this as someone who switched over to ChatGPT after (and largely because of the changes made in) the Fable release, it reads a bit naive. Not only do I find Sol to be as good, if not better than, Fable it is also faster, better behaved and has a much more coherent writing style. You also don't randomly get the Opus downgrade. OpenAI seems to be pulling this off due to their partnership with Cerebras so I wouldn't make any comparisons to Moore's law just yet considering it seems like we're just getting started in that department. Anthropic could (and should) do the same thing. It certainly feels like model development is at a point where it would be worthwhile building special purpose silicon for the models we have now since they are capable enough that they would still be useful even when/if further advancements are made. If anything, I think Anthropic's problem has more to do with their micromanagement of what users can do with their models, they're creating an undue amount of overhead for themselves by over-policing usage and capabilities.
r_lee 23 minutes ago|
Etched is doing this. it seems like in the near future they'll actually ramp up production. not sure how much faster/economical compared to Cerebras but..
aabhay 31 minutes ago||
This concept of a free lunch was never true. In a competitive dynamic, speed and performance were always worth optimizing, comparing, and improving.

One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.

mholm 1 hour ago||
As models train up the intelligence ladder, many common tasks will hit fully diminished returns, and instead it'll just get progressively cheaper to do that task. But the tasks that AI is capable of doing are also expanding. I'm not sure 'Some tasks don't require the peak of the frontier' is worth worrying about, from an AI finance perspective.
tyre 1 hour ago||
Yes. I use Opus for tasks that Sonnet could probably handle, but I'm not hitting my quota. Whatever minor incremental gain is "worth it", since marginal cost is zero.

Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.

Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.

jml78 55 minutes ago||
I operate mostly in the devops arena. Lots of things opus is fine for. But there is just things where I can hand hold Opus through changes, or I can ask Fable to do it and it gets it right on the first try. People will say let fable plan and validate with opus doing the work. I found that burns fable tokens even faster because opus makes so many mistakes, fable has to review things 4-5 times before opus gets it right. A single fable implementation at medium or low effort would have one shot it.
a2ff6eeb0 56 minutes ago|||
Exactly; so far, we've only replaced the need to design algorithms and hand-write code; what if we apply the same effort towards the skill needed for system architecture, project management, and the rest of the SDLC? Or even outside of software!

Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.

vineyardmike 41 minutes ago||
> Or even outside of software!

As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.

Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).

a2ff6eeb0 21 minutes ago||
I suspect the focus will probably shift once software engineering is no longer the biggest cost center for most AI company's clients, and we'll start working on getting rid of the next cost center.
dgellow 59 minutes ago||
It’s worth considering for companies paying API prices, and not relying on a subscription quota
zkmon 48 minutes ago||
I guess Moore's law analogy is weak. CPU speed has hit a limit in that case. What has hit a limit in AI case? Newer versions of the models are still flowing with more and more capability.

For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.

blfr 58 minutes ago||
What are all these rote coding tasks people do that they can farm it out to lesser models?
nicoburns 7 minutes ago||
One task I've found this useful for is writing example code. Release admin (updating version numbers, etc) as well.
denverllc 53 minutes ago|||
Write a detailed plan using a more expensive model and implement it using the cheaper one.
blfr 34 minutes ago||
How much are you saving once the more expensive model already has all the context loaded and ready to go?
csullivannet 26 minutes ago||
API calls get more expensive, not less, as you've loaded more context. This is exactly when you want to switch to cheaper models.
mattmanser 23 minutes ago||
Are you genuinely asking?

As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.

When you add a new module or whatever most of the code you have to write is rote code.

And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.

I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.

freepiai 50 minutes ago||
I've been offering Deepseek V4 Flash for free in www.freepi.ai and I've started using it as my main driver as well.

Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.

So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.

m3kw9 49 minutes ago|
looks like you haven't tried openai or Sol, or even luna (max)
Zylokloto 31 minutes ago||
He started with thinking were to send what.

I throw everything at claude Opus.

While some people start thinking like OP, A LOT of people just start exploring ai.

And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.

gpjanik 33 minutes ago|
"When Moore’s Law slowed in the mid-2000s" it did not, in fact, slow down in the mid 2000s, or at all.

https://ourworldindata.org/data-insights/moores-law-has-accu...

jbstack 30 minutes ago||
You've selectively quoted the article. The full quote (emphasis added):

"When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc."

Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.

sscaryterry 31 minutes ago||
It did in terms of the traditional more MHz (GHz) is better, but as you've correctly pointed out, not when it comes to actual compute.
More comments...