Top
Best
New

Posted by tosh 14 hours ago

Small Models Have Arrived(calv.info)
559 points | 250 commentspage 2
dgunay 7 hours ago|
Luna max is suitable for like 90% of the kinds of code changes I want to make. I only find myself actually reaching for a Sol or Fable tier model if the problem is very complex. If you're willing to build the guardrails and do some extra planning, Luna is very capable.
glimshe 13 hours ago||
> There's obviously a lot we can optimize here, but if you're charging what the WSJ or The Economist charges, you'd better be delivering similar value.

Gosh, watching paint dry has been a better value than reading The Economist in the last 5 years or so.

That aside, I had good results with Luna. I'd be interested in hearing about a comparison that takes into consideration response time (not TPS), cost and performance of the popular models at different settings. That chart has some of that. For instance, is Luna Max a better value than Terra Medium?

yousif_123123 13 hours ago|
[flagged]
my-next-account 11 hours ago|||
Are ya kidding me
butterisgood 2 hours ago||
I don’t care about the cost if the results are weak. I did a bunch of work with Sol only to find Opus 5 spot a bunch of bugs.

And it was correct. The results weren’t as good as they should have been for Sol. How am I going to trust Luna?

bartleeanderson 8 hours ago||
Nobody is questioning what is meant by small? When you started mentioning frontier models that don't run locally, I just go TB;DR "Too big, didn't read"
daquisu 2 hours ago|
I read the full post. It is more about cheap models with good enough performance rather than small models. "Local" was not mentioned at all
a13n 6 hours ago||
Regarding the Pareto frontier and related benchmarks, I have a hard time taking anything seriously that claims that Opus is anywhere near the intelligence of Fable. Are there any benchmarks that haven't just been benchmaxxed that more accurately represent actual usage?
kakugawa 6 hours ago|
FrontierCode is prob the closest. [1] It's closed source (so no direct benchmaxxing), and it was calibrated by 20+ open source maintainers. It shows Opus 5 (medium), beating out the other reasoning levels by a large margin. i.e. Opus 5 w/ higher reasoning levels actually reduces performance. [2]

However, you'll have to gauge for yourself how closely their tasks resemble your tasks.

1/ https://cognition.com/blog/frontier-code

2/ https://cognition.com/frontiercode

yipinwong 12 hours ago||
"Small models" nowadays work like someone who has IQ 100+ while SOTA ones are like 150, "relatively".

Given sheer number of turns I can make with small models, I can do a lotta stufff

- cheaper, and faster

Harness makes differences: There have been many HN posts about how one made tiny models work better at certain tasks using harnesses.

These "small" models with right context, and guidance, they work wonders.

---

I've been saying Luna has been my go-to AI in previous comments and why Luna is still more compelling than GLM-5.3-flash.

- https://news.ycombinator.com/item?id=49450353#49452248

Kim_Bruning 2 hours ago||
It's kinda fun that people get to experience what the 8-bit era was like!
pranav_tech26 11 hours ago||
Running small models locally beats wrestling with API latencies and rate limits. The compute trade-off is 100% worth the privacy and DX gains.
kokessch 1 hour ago||
you both deserve a hot place in the HELL
anuptalwalkar 8 hours ago|
I kind of agree with your assessment. Running models not just on local, but cheap and lightweight frameworks will drive the next phase.

Not trying to plug, but I do't know any other way. I wrote a piece couple of days ago on small models and memory usage on the edge devices- https://polign.com/blog-edge-agent-memory and https://news.ycombinator.com/item?id=49450816 closing on the same problem.

More comments...