Top
Best
New

Posted by tosh 17 hours ago

Small Models Have Arrived(calv.info)
608 points | 279 commentspage 5
Zigurd 13 hours ago|
I recently had some relevant experience: for a couple of months now I've been experimenting with on device models to summarize feeds in a Bluesky client I am developing. The feature extracts topic areas, categorizes posts, and creates a summary under each topic.

At first the results were hot garbage, and progress was slow. I hooked up the settings to download models from Hugging Face conveniently, so I could run experiments faster, and I massaged the prompts a bit. Last week this feature made a qualitative jump from science experiment to something I'd actually use.

The fact that all runs on the device means I've got no variable costs associated with adding this to what will be, at best, a pretty low revenue product. I've tested it on trailing edge devices like an M1 Mac and a Pixel 8, and performance is very tolerable.

The key is I'm not asking for open ended answers to open ended problems. When it proves to be useful it's not going to get less useful or more expensive.

There are vast domains of uses for LLM models with similar characteristics and likely similar results.

dzonga 15 hours ago||
small models + a good application layer - are more than enough, good for routine business tasks.

the application Layer i.e having a good graph RAG & connecting it up together is the missing piece for most.

sroerick 14 hours ago|
Can you elaborate on this?
lantry 13 hours ago||
The model doesn't have to be smart if all it's doing is pushing a few different buttons.

I don't have to be an automotive engineer to start my car and put it in drive.

sroerick 4 hours ago||
I mean - how specifically are you using knowledge graphs?
hartator 16 hours ago||
I have trouble seeing the points of using less capable models.

I just want the smartest, best, and most capable models. It feels smaller models for speed and cost are just transitions towards better hardware allowing the very best model.

krisoft 15 hours ago||
And that is why i always carry my groceries with an Antonov An-225 Mriya. Is it really needed? No, but i refuse to compromise on what is(was/will be) the best.
imp0cat 3 hours ago||
That plane was destroyed by Russia, wasn't it? I believe it was partially disassembled when Russia invaded Ukraine and so it wasn't possible to save it. :(
arjie 14 hours ago|||
My experience has been that responsiveness is value. For tasks where you need steering, responsiveness allows for better steering. For tasks which you want unattended, better models are just better.

There are still tasks that even Fable is bad at doing. And many are just mundane things. Because of the fact that you have to steer it on those tasks, you might as well steer an 80% model that is 5x faster. And those do exist.

Naturally there’s a bit of a gap because the faster models need steering on tasks the slower models don’t so there’s no smooth transition but I find it worth it. Especially if you want to stay in flow.

Ironically this sometimes means starting a plan with a great model, planning with a worse model, iterating, then submitting it to a better model for review, and then having the better model do the implementation.

trvz 15 hours ago|||
First, smaller models are fun for hackers: you can run them locally, or run them faster.

Second, when cloud models become unavailable or otherwise deteriorate, these will be all you have. May as well prepare.

breezybottom 14 hours ago||
If you're hacking a US-based entity, using a high-performance Chinese model through a VPN is probably safe enough. I doubt a local model is going to be sufficiently smart to hack any major company.
trvz 14 hours ago||
You misunderstood what I was referring to by “hacker” there.
ebiester 15 hours ago|||
It depends on what you're trying to do. For non-coding tasks luna is quite often enough. Flash models are more than enough for summarizing a text, for example, or whipping up a small script to save me fifteen minutes. If you're on a 200/month plan, I see your point. If you're on a dollar limit - or worse, paying per token out of your pocket - you look to be more efficient.
polotics 15 hours ago|||
Can you define your use of the word 'smartest' here just in case some of us don't quite know what you mean?
shafyy 15 hours ago|||
Some reasons: - Smaller models will always be cheaper - Smaller models will always use less energy, therefore better for the environment

It's a bit like saying you always want the fastest and best car; Sure, you can have it if you keep paying for it. But a small car will also get you from A to B, will use less gas and will be much cheaper.

tartuffe78 15 hours ago|||
Cost is the point
0xbadcafebee 14 hours ago||
There's a difference between want and need. I want a 650hp V8 supercar. I need a 150hp I4 toyota corolla. Why choose a less capable car? Because I don't want to spend 10x as much money to get groceries.
agcat 16 hours ago||
I like the analogy on ways to make small model useful.
kokessch 4 hours ago||
you both deserve a hot place in the HELL
nocodeexportcom 5 hours ago||
Nice!!
dev_awesome 7 hours ago||
how effective are the small model?
senectus1 6 hours ago||
not sure how this is really about model size.

I'm keenly interested in seeing super small (like 10's or 100's of mb) special purpose models starting to come into their own.

that might actually be transformational globally not just in rich countries.

kokessch 4 hours ago|
you both derve a special place in HELL
More comments...