Qwen is just crushing it overall. I regularly use 3.7-flash for everyday coding needs and it gets the job done.
quirino 1 hour ago||
A couple days ago they had published an overall score of 53 for this model, but that was removed and today it returned with a score of 56.
I wasn't able to find an explanation from them. Anyone knows what happened?
Art9681 1 hour ago|
A wire transfer happened.
ignoramous 1 hour ago||
The kind of distillation guaranteed to work.
h14h 1 hour ago||
This has me hopeful for Qwen3.8-27B!
aliljet 1 hour ago||
Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
Alpha3031 1 hour ago||
Depends on what you want to do. Some task specific models can be trained with a few ten or hundred thousand training examples so you can use a bigger model to produce synthetic training examples and then fine tune a smaller student model. I think that's the usual process. Whether you'd get acceptable performance this way depends, as mentioned, on what you're trying to do and what you'd consider acceptable.
teravor 1 hour ago||
once you are able to get the full probability distributions per token you can distill it on specific domains. distilling without that isn't generally a good idea unless you have invested millions in the requisite infrastructure.
esafak 55 minutes ago||
It is also the most expensive open source frontier model, per task; cf. Cost per Intelligence Index Task. If it is as good as the benchmarks indicate -- I'll never know because I see no reason to try it -- it is a positive indicator for Qwen and China.
looksjjhg 1 hour ago||
That took what 2 years? I love how the chip ban made them more efficient
steve-atx-7600 1 hour ago||
curious about methodology. ive seen them post results for claude/codex when they only ran over benchmarks 3 times per model...
atemerev 1 hour ago||
Well, that's the bad index then. It is barely usable in my opinion compared to other Chinese frontier models.
ramon156 7 minutes ago|
which one of the other chinese frontier models is better?
brcmthrowaway 1 hour ago||
Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?
colingauvin 1 hour ago||
DS4 Flash Q2/Q4 mixed quant fits on a DGX Spark (a $4000 device which is not particularly unheard of expense for Apple customers), and is indistinguishable for me from Opus for my personal daily use/assistant benchmarks[0].
Apple is already doing this... they worked with Gemini to distill the model into a smaller one that fits on your phone. If you have iOS 27 Beta, you're already using this
LPisGood 1 hour ago||
Almost surely. Apple is extremely well positioned to take advantage of this over the next decade.