Top
Best
New

Posted by volf_ 1 day ago

MiMo v2.6(mimo.xiaomi.com)
1094 points | 471 commentspage 5
toolswatcher 12 hours ago|
I don't overly trust benchmarks. Benchmarks can reflect a model's capabilities, but many models specifically adapt—or even overfit—to score high on benchmarks. DeepSeek V4.0 Pro is a good example: it achieved decent scores, but performed poorly on actual cybersecurity-related tasks, even worse than DeepSeek V4.0 Flash.
puszczyk 18 hours ago||
Very informative page vs the recent grok release
heyjstn 23 hours ago||
How could the flash model beat the pro on the Cyber benchmark?
T1ny 22 hours ago|
As far as I can tell, if you look at the open training dashboard it's because the flash was trained on some cyber while the pro wasn't. 4% of training on the flash was cyber and 0% of the training on the pro was cyber.
alfalfasprout 1 day ago||
The moat for OAI and anthropic seems to be very quickly shrinking. Chinese labs are now using RSI-like approaches and even without resorting to heavy distillation they're catching up in a couple of months vs. what would have been 6-12 months a year prior.

And as these models get better the pace of training is quickly speeding up too.

This doesn't bode particularly well for anthropic/OAI after they go public.

mdale 11 hours ago||
I don't know if we will look at OpenAI and Anthropic as moating on frontier models & selling tokens.

They are banking on the application layer and accumulated business and end user context. They have to quickly make that systems integrated value out weigh the model choice price value in the broader market

Unknown if they will be able to pull that off.

verdverm 1 day ago||
token vendors are headed to the same place mobile data vendors went, this is good for everyone but those who thought they could maintain exorbitant prices
pvab3 1 day ago||
Explain about mobile data vendors?
verdverm 1 day ago||
when mobile data first arrived, it was expensive and people bragged about their bills (token counts today)

with time, it became commoditized, people now have unlimited plans, and the money is made by the applications that sit on top (token generation is increasingly undifferentiated low-level infra)

This is not to say there has not been significant innovation in the time since, but it's a low margin business (tokens look to be headed this way)

nivance 23 hours ago||
it sounds greate. I still have 50% of my quota this month, so I'll continue renewing to give it a try.
sinan-faizal 13 hours ago||
mimo is imporving as well, great work!
Imanari 17 hours ago||
For simple edits it produces huge reasoning traces. Kind of disappointing. Super long repetitive reasoning always sits wrong with me. It feels like a way for the labs to brute-force higher benchmark scores but not actually like a smarter model. Disappointing, Mimo2.5 was such a nice model.
system2 1 day ago||
As I read this I am getting API overloaded errors from Claude. I will switch to something else soon. I hate Claude.
bertili 1 day ago||
They mixed up DeepSeek 4.1 Flash with something else on this page, possibly DeepSeek 4.1 Flash means Gemini 3.8 Flash.
esafak 1 day ago|
It tops the intelligence vs cost Pareto frontier and, uniquely for a Chinese model, does well in response time too.

https://artificialanalysis.ai/models/mimo-v2-6-pro#intellige...

That's pretty fast; I think I'll try it: https://openrouter.ai/xiaomi/mimo-v2.6-flash

One concern I have is that they allegedly do not discount cached tokens: https://www.reddit.com/r/opencodeCLI/comments/1t37dz3/xiaomi...

Can anyone comment?

More comments...