Top
Best
New

Posted by volf_ 1 day ago

MiMo v2.6(mimo.xiaomi.com)
1091 points | 470 commentspage 3
jjcm 1 day ago|
Here's an image->html test for it using 2.6 Pro Ultraspeed, along with comparisons for grok 4.7 and Astra.

Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

MiMo 2.6 Pro Ultraspeed (36min): https://html.non.io/annui-mimo/

Grok 4.7 (25min): https://html.non.io/Annui-grok/

Astra (19min): https://html.non.io/annui/

Overall this felt like the weakest of the three. Ultraspeed was fast as far as tokens per second goes, but it overthought quite a lot of things resulting it in having one of the longest build times. That overthinking didn't lead to better results either - note the statue with the cropped off arm. It's also the worst implementation of the dynamic lighting effect / displacement effect - the background especially has some significant distortion. Astra was the only one that seemed to understand that displacement should happen less the further something is in the distance.

Here's a vid of all 3 side by side with the source design: https://non.io/video/annui-comparison.mp4

faitswulff 1 day ago||
Astra's looks the worst to me on mobile, though
jjcm 1 day ago||
That's fair - worth noting that none of them were instructed to make a mobile variant or to test the mobile size.
Imustaskforhelp 1 day ago|||
For me it seems like this model’s overthinking should be tamed.

Perhaps instead of giving it a one shot task with vague prompt. I wonder how it can perform with much more detailed and constrained prompt (can you please elaborate more on the level of detailness and ambiguity that the prompt is and where does this model seem to overthink the most?)

Also are there any ways to tame such overthinking of models in general?

I hope that once models start becoming smart enough (I think for me it’s already there) or becoming genuinely the Sota. They then start focusing a lot more on optimizing token usage

Saline9515 16 hours ago||
I used 2.6 Pro and it seems to overthink way too much, leading indeed to very slow build time.
Alien1Being 17 hours ago||
Training cost $ 3.47 million....

Staggeringly low for a frontier model.

leothetechguy 16 hours ago||
This figure doesn't include pretraining cost. Which is probably higher.
tw1984 16 hours ago||
that is just for the RL
dom96 1 day ago||
Very capable model. I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

KillSwitch-Bench 1.0

  Claude Opus 5           66.9
  GPT-6 Astra             57.9
  Claude Fable 5.1        46.7
  MiMo-V2.6-Pro           38.8
  Muse Spark 1.3          36.5
1 - https://bench.killswitch-lang.org/
bel8 19 hours ago|
I appreciate that this benchmark is different but it is in no way how most people use LLMs or promote equal grounds when benchmarking:

- capped per-task budget and time limit

- No internet access

- different harnesses mixed

dom96 17 hours ago|||
How do you think other benchmarks work? Every single one is going to have a budget and time limit. There has to be a cap on those, you can't just let it spin forever and use unlimited funds.

I would argue that this benchmark is uniquely suited to how most people use LLMs because it actually tests common harnesses and it is a true coding benchmark for a language that is unseen, thus testing the LLMs actual ability to understand nuance and learn.

Internet access is restricted to ensure that over time models cannot just look up the source code of KillSwitch, which would allow them to cheat.

bel8 11 hours ago||
I didn't expect there to be a time limit and I do know benchmarks that don't have a time limit. They just penalize slower models which I think is fair.

In my experience models just don't take forever to mark tasks as done.

For my coding usage I don't set time limit so benchmarks that do so are providing me less interesting use cases.

With that said, I can understand having budget limits for expensive LLMs for those of us that don't have infinite VC money.

As for no internet access, the issue is that for most problems we do want Llms to be able to search docs, API specs, GitHub issues, etc. So not allowing that usually just favours larger LLMs that were able to memorize more data, not necessarily smarter ones when both have internet access.

yt1998 12 hours ago|||
[dead]
geokon 18 hours ago||
is there a good metric of model degredation over time?

Im a bit lazy and only use the free models different companies host and the biggest difference i see is that some models (Gemini, OpenAI) get progressively stupid in long chats. You end up having to start a new session every oncr in a while. Or they get really hung up on a theme and cant shift to a new topic.

By contrast, Ive been impressed with Qwen. I have some chats on research and code architecture that have stretched for weeks without any noteable change in quality (though occassionally it seems to "rush" to an answer)

Im just looking at all the listed benchmarks and im unsue which i should be looking at

lwansbrough 1 day ago||
Anyone else more excited about Chinese models than American models these days? Big thing for me is affordability.
tacomagick 1 day ago||
Absolutely! Chinese models are both cheaper and more capable in many cases, compared to the American models and their makers continuously fumbling or reducing model capability with each update. Deepseek decreased costs when they released Flash 4.1 you would not see any American company do this, in reverse they would try charge you more.
user43928 1 day ago|||
OpenAI decreased prices with the 5.6 model family.

And later they further cut Sol and Terra pricing by 20% (maybe only in the API) and Luna by 80%.

In fact Luna still outperformed DeepSeek Flash 4.1 in cost per task on Artificial Analysis when I last checked.

However, Luna is slightly less intelligent. I have a feeling that it's pretty dumb and prone to hallucination unless running at xhigh or max effort, where it somehow manages to work quite well.

I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.

The competition is great, and I hope Chinese models will continue to force leading US labs to offer models at a low price point.

That said, I don't think the Chinese labs have anything over OpenAI and Anthropic when it comes to capability or efficiency - I have no reason not to believe the US labs have even lower cost to serve the models.

Implicated 1 day ago|||
> I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.

So you don't have much perspective on things, it seems. Let me introduce you to the GLM 5.2 and then 5.3/5.3 flash series of... "oh, wow, I should have bought some RTX PRO 6000's while they were 'cheap'" stage of progression.

As someone carrying multiple max subscriptions to both claude and codex - primary workhorse is glm 5.3 flash running on rented GPUs for less than a latte/hr.

I also found qwen 3.6 27B nearly useless for my own needs. DS4 flash 0731 and then 4.1 have been nearly as eye opening as glm 5.3 flash, but have their own warts.

tacomagick 1 day ago|||
OpenAI had to cut costs because of Anthropic. I also do not trust the benchmarks when it comes to models anymore. I have tried both Claude and OpenAI models and while it is true that the 5.6 series is smarter than Deepseek (at the time i tested it against 4.0) at that price it is still not worth it and sometimes randomly refuses to do tasks or stops midway etc.

Do also remember China is this far in the AI race despite all chip restrictions from America. If they were in equal standards I truly think Chinese models would have long surpassed American ones. Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling meanwhile their own models claimed to be Qwen¹ and their stance against open models is negative² and they still keep blaming China for it.

1- https://news.ycombinator.com/item?id=48671252

2-https://www.anthropic.com/news/position-open-weights-models

goosejuice 1 day ago|||
> Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling

Why wouldn't he? If there really was 25,000 accounts breaking ToS any CEO would at minimum be upset. Evidence of Claude distilling qwen would be damning but that a) makes no sense b) doesn't exist afaik.

user43928 1 day ago|||
Not sure about that.

Given the difference in compute, it seems plausible.

However, the researchers at the US labs are surely no less talented, and they have better access to hire talent globally.

They too have to serve their models efficiently at a large scale, and with current capacity constraints this must be a top priority.

goosejuice 1 day ago|||
> Deepseek decreased costs when they released Flash 4.1 you would not see any American company do this, in reverse they would try charge you more.

OpenAI reduced prices and Anthropic increased weekly usage limits.

bellowsgulch 1 day ago|||
[delayed]
SyneRyder 1 day ago|||
Yep, I'm trending in that direction, and I'm someone with Claude stickers all over my laptop. My main app dev work is still going to Claude, but everything else is going to China even at API rates now.

One simple task: I needed an LLM to go through and clean up a few thousand page descriptions and titles in my personal search engine index, where the human web page authors had put in no effort sigh. I did a shoot out between Claude, Luna, GLM 5.3 Flash and Deepseek. Despite the high cost, Claude's descriptions were terrible, and even Opus warned me that the descriptions coming back from Haiku were "generalized, not accurate". I expected I would choose Luna because of price, and occasionally it did have wonderful descriptions (one captured emotion in a way no other model did). But in the end, the GLM 5.3 Flash descriptions were the easiest to read, they flow well while also being accurate & including necessary keywords, and being highly affordable. So it won out. It's a task that is nowhere near frontier, but a task where somehow China is better than frontier.

rapind 1 day ago||
API rates still aren’t quite competitive with the OpenAI x20 accounts, but they are definitely getting close with deepseek 4.1 flash. I spent a few days with only 4.1 and was very impressed.
verdverm 1 day ago|||
I have a contrarian opinion that China passing America in Ai is the Sputnik moment we need to leave the hubris behind and get our mojo back

debatable if a turn around is possible before '29

swingandamiss 1 day ago||
No, because I'd rather not support our economic and military rivals.
lwansbrough 1 day ago|||
I'm Canadian so this sentiment has little value in 2026 unfortunately.
zemvpferreira 1 day ago|||
As much as the US has been easy to hate lately, I don't hesitate to say Xi Jinping as the most powerful man on Earth would be much, much worse.
ActionHank 1 day ago||||
Also, frankly, as a fellow Canadian it's pretty clear that the biggest "rival" the US has right now is itself. Just passed out in the corner puking on itself shouting about all the foreigners who won't talk to it.
tancop 1 day ago||||
I'm from Europe and I hate America way more than China now. Used to be about equal but then Trump started extorting Ukraine, threatening their own allies and sending billions to Israel to help with a genocide. I think that exposed America for what it really is.
scottyah 1 day ago||||
[flagged]
lwansbrough 1 day ago|||
Because at present the pedophile US president is making it his mission to molest my country. China, for all its faults (including espionage, which the US is also guilty of) is mostly focused on conducting trade.
verdverm 1 day ago|||
Half of Canada now uses the word 'enemy' when asked for an adjective to describe America or China. We're equivalent in their eyes now because we elected Trump a second time and all that he has said and done in 2.0
cwillu 1 day ago||
It's closer to a cousin you used to be close with despite some moral failings, but who has now has a substance abuse problem and is lashing out at family and friends.

Not an enemy, just a danger.

verdverm 1 day ago||
I'm relaying a poll of Canadians, their word choice, not mine

"plurality" would have been accurate over "half" on my part

https://www.commondreams.org/news/canadians-us-enemy-poll

rayiner 1 day ago|||
Canadians warming up to China makes me think of Germany becoming increasingly reliant on Russia in the 2010s.
rapind 1 day ago|||
Murica just has a MAGA problem. We can still be friends if and when you sort that out. Us Canadians like most of you quite a lot.
cgio 1 day ago||||
Yes, someone can still blow up a pipe and they look the other way. On the other hand, you can also draw parallels to themselves becoming increasingly reliant on US vs UK in the past.
Freedom2 1 day ago|||
Agreed, and also because I support freedom of speech!
girvo 1 day ago||
Neither the US nor the Chinese companies are on your side then. They both censor, just different topics.

But at least I can run Chinese models locally, and strip a lot of that censorship/refusal.

Havoc 1 day ago||
3.5 mil cost for a opus level model? Even if excluding salaries that seems very cheap
redox99 20 hours ago|
That's the RL.
wmedrano 1 day ago||
Any idea on how they get the pricing so efficient? Their artificialanalysis graph has them on par with GLM 5.3 but at less than 10% of the price despite being larger than GLM 5.3.
system2 23 hours ago|
It sounds like they are pricing very well because OpenAI and Anthropic have been scamming us for years. Once the infrastructure is in place, electricity cost is the only concern. China supports businesses and gives them a lot of incentives to lower their costs. And who knows what's being provided to them without anyone knowing. All in the name of winning the race.
edg5000 14 hours ago||
If that is true, why are the open model inference providers on OpenRouter (and on their own website) so expensive then? It must be the hardware. I've checked and a lot of it it just the HBM, with the NVidia tax being a smaller but also large factor.
system2 7 hours ago||
Nvidia's top AI chip, Rubin, sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi's MiMo V2.6 Pro could generate roughly 150–300B tokens a day, worth about $130–260k at Xiaomi's API price. That's a payback of a few weeks in theory.

So yes, we are getting scammed by American SOTA.

MisterMunchkin 1 day ago||
I really liked MiMo 2.5, it was really affordable and actually had vision, unlike DeepSeek. (DeepSeek has only recently added it)

Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.

perrygeo 1 day ago||
Can we afford to look past it? If/when claudeslop starts infecting every new model to such an extent, that model will produce its own slop, infecting new models... At what point do we lose all reliable methods for establishing "truth"? This is epistemic collapse waiting to happen. I honestly thought it would take longer... holding out for a coherent shared reality in 2030 seems optimistic.
omani 1 day ago||
how do you recognize "claudeslop"?
Bluestein 1 day ago||
It's an honest, load-bearing, simple thing.-
SSLy 1 day ago|||
that's belt and suspenders too
pimeys 1 day ago||
a smoking gun
Bluestein 17 hours ago||
... and a caveat worth flagging.-
nullc 19 hours ago|||
You're right to push back. That's on me. There is one thing I must flag which will move the needle. My honest take: It's all about the shape of your priors. That's the lever, and that's not nothing. It's worth your attention before you land your next remark. Here's why that matters: It re-contextualizes everything.
eriquesito 1 day ago|
Funny that all but one video has audio, the house 3D model one, where you can hear (what I assume are) Xiaomi's engineers talking about who knows what.
More comments...