Top
Best
New

Posted by krackers 13 hours ago

Xiaomi Mimo 2.6 live post-training dashboard(mimo.xiaomi.com)
424 points | 109 comments
joelwallis 12 hours ago|
I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.

The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.

-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

ehsankia 14 minutes ago||
> late last year/early this year

That's an eternity when it comes to coding models.

In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.

miyuru 3 hours ago|||
Same here. It’s the first AI provider I actually gave money to, since they offered the model for free with a Mimo code for the first month or so, and it was great.

These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.

alwinaugustin 9 hours ago|||
I am also using 2.5 and it is giving me solid results. Its available free on Openrouter
rapind 7 hours ago|||
I’ve been very pleased with DS 4.1 flash. Not so much the 4.0 models, but for coding (Rust) it’s been great so far (3 solid days of work).

I’ll give Mimo a try.

trollbridge 7 hours ago||
MiMo is my backup whenever DeepSeek is down, had the price bump, is slow, etc.

UltraSpeed was absolutely awesome. I miss it.

DS 4.1 Flash is amazing. Well worth the extra cost.

baxtr 2 hours ago|||
Could you elaborate on how you check daily? Do you swap models for certain tasks?
walrus01 12 hours ago|||
I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.
girvo 6 hours ago|||
The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.
jonsoft 4 hours ago|||
I made this 3D game in a day on the same setup with Qwen Code as agent: https://games.jonathanpage.com/

And I am not a web developer! It's an extraordinary model.

(Mouse and keyboard required)

walrus01 6 hours ago|||
Yeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it.
girvo 5 hours ago||
Yep, the engrams are on NVMe (the speed penalty was lower than I expected) and it is quantised to fit.

It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it

jwpapi 11 hours ago|||
May I ask why you ended up there instead of just using the heavy subsidized subscription. I’m actually curious.
eli 9 hours ago||
Mimo has subsidized subscriptions too
flexagoon 9 hours ago|||
How does it compare with DS 4.1 Flash in your experience, if you ignore the cost?
james2doyle 12 hours ago|||
2.5 Pro or the regular 2.5?

I always found that those Mimo models to be really good at tool calling and following instructions

ignoramous 2 hours ago|||
> GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.

miroljub 2 hours ago||
I wouldn't call that inexpensive.

For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.

esafak 11 hours ago|||
How fast is it compared with the other Chinese models?
ricardobeat 10 hours ago||
They both are in the 50-100 tok/s range. The Mimo v2.5 Pro Ultraspeed beta could reach 1000 tok/s, hoping they can do something similar for the new model, it was amazing.
wangxili1997 1 hour ago|||
[flagged]
electroglyph 9 hours ago|||
[flagged]
NuclearPM 9 hours ago||
Real?
electroglyph 6 hours ago||
mimo 2.5 has been a big underperformer since shortly after it's release imo. i cancelled my sub after the first month. purposefully using 2.5 right now is just handicapping yourself for no reason.
NuclearPM 5 hours ago||
I understand now. You used the wrong word.
yeeeloit 11 hours ago||
[flagged]
senordevnyc 11 hours ago|||
Yeah, this Brazilian dude who has been a contributor here on HN longer than your anonymous account is shilling for a Chinese model company. Makes sense.
platinumrad 11 hours ago|||
Are you accusing them of astroturfing? Why is it strange for someone to say something topical?
dr_dshiv 10 hours ago||
Well, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.
dzonga 9 hours ago||
the open burial started when zAI served their latest model on all Chinese chips.

now we r just noticing the grave getting dug deeper.

skybrian 8 hours ago|||
For my own usage, Luna is cheap enough that I don't care if other models are cheaper. I'm interested if another model is in some way better and not too expensive.
rapind 7 hours ago|||
Luna is great but makes a lot of mistakes at high and lower in my experience (large rust codebase). I use Luna Max for asynchronous subagent reviews and am very happy with its work, but it’s slow af.
ijidak 6 hours ago|||
What plan are you on?

Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.

teki_one 6 hours ago||
Sol is useless atm on the Plus plan, 1-2 questions 5-10m to get through the 5h allowance. (used to be good, can change any day)
SlightlyLeftPad 8 hours ago||
I think it might be closer to this:

https://www.debtdefaultclock.us/

passive 10 hours ago||
Neat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far.

I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.

The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.

ricardobeat 10 hours ago||
For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.

Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).

https://deepswe.datacurve.ai/blog/deepswe-v1-1

markasoftware 20 minutes ago||
gemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure
Cookingboy 8 hours ago||
2.6-pro just reached 63.7% by step 10, it's on step 11 right now.

Even flash reached 60.7% by step 12, and it's on step 16 now.

This is so exciting lmao.

krm01 12 hours ago||
This is pretty neat. What would be a good reason for the other Model providers to not do this?
kibae 12 hours ago||
Speculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed.

Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.

Bolwin 3 hours ago|||
I don't really remember a situation, which of those models supposedly beat the other?

I still opus 4.6 though not for code

jwpapi 11 hours ago|||
I think first of all it’s not an obvious idea, also the marketing surplus for other providers is not as big for openai/anthropic as for xiaomi and last but not least I’m pretty sure you can withdraw methodology from here.

I’m saying who has a million dollars for me, so I can make my own model?

nikcub 5 hours ago||
this is remarkable transparency in an otherwise hyper competitive and secretive industry
liuliu 12 hours ago||
When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
jampekka 12 hours ago||
Kinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data.

I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.

https://en.wikipedia.org/wiki/Training,_validation,_and_test...

liuliu 12 hours ago||
Correct. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.
lucrbvi 12 hours ago|||
They are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.
nodja 10 hours ago|||
They exist to detect degradation. Datasets are not perfect and if a batch contains too much bad data it can ruin a run, also an opportunity to find bad data and improve the dataset filtering.
SwellJoe 12 hours ago|||
You gotta have something to aim at. And, presumably, the benchmark is not part of the training data, it is the test against which the model is tested at each stage; is behavior moving in the right direction?
esafak 11 hours ago||
Not if you don't train against them.
kingstnap 10 hours ago||
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.

It's not the direct feedback loop of RL but its not far.

fzysingularity 11 hours ago||
Very cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.
ProfessorLayton 12 hours ago||
2.6 Pro: >started 2026-09-15 10:32 UTC

For some reason I thought training took much, much longer than what the progress bar suggests.

This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.

GaggiX 12 hours ago|
These are post-training reinforcement learning steps.
krackers 12 hours ago||
Yes, updated the submission title to say "post-training" to hopefully prevent further confusion
ssn2000 4 hours ago||
Total run cost is $1.2M until now, what resources are they using to train their model? Wish they shared more details on that and what the MFU metrics are.
ttul 7 hours ago|
$5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.
stymaar 10 minutes ago|
Which isn't that much when you compare to the kind of DC that US actors are using.
More comments...