Top
Best
New

Posted by volf_ 1 day ago

MiMo v2.6(mimo.xiaomi.com)
1092 points | 471 commentspage 4
ddxv 1 day ago|
This looks great in terms of cost and capabilities, truly pushing the frontier forward in terms of open weight light weight models.
informal007 1 day ago||
I'm thinking if this is a more fair comparison among other models, like ChatGPT, Claude and Deepseek
drob518 1 day ago||
Conspicuous that there’s no reference to GLM 5.3/Flash in the reported benchmarks. Just Deepseek and Kimi.
wren6991 21 hours ago||
Granted this is an awesome release and I loved watching the livestreamed RL dashboard, I found this message on the dashboard (https://mimo.xiaomi.com/rl/) quite funny:

> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.

And this in the model card (emphasis mine):

> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.

Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?

coss 1 day ago||
Anyone know what game engine its using to make the 3d game?
DanMcInerney 1 day ago||
This is a big week. Probably getting next OpenAI and Anthro models, Grok 4.7, Mimo, etc. These open source model releases are why I can't take the "slow down" crowd seriously. I pitted older Mimo, qwen, step, gpt-oss, and other models against each other playing games like Werewolf and Sketch.io-like games where I let them talk shit while they played against each other. Mimo was by far pareto frontier of game-playing for the models that were <$0.15/m input tokens on OpenRouter. Qwen was pareto frontier in the shit talking game though. Qwen's hilarious. https://www.tiktok.com/@clankerfights/video/7642862917582425...
Zaraif13 23 hours ago||
> This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback.

Gotta love a capable open model. BUT, how can they just casually throw in that they're actively exploring RSI as if it's just another technique? Is this not alarming at all?

est 23 hours ago||
This model & pricing fits LeiJun's visio for xiaomi: The costco wholesale of tech companies.
algoth1 1 day ago||
Finally a lab that doesn't cheat on the charts
stemlord 1 day ago|
Stupid question: in the benchmark diagrams I'm assuming the values are percentiles, so what does 100% represent?
More comments...