SØNDAG
2026-10-04

Too many projects, too many ideas, too few hours — one learning a day anyway

MiMo-V2.6-Pro: the cache-hit price is the real headline

The benchmark rank gets the headline. The line I care about is the cache-hit price: $0.0036 per million input tokens. Agent loops resend the same long context constantly, so that’s where a coding agent’s bill actually lives. At $0.13 per index task, it’s worth a side-by-side run against what I use now.

Self-hosting is another story. Pro is 1.02 trillion parameters, Flash 309 billion. Neither is going on a box under my desk, MIT license or not. What I’d actually pull is MiMo-V2.6-Distill-Qwen-9B, plus the released RL code and training environments. Flash I’d hold: Xiaomi says it’s near Pro on agent benchmarks, but Artificial Analysis hasn’t evaluated it yet.


The story — Xiaomi released MiMo-V2.6, led by MiMo-V2.6-Pro, which now tops open-weight models on Artificial Analysis’s Intelligence Index with 46 points, ahead of GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash. API pricing is $0.435 per million input tokens and $0.87 per million output tokens. Pro and Flash are mixture-of-experts models handling text, images, video and audio with up to one million tokens of context. Xiaomi credits heavily scaled reinforcement learning for the gains. All three models’ weights are on Hugging Face under MIT. (Source)