Qwen 3.8 Flash Next vs MiMo 2.6 Flash
Sadly for a 2x DGX Spark setup MiMo isnt the next great model
October 07, 2026
MiMo 2.6 Pro and Flash made a lot of buzz a couple weeks ago as a new benchmark darling. Flash was particularly interesting because it joined that new family of models that run pretty well on a pair of DGX Sparks. I decided to give it a run through the benchmarks we put Qwen through while performance tuning a couple weeks back.
I'll share the repo here, as I continue to test more model releases that come out.
The tl;dr, MiMo is considerably slower even with DFlash that appears to be the faster way to run it currently
| ctx | task | Qwen C tok/s (acc.len) | MiMo tok/s (acc.len) | Qwen / MiMo |
|---|---|---|---|---|
| 1K | code | 66.3 (2.55) | 37.2 (3.50) | 1.78× |
| 1K | code_ts | 68.0 (2.62) | 40.7 (3.93) | 1.67× |
| 1K | edit | 70.1 (2.73) | 52.2 (5.00) | 1.34× |
| 128K | code | 65.2 (2.56) | 24.6 (3.45) | 2.65× |
| 128K | code_ts | 65.1 (2.55) | 29.7 (4.17) | 2.19× |
| 128K | edit | 67.6 (2.72) | 30.4 (4.30) | 2.22× |
MiMo is slower per step, and gets slower as context increases. Qwen stays flat. Those speeds are a bit too slow for my coding agent usage for this class of model, but maybe worth it if it performs significantly "smarter". Unfortunately not:
| Test | n | Qwen C | MiMo | Qwen-only / MiMo-only passes | McNemar p |
|---|---|---|---|---|---|
| HumanEval+ (base) | 164 | 95.1 | 90.9 | ||
| HumanEval+ (plus), greedy, thinking off | 164 | 93.3 | 88.4 | 11 / 3 | 0.057 |
| MBPP+ (base) | 378 | 93.1 | 88.9 | ||
| MBPP+ (plus), greedy, thinking off | 378 | 79.9 | 74.3 | 31 / 10 | 0.001 |
| GSM8K, greedy, thinking off | 1319 | 96.4 | 96.4 | 21 / 22 | 1.0 |
| HumanEval+ (plus), thinking, 16K budget | 164 | 92.7 | 89.0 | 11 / 5 | 0.21 |
| LiveCodeBench (80 recent problems), thinking, 16K budget | 80 | 41.2 | 46.2 | 3 / 7 | 0.34 |
| Long context: retrieve / two-hop / reverse at 64K + 128K | 144 | 95.1% | 90.3% | two-hop 41/48 vs 34/48 | — |
Oh well. Still exciting to see another model in this range out there. Between Qwen, Deepseek, GLM, and MiMo there is a ton of awesome development happening. Eagerly awaiting what comes next :)