· via dev.to (home feed)
Xiaomi's 1.02T MiMo-V2.6 Pro and DeepSeek V4.1 Flash make MoE the open-weight frontier
Two MIT-licensed open-weight releases arrived twelve days apart: DeepSeek's 552B V4.1 Flash and Xiaomi's 1.02T MiMo-V2.6 Pro, each computing with a sliver of its parameters per token.

Two MIT-licensed releases, twelve days apart
September brought two open-weight Mixture-of-Experts (MoE) releases that bracket the current state of the technique, according to a dev.to analysis published October 8. DeepSeek V4.1 Flash arrived September 10 with 552 billion total parameters, of which roughly 8 billion activate per input token and 16 billion per output token. Xiaomi's MiMo-V2.6 Pro followed on September 22 with 1.02 trillion parameters and just 42 billion active per token, organized as 384 experts under top-8 routing. Both carry MIT licenses, and both undercut far smaller dense models on price: the post lists V4.1 Flash at $0.15/$0.60 per million tokens off-peak and MiMo-V2.6 Pro at $0.435/$0.87.
The spread between capacity and compute is the story. V4.1 Flash stores 552B parameters of knowledge but computes like a roughly 16B model on output tokens, a 34x gap — about 3% of the network participates in any given step. MiMo-V2.6 Pro's spread is roughly 24x.
How the sparse trick works
In a dense transformer, every parameter is evaluated for every token. MoE replaces each block's feed-forward network with a set of smaller expert networks plus a router. The router scores every expert against the current token using a softmax, keeps the top k (typically 2 to 8), and merges the selected experts' outputs as a weighted sum. Modern variants add a shared expert that every token passes through regardless of routing — DeepSeek's design pairs 256 routed experts with one shared expert — so the routed experts can specialize.
Training has a notorious failure mode: router collapse, where the router locks onto a few favorite experts and starves the rest. The standard remedy is an auxiliary load-balancing loss, inherited from the Switch Transformer work, that penalizes the router whenever traffic concentrates.
The idea itself dates to 2017, when Shazeer et al. described a sparsely-gated 137B-parameter MoE. Per the dev.to tally, six of the seven frontier open-weight models actually deployed in 2026 use the architecture.
Compute savings, memory bill
The catch, as the post frames it, is that sparsity cuts the compute bill, not the memory bill. Every expert sits in VRAM whether or not a given token uses it. DeepSeek V3, at 671B parameters with about 37B active, needs roughly 671 GB of FP8 weights — a multi-GPU server — even though each token computes like a 37B model. V4.1 Flash's checkpoint alone is about 510 GB. Open weights at 552B to 1T parameters mean provisioning serious hardware even when the per-token cost is tiny.
What 2026 routing research found
The dev.to piece surveys three lines of work attacking MoE's chronic routing problems:
- UniPool (arXiv:2605.06665, May 2026) found deep-layer routers mostly redundant: swapping a learned top-k router for uniform random routing in deep layers costs only 1.0 to 1.6 accuracy points. Its fix — one globally shared expert pool across all layers instead of per-layer expert sets — matches or beats vanilla MoE using 41.6% to 66.7% of the expert-parameter budget, with validation loss improving by up to 0.0386.
- Latent Prototype Routing reframes routing as clustering and cuts the Gini coefficient of expert load from 0.70 to 0.035 on DeepSeek-V3, Qwen3-MoE and Mixtral, reportedly with no quality loss.
- An audit dubbed COMMITTEEAUDIT (arXiv:2601.03425) found roughly 6 of 64 experts appear in the top-k for the overwhelming majority of tokens regardless of domain — code, poetry or math — carrying 60% to 67% of total routing weight. Specialization is real, but a small generalist "standing committee" does most of the work.
Serving tax and a verification hole
Because no single GPU holds every expert, serving shards them across GPUs via expert parallelism: each token travels to the GPUs hosting its experts, then the weighted outputs travel back. That all-to-all shuffle is pure overhead, with bandwidth and stragglers capping throughput — the post argues MoE inference resembles HPC work more than dense-model serving.
It also flags a verification blind spot. A June 2026 analysis, "The Expert Shuffle," argues that a decentralized provider can quietly cut top-k from 8 to 4 — "k-shunting" — and stay nearly undetectable from outputs alone, because the same standing-committee experts dominate both configurations. Only measurement of intermediate computation catches it; anyone buying third-party MoE inference should ask what the provider actually measures.
Why it matters
The September pair shows where frontier open weights are heading: 552B-to-1T parameter capacity with double-digit-billion active counts, MIT licenses, and per-token prices that resemble small dense models. For practitioners, capacity and cost are now separate line items — total parameters set the memory bill, active parameters set the compute bill, and deployment planning has to budget VRAM for the whole model, not the active slice. The research also suggests much of the advertised expert count may be decorative, given standing committees and deep-layer redundancy. And with friendly licenses making the weights themselves near-commodity, the dev.to analysis concludes that the competitive edge has shifted to the serving layer — and to being able to prove what compute a provider actually delivered.
- #deepseek
- #xiaomi
- #mixture-of-experts
- #open-weights
- #llm