CURRICULUMSIZING AND SERVING SPARSE MODELS03

Expert routing in production

Load skew, dropped tokens and all-to-all collectives. Why a model that does a fraction of the arithmetic does not deliver a proportional fraction of the latency.

Reading time
4 min
Words
908
Sources cited
4
Updated
2026-09-05
Prerequisites
Total, active, and resident: three different numbers

A fraction of the FLOPs, nothing like the speedup

A team moves from a dense model to a sparse one with a small fraction of the active parameters, expects a large throughput gain, and measures something far smaller. Nothing is misconfigured. The arithmetic did shrink. The bottleneck moved.

Three things about routing determine what a sparse model does under load, and none of them appears on a model card.

Routing is data-dependent, so load is a property of your traffic

Which expert a token goes to depends on the token. Expert load is therefore a random variable whose distribution depends on the prompt mixture you actually serve, not on anything the publisher measured.

A model balanced on a general pretraining mixture can be badly skewed on a single-domain workload. Code-heavy traffic, one language, one document template — each concentrates routing. The skew shows up as one device finishing late while the rest wait, which turns a throughput problem into a tail-latency problem.

This is measurable in an afternoon: instrument per-expert token counts per layer over a representative prompt sample and look at the ratio of maximum to mean load. It is the single most informative number about a sparse deployment, and almost nobody collects it.

Capacity factor, and the tokens that quietly disappear

Switch Transformer introduced the expert capacity factor: a cap on how many tokens each expert will process in a step. Tokens beyond the cap overflow. Depending on the implementation they are routed to the next-best expert or dropped entirely and passed to the next layer through the residual connection.

A larger capacity factor reduces dropping and costs memory and compute. Switch Transformers were reported to perform well at low capacity factors of 1.0 and 1.25, which is precisely the regime where drops occur.

The operational hazard is silence. A dropped token is not an error. Nothing in the response says a layer skipped it. Quality degrades under exactly the traffic conditions that cause skew — which is to say, under load — and there is no log line unless you added one.

Balancing at training time and at serving time

At training time the classical fix is an auxiliary load-balancing loss: a scaled dot product between the fraction of tokens dispatched to each expert and the fraction of router probability assigned to each expert, which pushes routing towards uniformity. DeepSeek-V3 instead uses an auxiliary-loss-free balancing strategy, on the argument that the auxiliary loss buys balance by degrading the model.

At training time DeepSeek-V3 also constrains topology directly, limiting how many nodes any single token may be dispatched to, which bounds cross-node traffic by construction rather than by hope.

At serving time the problem returns, because your traffic is not the training mixture. DeepSeek's Expert Parallelism Load Balancer replicates heavily loaded experts and packs the replicas across devices so that the aggregate load per device evens out. Redundant experts cost memory to buy balance. That trade is the serving-time version of the capacity-factor trade.

All-to-all is the cost nobody budgets for

Under expert parallelism, every MoE layer performs two collective communications: a dispatch that sends each token to the devices holding its chosen experts, and a combine that brings the results back. Both are all-to-all patterns, and both happen on every layer of every forward pass.

DeepSeek-V3's training pipeline splits each chunk into four components — attention, all-to-all dispatch, MLP, all-to-all combine — specifically so that computation and communication can be overlapped. The fact that this is an architectural concern in the training schedule is the clearest available statement of how large the communication term is.

Cross-node routing runs at the speed of the slowest link in the path. A configuration that looks identical on paper can differ by a large factor in practice depending on whether experts for a typical token sit within one node or across a slower fabric.

What to measure before committing

A sparse model's serving profile is not derivable from its checkpoint. These are the measurements that make it knowable.

  1. Per-expert token counts per layer over a representative prompt mixture, reported as maximum-to-mean load ratio.
  2. Token drop rate at your configured capacity factor, at your real concurrency rather than at batch size one.
  3. Share of step time spent in dispatch and combine, separated from compute.
  4. Time to first token and inter-token latency at the concurrency you intend to run, not at the concurrency that benchmarks well.
  5. All of the above again after any change to the traffic mixture. Routing behaviour is a function of the inputs, so a workload change is a new measurement.

The rule

Sparsity moves the bottleneck from arithmetic to memory and interconnect. A model that does a fifth of the FLOPs will not run five times faster, and the difference is decided by your traffic, not by the architecture.

Every figure in this lesson is sourced or computed. Where a claim would not verify against its source, it was removed rather than qualified.