CURRICULUMSIZING AND SERVING SPARSE MODELS01

Total, active, and resident: three different numbers

Active parameters describe arithmetic per token. They say nothing about what has to be in memory, and the gap between the two is where capacity plans die.

Reading time
4 min
Words
863
Sources cited
4
Updated
2026-09-05
Prerequisites
None

Thirty-seven billion active, and it still needs a rack

DeepSeek-V3 activates 37B parameters per token out of 671B total. A capacity plan that reads the first figure and stops concludes that the model runs on a single 80GB accelerator. It does not run on eight.

The confusion is understandable, because three different quantities are all called the parameter count, and vendors quote whichever one flatters the release.

Three parameter counts, and what each one is evidence about.
CountDefinitionWhat it predicts
TotalEvery parameter in the checkpoint.Storage, and the floor on accelerator memory.
ActiveParameters touched computing one token.Arithmetic per token, and therefore compute cost.
ResidentBytes that must be in accelerator memory to serve.Whether the deployment is possible at all.

Active is a statement about FLOPs

Routing happens per token, per layer, at run time. Over any realistic batch the union of experts touched approaches the full set, and even for a single token the router chooses after the weights would have had to be loaded. There is no schedule that lets you keep only the active fraction resident and still answer arbitrary prompts at speed.

Mixtral 8x7B is the clean illustration because its arithmetic is public. 46.7B total parameters, 12.9B used per token: the model processes and generates at the cost of a 12.9B model. Attention is shared across experts at roughly 5B parameters, and the eight expert feed-forward blocks contribute about 8 × 5.25B = 42B. Serving it means holding 46.7B parameters. Paying for it means paying for 12.9B.

Both numbers are true. They answer different questions, and only one of them is about memory.

What stays dense in a sparse model

Sparsity applies to the feed-forward experts. It does not apply to attention projections, embeddings, the language-model head, layer norms, the router itself, or shared experts that are active on every token.

DeepSeek-V3 has 256 routed experts per MoE layer with 8 activated per token, plus 1 shared expert that runs unconditionally. The shared expert is dense by construction. So the active parameter count never falls to the routed fraction of the total, and the dense remainder sets a hard floor under per-token compute no matter how aggressive the routing gets.

This is also why routing sparsity — routed experts divided by experts per token — is a useful architectural fact and a bad memory heuristic. A 32× routing sparsity does not divide anything by 32.

Weight bytes are not resident bytes

Even once you have the right parameter count and the right byte width, weights are not the whole allocation. The vLLM paper measured a 13B model on a 40GB A100: roughly 65% of memory went to model weights, close to 30% to the dynamic per-request state — the KV cache — and the small remainder to activations and other ephemeral data.

The KV cache is the part that scales with your traffic rather than with the model. It grows with sequence length and with concurrency, and it is the reason two deployments of the same checkpoint can have completely different memory profiles.

It is also historically the most wasted region. Before paged allocation, systems that reserved contiguous KV memory by maximum sequence length used only 20.4% to 38.2% of that region for actual token state; the rest was reservation and fragmentation. Modern runtimes recover most of it, but the general lesson survives: the allocator's behaviour is part of your capacity plan.

On top of that sit CUDA graph capture buffers, which trade memory for launch latency, and the runtime's own working set.

How to use a residency floor

The catalogue publishes a residency floor for each artifact: the ceiling of checkpoint bytes divided by advertised accelerator memory. It is arithmetic on the weights, and it is deliberately not a deployment recommendation.

Use it as a disqualifier, not a plan. If the floor already consumes most of a node, there is no room left for the KV cache, which means there is no room left for concurrency, which means you do not have a serving configuration — you have a loading configuration.

The number that matters, throughput at your latency target under your traffic mix, cannot be derived from a checkpoint manifest. It has to be measured. Everything before that measurement is a filter that tells you which configurations are worth measuring.

The rule

Total parameters bound the memory. Active parameters bound the arithmetic. Neither bounds what a runtime will actually allocate, and only the last one decides whether the deployment works.

Every figure in this lesson is sourced or computed. Where a claim would not verify against its source, it was removed rather than qualified.