Same parameter count, half the bytes
Two repositories publish the same model at the same parameter count. One occupies roughly two bytes per parameter, the other roughly one. Nothing was pruned and no layer was removed. One shipped BF16 tensors and the other shipped FP8.
The BF16 build is close to twice the size of the FP8 build of the same weights, and the live table further down shows the pair. At frontier scale that ratio is the difference between eight accelerators and sixteen. It is invisible in every summary that reports a parameter count, which is most summaries.
The parameter count is the wrong unit. The unit that decides deployments is bytes, and bytes are not derivable from parameters without knowing the width of every tensor.
Sum per dtype, not per parameter
The safetensors index reports parameter counts keyed by dtype. The correct size calculation multiplies each dtype's element count by that dtype's width and sums the results. There is no shortcut through the total.
The shortcut fails because real checkpoints mix widths. An FP8 release commonly keeps layer norms, embeddings and the router in BF16 while quantising the expert and attention projections, so its measured bytes-per-parameter lands slightly above one rather than exactly at one. A parameter count multiplied by a single assumed width cannot reproduce that, and the error is not small.
This is why every catalogue record here carries a per-dtype breakdown rather than an estimate, and why records whose dtype map contains a width the project cannot establish are marked undetermined instead of being filled with a plausible default.
The packing trap
Low-bit quantization does not store low-bit elements. It packs several quantized weights into one wider integer container, because there is no hardware type for a four-bit number.
MLX documents the layout exactly: quantized weights are packed into unsigned 32-bit integers, and at four bits per weight it fits eight elements in an unsigned 32-bit integer, where the first element occupies the four least significant bits, the second bits four to seven, and so on. GPTQ and AWQ store their qweight tensors the same way.
The safetensors index does not see the packing. It reports the container element count and the container dtype. Multiply those two together and you get container arithmetic: for 4-bit weights in a 32-bit container, a figure eight times the real weight bytes. The number is deterministic, reproducible and completely wrong as a size.
FP8 is not affected. FP8 is a real hardware dtype: one element occupies one byte, and element count multiplied by width is the truth. The trap belongs to formats that pack, and packing is not confined to four bits — an eight-bit MLX tensor is four elements to a 32-bit container, so the same overstatement applies at a factor of four.
Signals that a checkpoint is packed
- A
quantization_configdeclaring an integer weight bit-width of eight or fewer. A float type of the same nominal width, such as FP8, is not the same thing. - An integer container dtype —
I32,U32,I8used as a container — appearing in the dtype map. - A packing library named in the repository tags or library field: MLX, GPTQ, AWQ and their derivatives.
- A bytes-per-parameter figure near 4.0 computed from the dtype map of a repository whose name says 4-bit. That is the container width, not the weight width.
One model family, several artifacts
Every row below is the same base model at the same parameter count, resolved live from the catalogue snapshot at build time. The basis column is the one to read first: an exact per-dtype sum where the safetensors index can be taken literally, and measured repository storage where the checkpoint packs sub-byte weights into wider containers and the tensor product would overstate the size.
| Repository | Revision | Parameters | Weight bytes | Bytes / parameter | Basis |
|---|---|---|---|---|---|
| Qwen/Qwen3-30B-A3B | ad44e77 | 30.5B | 61.1 GB | 2.00 | Summed per dtype from the safetensors index |
| Qwen/Qwen3-30B-A3B-FP8 | d206ba7 | 30.5B | 32.4 GB | 1.06 | Summed per dtype from the safetensors index |
| Qwen/Qwen3-30B-A3B-GPTQ-Int4 | 9b534e4 | 30.5B | 16.9 GB | 0.55 | Repository storage. Weights are packed into wider containers, so tensor arithmetic would overstate the size. |
| QuixiAI/Qwen3-30B-A3B-AWQ | 1ba5586 | 30.5B | 16.8 GB | 0.55 | Repository storage. Weights are packed into wider containers, so tensor arithmetic would overstate the size. |
| mlx-community/Qwen3-30B-A3B-4bit | d388dea | 30.5B | 17.2 GB | 0.56 | Repository storage. Weights are packed into wider containers, so tensor arithmetic would overstate the size. |
| mlx-community/Qwen3-30B-A3B-8bit | 7d5b2e5 | 30.5B | 32.5 GB | 1.06 | Repository storage. The tensor index does not establish an exact weight total. |
Rows are dropped when the identifier is absent from the current snapshot, so this table shrinks rather than lying when the Hub changes underneath it.
Use the measured file size when the map is packed
When the dtype map cannot be read literally, the honest number is the repository's stored bytes as measured on disk. It is a real measurement rather than an inference, and it is what you will actually download.
It is an upper bound on weight bytes rather than a substitute for them. A repository holds more than tensors: tokenizer files, configuration, generation config, a model card, and frequently more than one copy of the model. Publishers routinely ship legacy PyTorch .bin shards alongside safetensors, or a consolidated single-file build next to the sharded one, or several quantizations in one repository. Every one of those is counted in stored bytes and none of them is loaded twice at run time.
So the discipline is to say which number you are quoting. Exact tensor bytes summed per dtype where the index can be read. Measured repository storage where it cannot, labelled as storage and as an upper bound. Never the two silently interchanged.
The wrong move is to keep the dtype product and add a disclaimer. A number that is eight times too large is not a conservative estimate; it disqualifies hardware that would have worked, and it does so with the authority of an exact-looking figure.
The rule
Never derive a checkpoint's size from its parameter count. Sum per dtype where the index permits it, use measured storage where it does not, and state which of the two you published.