CURRICULUM / 3 TRACKS
Learn to read the number.
Most disagreements about model quality are not disagreements about models. They are disagreements about measurement conventions, sample sizes, byte accounting and licence chains that nobody computed.
Every load-bearing figure below carries the paper, repository or licence clause that establishes it. Where a claim would not verify, it was cut rather than softened.
Reference, not commentary.
These lessons exist because the questions they answer keep being answered wrongly in public: that two published scores for one model can be compared, that a two-point gap on a two-hundred-question benchmark is a result, that active parameters tell you what fits, that a merge inherits the best of its constituents, that an open licence discharges the obligations attached to a derived model.
Each one opens on a specific situation, works through the arithmetic or the clause, and ends with a rule you can apply to the next artifact that crosses your desk. Reading times are counted from the text itself.
TRACK 01
Reading a benchmark honestly
What a published score is evidence of, what it is not evidence of, and the arithmetic that separates a result from a rounding error.
| No. | Lesson | Read | Prerequisite |
|---|---|---|---|
| 01 | Why two labs report different scores for the same model The same model, the same dataset, and a fifteen-point spread. Prompt format and scoring convention are not implementation details; they are the measurement. | 4 min | None |
| 02 | How many questions does it take to tell two models apart Most published two-point wins are unfalsifiable. The arithmetic that says so takes one line, and almost nobody runs it. | 5 min | Why two labs report different scores for the same model |
| 03 | Contamination, and how to test for it cheaply Four tests, none of which needs the pretraining corpus, and one of which needs nothing but a release date. | 4 min | Why two labs report different scores for the same model |
| 04 | LLM-as-judge and its failure modes A judge that agrees with itself a quarter of the time when you swap the answers is not measuring quality. It is measuring position. | 4 min | How many questions does it take to tell two models apart |
TRACK 02
Sizing and serving sparse models
Three different parameter counts, the bytes a checkpoint actually occupies, and why a mixture-of-experts model does not behave like a dense one under load.
| No. | Lesson | Read | Prerequisite |
|---|---|---|---|
| 01 | Total, active, and resident: three different numbers Active parameters describe arithmetic per token. They say nothing about what has to be in memory, and the gap between the two is where capacity plans die. | 4 min | None |
| 02 | Reading a checkpoint's real size Per-dtype accounting, why an FP8 release is half the size at the same parameter count, and the packing trap that makes a 4-bit checkpoint look eight times larger than it is. | 4 min | Total, active, and resident: three different numbers |
| 03 | Expert routing in production Load skew, dropped tokens and all-to-all collectives. Why a model that does a fraction of the arithmetic does not deliver a proportional fraction of the latency. | 4 min | Total, active, and resident: three different numbers |
TRACK 03
Deriving your own model
Post-training methods that are still maintained, what composition really costs, and the licence obligations that travel with a derived artifact.
| No. | Lesson | Read | Prerequisite |
|---|---|---|---|
| 01 | Post-training in 2026: what is actually live Two of the most-cited fine-tuning libraries no longer take bug fixes. Here is what replaced them, and what each method can and cannot change. | 4 min | None |
| 02 | Merging and mixture-of-experts composition, honestly A four-expert merge that scores below one of its own constituents, the token budget real upcycling costs, and why no metadata-only compatibility score is defensible. | 4 min | Post-training in 2026: what is actually live |
| 03 | The licence chain nobody computes A derived model inherits the base licence, the dataset licence, and the terms of whatever generated its synthetic data. Four traps, and one regulatory duty that open release does not remove. | 6 min | Post-training in 2026: what is actually live |
The curriculum is deliberately narrow. It covers the places where a confident number is routinely wrong, and it stops there. It is not an introduction to transformers, and it does not rank models.
Corrections are welcome and are applied to the record rather than to its history. If a cited figure has been superseded, the lesson changes and the source changes with it.