CURRICULUM / 3 TRACKS

Learn to read the number.

Most disagreements about model quality are not disagreements about models. They are disagreements about measurement conventions, sample sizes, byte accounting and licence chains that nobody computed.

Every load-bearing figure below carries the paper, repository or licence clause that establishes it. Where a claim would not verify, it was cut rather than softened.

10LESSONS
3TRACKS
43MINUTES END TO END
41CITED SOURCES

Reference, not commentary.

These lessons exist because the questions they answer keep being answered wrongly in public: that two published scores for one model can be compared, that a two-point gap on a two-hundred-question benchmark is a result, that active parameters tell you what fits, that a merge inherits the best of its constituents, that an open licence discharges the obligations attached to a derived model.

Each one opens on a specific situation, works through the arithmetic or the clause, and ends with a rule you can apply to the next artifact that crosses your desk. Reading times are counted from the text itself.

TRACK 01

Reading a benchmark honestly

What a published score is evidence of, what it is not evidence of, and the arithmetic that separates a result from a rounding error.

Lessons in Reading a benchmark honestly
No.LessonReadPrerequisite
01Why two labs report different scores for the same model

The same model, the same dataset, and a fifteen-point spread. Prompt format and scoring convention are not implementation details; they are the measurement.

4 minNone
02How many questions does it take to tell two models apart

Most published two-point wins are unfalsifiable. The arithmetic that says so takes one line, and almost nobody runs it.

5 minWhy two labs report different scores for the same model
03Contamination, and how to test for it cheaply

Four tests, none of which needs the pretraining corpus, and one of which needs nothing but a release date.

4 minWhy two labs report different scores for the same model
04LLM-as-judge and its failure modes

A judge that agrees with itself a quarter of the time when you swap the answers is not measuring quality. It is measuring position.

4 minHow many questions does it take to tell two models apart

TRACK 02

Sizing and serving sparse models

Three different parameter counts, the bytes a checkpoint actually occupies, and why a mixture-of-experts model does not behave like a dense one under load.

Lessons in Sizing and serving sparse models
No.LessonReadPrerequisite
01Total, active, and resident: three different numbers

Active parameters describe arithmetic per token. They say nothing about what has to be in memory, and the gap between the two is where capacity plans die.

4 minNone
02Reading a checkpoint's real size

Per-dtype accounting, why an FP8 release is half the size at the same parameter count, and the packing trap that makes a 4-bit checkpoint look eight times larger than it is.

4 minTotal, active, and resident: three different numbers
03Expert routing in production

Load skew, dropped tokens and all-to-all collectives. Why a model that does a fraction of the arithmetic does not deliver a proportional fraction of the latency.

4 minTotal, active, and resident: three different numbers

TRACK 03

Deriving your own model

Post-training methods that are still maintained, what composition really costs, and the licence obligations that travel with a derived artifact.

Lessons in Deriving your own model
No.LessonReadPrerequisite
01Post-training in 2026: what is actually live

Two of the most-cited fine-tuning libraries no longer take bug fixes. Here is what replaced them, and what each method can and cannot change.

4 minNone
02Merging and mixture-of-experts composition, honestly

A four-expert merge that scores below one of its own constituents, the token budget real upcycling costs, and why no metadata-only compatibility score is defensible.

4 minPost-training in 2026: what is actually live
03The licence chain nobody computes

A derived model inherits the base licence, the dataset licence, and the terms of whatever generated its synthetic data. Four traps, and one regulatory duty that open release does not remove.

6 minPost-training in 2026: what is actually live

The curriculum is deliberately narrow. It covers the places where a confident number is routinely wrong, and it stops there. It is not an introduction to transformers, and it does not rank models.

Corrections are welcome and are applied to the record rather than to its history. If a cited figure has been superseded, the lesson changes and the source changes with it.