CURRICULUMDERIVING YOUR OWN MODEL01

Post-training in 2026: what is actually live

Two of the most-cited fine-tuning libraries no longer take bug fixes. Here is what replaced them, and what each method can and cannot change.

Reading time
4 min
Words
866
Sources cited
7
Updated
2026-09-05
Prerequisites
None

The tutorial runs; the library is gone

A large share of published fine-tuning guidance points at tooling that no longer accepts fixes. The code still installs. The examples still run. Nobody is maintaining the path underneath them.

torchtune's README now opens with a notice that the project is no longer actively maintained and that development wound down in 2025. Hugging Face AutoTrain's README states that the project is no longer maintained, that no new features will be added and bugs will not be fixed, and directs users to Axolotl, TRL or transformers.Trainer instead.

Neither project failed. Both were superseded. But a 2024 tutorial does not know that, and a first search result does not either. Read the repository's own README before following any post-training guide, and check the date of the last release rather than the date of the article.

Parameter-efficient adaptation

LoRA freezes the base weights and trains a pair of low-rank matrices whose product is added to selected projections. The trained artifact is small, training is cheap, and the adapter can be removed, swapped or composed without touching the base checkpoint.

QLoRA extends this by quantising the frozen base to 4-bit and training the adapter on top, which is what brought single-accelerator fine-tuning of large models within reach.

What adaptation gives you is style, format, domain vocabulary and task shape. What it does not give you is capability the base model does not have. If the base model cannot do the reasoning the task requires, an adapter will teach it to fail in your house style.

Preference optimisation without a reward model

Direct Preference Optimization derives a training objective directly from preference pairs, removing the separate reward model and the online sampling loop that classical RLHF requires. In practice it is close to supervised training in operational complexity, which is most of why it is used.

Its ceiling is the preference data. DPO can only express distinctions that appear in the pairs, and it inherits whatever biases the annotators or the judge model brought — including, if a language model produced the preferences, the position and verbosity effects covered in the benchmark track.

Group-relative policy optimisation

GRPO, introduced in DeepSeekMath, keeps the PPO-style surrogate objective and removes the learned value function. Instead of a critic estimating a baseline, it samples a group of completions for each prompt and uses the group's own scores as the baseline.

The practical consequence is memory. PPO requires the policy, a reference copy and a value network resident simultaneously; GRPO removes one full copy of model weights from that budget. That is frequently the difference between a training run that fits and one that does not.

Reinforcement learning from verifiable rewards

RLVR, as formulated in Tulu 3, replaces the reward model with a verification function. Completions are sampled from the policy, checked by a deterministic function, and rewarded only when verifiably correct — otherwise the reward is zero. The policy is then trained against that signal.

The constraint is the entire point. RLVR applies where correctness is mechanically checkable: a numeric answer, a passing test suite, a satisfied output schema, a followed formatting instruction. It does not apply to helpfulness, tone or judgement, and attempting to stretch it there reintroduces exactly the judge problems it was designed to avoid.

It is also the method with the cleanest evaluation story, because the verifier that trains the model can also grade it, and that grade is not a model's opinion.

What each live method needs, and what it can change.
MethodNeedsChangesArtifact
LoRA / QLoRASupervised examplesStyle, format, domain vocabularySmall adapter
DPOPreference pairsResponse selection within existing abilityFull weights or adapter
GRPOA scoring function and sampling budgetPolicy behaviour, without a critic networkFull weights
RLVRA deterministic verifierVerifiable-task accuracyFull weights

Choosing

Start with the cheapest method that could possibly work, and prove the base model can do the task at all before spending on reinforcement learning. If a well-constructed prompt with a handful of examples cannot get the behaviour some of the time, no post-training method will get it reliably.

Then match method to signal. If you have examples, adapt. If you have preferences, use DPO. If you have a checker, use RLVR — it is strictly the strongest signal available, because it does not require anyone to be right about what good looks like.

The rule

Check the maintenance status of the library before the quality of the tutorial. A well-written guide to an unmaintained tool is a well-written dead end.

Every figure in this lesson is sourced or computed. Where a claim would not verify against its source, it was removed rather than qualified.