The tutorial runs; the library is gone
A large share of published fine-tuning guidance points at tooling that no longer accepts fixes. The code still installs. The examples still run. Nobody is maintaining the path underneath them.
torchtune's README now opens with a notice that the project is no longer actively maintained and that development wound down in 2025. Hugging Face AutoTrain's README states that the project is no longer maintained, that no new features will be added and bugs will not be fixed, and directs users to Axolotl, TRL or transformers.Trainer instead.
Neither project failed. Both were superseded. But a 2024 tutorial does not know that, and a first search result does not either. Read the repository's own README before following any post-training guide, and check the date of the last release rather than the date of the article.
Parameter-efficient adaptation
LoRA freezes the base weights and trains a pair of low-rank matrices whose product is added to selected projections. The trained artifact is small, training is cheap, and the adapter can be removed, swapped or composed without touching the base checkpoint.
QLoRA extends this by quantising the frozen base to 4-bit and training the adapter on top, which is what brought single-accelerator fine-tuning of large models within reach.
What adaptation gives you is style, format, domain vocabulary and task shape. What it does not give you is capability the base model does not have. If the base model cannot do the reasoning the task requires, an adapter will teach it to fail in your house style.
Preference optimisation without a reward model
Direct Preference Optimization derives a training objective directly from preference pairs, removing the separate reward model and the online sampling loop that classical RLHF requires. In practice it is close to supervised training in operational complexity, which is most of why it is used.
Its ceiling is the preference data. DPO can only express distinctions that appear in the pairs, and it inherits whatever biases the annotators or the judge model brought — including, if a language model produced the preferences, the position and verbosity effects covered in the benchmark track.
Group-relative policy optimisation
GRPO, introduced in DeepSeekMath, keeps the PPO-style surrogate objective and removes the learned value function. Instead of a critic estimating a baseline, it samples a group of completions for each prompt and uses the group's own scores as the baseline.
The practical consequence is memory. PPO requires the policy, a reference copy and a value network resident simultaneously; GRPO removes one full copy of model weights from that budget. That is frequently the difference between a training run that fits and one that does not.
Reinforcement learning from verifiable rewards
RLVR, as formulated in Tulu 3, replaces the reward model with a verification function. Completions are sampled from the policy, checked by a deterministic function, and rewarded only when verifiably correct — otherwise the reward is zero. The policy is then trained against that signal.
The constraint is the entire point. RLVR applies where correctness is mechanically checkable: a numeric answer, a passing test suite, a satisfied output schema, a followed formatting instruction. It does not apply to helpfulness, tone or judgement, and attempting to stretch it there reintroduces exactly the judge problems it was designed to avoid.
It is also the method with the cleanest evaluation story, because the verifier that trains the model can also grade it, and that grade is not a model's opinion.
| Method | Needs | Changes | Artifact |
|---|---|---|---|
| LoRA / QLoRA | Supervised examples | Style, format, domain vocabulary | Small adapter |
| DPO | Preference pairs | Response selection within existing ability | Full weights or adapter |
| GRPO | A scoring function and sampling budget | Policy behaviour, without a critic network | Full weights |
| RLVR | A deterministic verifier | Verifiable-task accuracy | Full weights |
Choosing
Start with the cheapest method that could possibly work, and prove the base model can do the task at all before spending on reinforcement learning. If a well-constructed prompt with a handful of examples cannot get the behaviour some of the time, no post-training method will get it reliably.
Then match method to signal. If you have examples, adapt. If you have preferences, use DPO. If you have a checker, use RLVR — it is strictly the strongest signal available, because it does not require anyone to be right about what good looks like.
The rule
Check the maintenance status of the library before the quality of the tutorial. A well-written guide to an unmaintained tool is a well-written dead end.