TRAINING STUDIO / RECIPE COMPILER
This page compiles the run. It does not run it.
A browser cannot train a model, and a product that implies otherwise has already lied to you once. What a browser can do is bind an exact base checkpoint to an exact dataset and an exact framework version, compute what the run will need, name the obligations it inherits, and hand you a configuration that is byte-identical every time you generate it.
Nothing here is executed. Every recipe is emitted with status generated_not_executed, and no loss curve, accuracy or duration is shown for a run nobody has performed.
WHY A RECEIPT
Every fine-tuning product returns a model ID.
None of them returns a record binding that model to the revision of the checkpoint it started from, the content hash of the data it read, the digest of the configuration it ran, and the versions of the code that executed it. Without those, a fine-tuned model is a claim about a moving repository name.
This studio emits the configuration and specifies the record. The inference side of the same discipline already exists here as the Deployment Passport.
A training receipt is a tamper-evident provenance record, not a reproduction guarantee. Training is not bitwise reproducible across accelerator types, so the word “reproducible” does not belong next to it and is not used here.
Base checkpoint
Method and framework
Axolotl is pinned at 0.18.0, read from its package index on 2026-09-05. Methods listed below are the ones the project’s own documentation establishes; absence means unestablished, not unsupported.
Architecture dimensions
The Hub’s config summary omits these for most architectures, so they are inputs rather than derivations. Read them from the repository’s config.json. Leaving one blank makes the components that depend on it undetermined instead of guessed.
Dataset and hardware
Memory floor
| Component | Bytes | Basis |
|---|---|---|
| Base weights (4-bit) | 376 MB | Total parameters × 0.5 bytes. QLoRA holds the frozen base at 4 bits; the per-block quantisation constants and any modules the framework keeps at a wider dtype are not counted here. |
| Gradients (BF16) | undetermined | The adapter's parameter count depends on which projection matrices the framework adapts and on the model's hidden and intermediate dimensions. Neither is carried by the catalogue record. Start the run once, read the trainable-parameter line the framework prints on the first step, and enter it above. |
| Optimiser state (AdamW 8-bit) | undetermined | The adapter's parameter count depends on which projection matrices the framework adapts and on the model's hidden and intermediate dimensions. Neither is carried by the catalogue record. Start the run once, read the trainable-parameter line the framework prints on the first step, and enter it above. |
| Adapter weights (BF16) | undetermined | The adapter's parameter count depends on which projection matrices the framework adapts and on the model's hidden and intermediate dimensions. Neither is carried by the catalogue record. Start the run once, read the trainable-parameter line the framework prints on the first step, and enter it above. |
| Activations (checkpointed) | undetermined | With full activation checkpointing the framework stores one layer-input tensor per layer: micro-batch × sequence length × hidden size × layers × bytes per element. Hidden size and layer count are missing — both sit in the repository's config.json as `hidden_size` and `num_hidden_layers`. Enter them above and this row resolves. |
| Loss logits (fp32) | undetermined | Cross-entropy over the vocabulary materialises a micro-batch × sequence × vocabulary tensor. The vocabulary size sits in config.json as `vocab_size`. Enter it above and this row resolves. |
| Accounted floor | 376 MB | Sum of the components above that could be computed. Not a total. |
| Per device | 80.0 GB | Budget 80.0 GB across 1 device |
The accounted components fit, but 5 components could not be computed from these inputs. Whether the run fits is undetermined, and no arrangement of unknowns makes it a yes.
5 components are undetermined, so the accounted floor is lower than the real requirement by an unknown amount. Fill in the architecture dimensions above to close the gap.
Guardrails
MaterialThe Hub dataset is not pinned to a revision
A Hub dataset repository is mutable in exactly the way a model repository is. Pin it with `--revision <sha>` and record that SHA, or download it once and record the SHA-256 of the bytes on disk. Otherwise 'trained on this dataset' names a moving target.
Source ↗MaterialNo dataset content digest
The receipt binds the run to the exact bytes the training loop read. Compute `shasum -a 256` over the file the run will open and carry that digest through to the receipt. A dataset name is not a dataset.
MaterialOverlap with the benchmarks you intend to report is unchecked
If any part of this dataset overlaps a benchmark you later quote, the number you quote measures memorisation rather than capability. Instruction sets assembled from the open web routinely contain benchmark items verbatim, and synthetic data generated by a model that saw those benchmarks reproduces them at a rate nobody's upload flow measures. Run an n-gram overlap pass against every evaluation you plan to report, and record the result — including a null result — in the receipt.
MaterialNo held-out split
Without a split the training loop never sees, a falling loss curve is a statement about fit to the training set and nothing else. Hold data out before the first run, not after a result you like.
MaterialA run without a baseline cannot tell you whether it helped
Evaluate the pinned base checkpoint on the same harness, at the same version, with the same prompts and decoding parameters, before training starts — and keep the raw output. A post-training score with no matched pre-training score on the same harness is uninterpretable: it cannot distinguish a real improvement from a prompt-format change, a harness version bump or noise. Report both numbers with their error bars or report neither.
MaterialThe output inherits every licence in the chain
A fine-tuned model is a derivative work of everything that went into it. This run's output is constrained by the base model's licence (apache-2.0), the dataset's licence. The most restrictive term in that chain governs the result — not the most convenient one — and a permissive base does not launder a restricted dataset.
NoteHalf the public fine-tuning corpus points at stopped projects
Axolotl was verified live on 2026-09-05 against its own repository. torchtune wound down in 2025 and its README says so; Hugging Face AutoTrain Advanced states that it is no longer maintained, that no new features will be added and that bugs will not be fixed. Both are still the target of tutorials, course material and model cards published after they stopped. Check the liveness matrix below before following any fine-tuning guide.
Source ↗
Compiled recipe
# MOEModels training recipe
# Base repository : Qwen/Qwen3-0.6B
# Base revision : c1899de289a04d12100db370d81485cdf75e47ca
# Method : QLoRA (4-bit base)
# Framework : Axolotl 0.18.0
# Dataset : HuggingFaceH4/ultrachat_200k (train)
# Dataset digest : NOT COMPUTED
# Seed : 3407
#
# This file carries no timestamp and no random value, so identical inputs produce identical bytes. That is what makes the digest in recipe.lock.json worth recording.
# MOEModels compiled this run. It did not run it, and it makes no claim about the result.
base_model: "./base/Qwen3-0.6B"
sequence_len: 4096
adapter: qlora
lora_r: 32
lora_alpha: 64
lora_dropout: 0.0
lora_target_linear: true
datasets:
- path: "HuggingFaceH4/ultrachat_200k"
split: "train"
type: "chat_template"
val_set_size: 0.05
output_dir: "./out/Qwen3-0.6B-qlora"
micro_batch_size: 1
gradient_accumulation_steps: 8
num_epochs: 2
learning_rate: 0.0002
lr_scheduler: cosine
warmup_ratio: 0.03
optimizer: adamw_bnb_8bit
gradient_checkpointing: true
bf16: true
fp16: false
seed: 3407
logging_steps: 10
saves_per_epoch: 1
evals_per_epoch: 2
load_in_4bit: true
# Versions MOEModels did not read from a package index are left unpinned rather than
# guessed. After the first successful install, freeze the environment and record it.
Commands
# 1. Authenticate. Required for a gated repository, harmless otherwise.
hf auth login --token "$HF_TOKEN"
# 2. Pin the base checkpoint to an immutable revision.
hf download Qwen/Qwen3-0.6B --revision c1899de289a04d12100db370d81485cdf75e47ca --local-dir base/Qwen3-0.6B
# 3. Pin the dataset and record the bytes the run will actually read.
hf download HuggingFaceH4/ultrachat_200k --repo-type dataset --local-dir data/ultrachat_200k # unpinned: add --revision <SHA>
find data/ultrachat_200k -type f -exec shasum -a 256 {} \; | tee dataset.sha256
# 4. Pin the toolchain. Unpinned entries are ones MOEModels did not verify.
python -m venv .venv && . .venv/bin/activate
pip install "axolotl==0.18.0" datasets accelerate bitsandbytes
pip freeze > environment.lock.txt
# 5. Measure the base BEFORE training, on the harness you intend to report.
# Keep the raw output. A post-training score without this is uninterpretable.
# 6. Record the recipe digest that the receipt will carry.
shasum -a 256 recipe.lock.json
# 7. Run, with an explicit wall-clock ceiling.
timeout 10800 axolotl train axolotl.yaml 2>&1 | tee train.log
# 8. Measure the result on the same harness, at the same version, same prompts.
# 9. Publish, or keep it local.
# result saved to ./out/Qwen3-0.6B-qlora — no Hub repository configured- Recipe fingerprint
0726cf1ca8deec5eNon-cryptographic. It detects accidental drift between two recipes; it is not a security property and must not be presented as one.- Recipe digest (SHA-256)
- Unavailable in this context. WebCrypto is absent, so compute it where the recipe runs with
shasum -a 256 recipe.lock.json.
THE STATE OF THE STACK
Half the published guidance points at finished projects.
Tutorials, course material and model cards still recommend libraries whose own documentation says they have stopped. Each status below is the project’s own wording, not a judgement about quality, with the date it was read.
| Project | Status | Licence | Hardware floor | Methods | Checked |
|---|---|---|---|---|---|
Unsloth2026.9.2 | LiveRepository documents active development; version below published to PyPI. | Apache-2.0 core, AGPL-3.0 Studio UI | NVIDIA, AMD or Intel GPU | LoRA, QLoRA, full fine-tune, DPO, GRPO, FP8 | 2026-09-05 |
Axolotl0.18.0 | LiveRepository documents changes through August 2026, including MoE LoRA training. | Apache-2.0 | NVIDIA (Ampere or newer for bf16) or AMD GPU | Full, LoRA, QLoRA, QAT, DPO, IPO, KTO, ORPO, GRPO, GDPO | 2026-09-05 |
TRL1.12.0 | LiveVersioned documentation published for the release below; repository under active development. | Apache-2.0 | Single GPU to multi-node; DeepSpeed and FSDP supported | SFT, DPO, GRPO, KTO, reward modelling; LoRA and QLoRA via PEFT | 2026-09-05 |
LLaMA-Factory0.9.5 | LiveChangelog documents a Megatron-core training backend added October 2025; repository active. | Apache-2.0 | NVIDIA (CUDA 11.6+), AMD ROCm, or Ascend NPU | Full, freeze, LoRA, QLoRA, OFT, reward modelling, PPO, DPO, KTO, ORPO, SimPO | 2026-09-05 |
MLX-LM0.31.3 | LiveRepository under active development; package published to PyPI. | MIT | Apple silicon only | LoRA, DoRA, full fine-tune; QLoRA by pointing at a quantised base | 2026-09-05 |
NVIDIA NeMo RL0.7.0 | LiveRelease 0.7.0 published 25 July 2026; documented roadmap. | Apache-2.0 | NVIDIA GPU with CUDA; single GPU to multi-node cluster | SFT, LoRA/PEFT, DPO, GRPO, DAPO, GDPO, distillation | 2026-09-05 |
| torchtune | Wound downThe README carries a maintenance notice stating that torchtune is no longer actively maintained and that development wound down in 2025. | BSD-3-Clause | NVIDIA GPU | LoRA, QLoRA, full fine-tune, DPO, PPO | 2026-09-05 |
| Hugging Face AutoTrain Advanced | UnmaintainedThe README states the project is no longer maintained, that no new features will be added and that bugs will not be fixed, and directs users to Axolotl, TRL or transformers.Trainer. | Apache-2.0 | NVIDIA GPU, or Hugging Face Spaces hardware | SFT, LoRA, DPO, ORPO, reward modelling | 2026-09-05 |
THE TRAINING RECEIPT
What a trustworthy fine-tune record has to carry.
This is the specification, not a claim that any run has produced one. A record missing any field below cannot support the statement “this model was trained from that checkpoint on that data”.
- Base checkpoint
- Repository, immutable commit revision, and the digest of the weights actually read.
- Dataset
- Identifier, revision, and a content hash over the exact bytes the run consumed.
- Recipe
- A digest over the canonical configuration, commands and pinned versions.
- Framework
- Name and exact version, not a range.
- Hardware
- Accelerator model, device count and topology.
- Seed
- Every seed the framework accepts, recorded whether or not it changed anything.
- Before and after
- The same evaluation on the same pinned harness, run on the base and the result, with intervals.
- Author
- An optional signature over the record, so a third party can tell who asserts it.