CURRICULUMDERIVING YOUR OWN MODEL02

Merging and mixture-of-experts composition, honestly

A four-expert merge that scores below one of its own constituents, the token budget real upcycling costs, and why no metadata-only compatibility score is defensible.

Reading time
4 min
Words
896
Sources cited
6
Updated
2026-09-05
Prerequisites
Post-training in 2026: what is actually live

A merge that lost to one of its own experts

Beyonder-4x7B-v3 is a four-expert mixture assembled with mergekit from four finished 7B models. One of those four is AlphaMonarch-7B. On the Nous benchmark suite, the merge averages 61.91 and AlphaMonarch-7B averages 62.74.

The composite is roughly four times the size of one constituent and scores below it on the headline average, on the suite its own author published.

Nous benchmark suite, as published on the Beyonder-4x7B-v3 model card. Three of the six rows are experts inside the merge. These are self-reported figures under a harness configuration the card does not state, so treat the ordering as an illustration rather than a measurement.
ModelAverageAGIEvalGPT4AllTruthfulQABigbench
AlphaMonarch-7B (expert)62.7445.3777.0178.3950.20
Beyonder-4x7B-v3 (merge)61.9145.8576.6774.9850.12
NeuralDaredevil-7B (expert)59.3945.2376.2067.6148.52
Kunoichi-DPO-v2-7B (expert)58.2944.7975.0565.6847.65
Beyonder-4x7B-v257.1345.2975.9560.8646.40
CodeNinja-1.0-OpenChat-7B (expert)50.3539.9871.7748.7340.92

Read the row, not the average

The merge is not uniformly worse. It leads AlphaMonarch on AGIEval, 45.85 against 45.37, and is within a third of a point on Bigbench. It loses the average on TruthfulQA, 74.98 against 78.39.

That is the characteristic signature of weight-time composition: the merge preserves broad knowledge and gives up the specific alignment behaviour that the strongest constituent had been tuned into. Reporting only the average hides which half you lost, and which half you lost is the entire question when you are deciding whether to ship it.

Beyonder-4x7B-v3 also beats every other constituent, and comfortably beats the previous version of itself. Composition is not worthless. It is just not free, and it is not reliably additive.

Why weight-time composition is not upcycling

A weight-time frankenMoE copies feed-forward blocks out of finished models into a mixture-of-experts shell and initialises a router, commonly from a handful of representative prompt embeddings per expert. Everything else — attention, embeddings, norms — is taken from one base.

Nothing in that procedure trains the router on the token distribution it will actually route. The experts were each trained as complete models with their own attention in front of them, and they are now being asked to operate behind someone else's. Sometimes this works. There is no mechanism that makes it work by construction.

What real upcycling costs

Upcycling proper — initialising a sparse model from a dense checkpoint and then continuing to train it — is a training run, not a merge.

NVIDIA upcycled Nemotron-4 15B on 1T tokens and compared it against continued dense training of the same model on the same 1T tokens: 67.6% MMLU for the upcycled model against 65.3% for the dense continuation. The gain is real and it cost a trillion tokens.

The Drop-Upcycling work used 500B tokens of continued training at 8×152M, 8×1.5B and 8×3.7B scales, against dense baselines trained on 1,000B tokens. It also reports that naive upcycling starts ahead and then converges most slowly of the methods compared — the initial advantage from copied weights is paid back later as reduced expert specialisation.

So the honest budget for converting dense checkpoints into a working sparse model is hundreds of billions to a trillion tokens of continued training. It is not an afternoon with a YAML file, and a merge configuration is not a cheaper version of it. It is a different operation with a different expected outcome.

Compatibility is not in the metadata

The obvious product is a compatibility score: given two model identifiers, predict whether merging them will work. Every input such a score could cheaply use is the wrong input.

A 2026 study of task-level merging collapse tested both families of predictor. Hidden-state distance — the average distance between two models' internal representations of the same inputs — correlates strongly with merging success, with p-values from 0.001 to 0.006 across merging techniques. Parameter-space conflict metrics do not: parameter magnitude change ratio, sign change ratio, conflicting parameter magnitude and average cosine similarity all returned p-values above 0.05.

The metrics that fail are the cheap ones. The metric that works requires running both models on the same inputs and comparing what happens inside them, which means you need the weights, the compute and a representative input set before you can say anything.

This is why this site publishes no merge-compatibility score. A score computed from architecture family, parameter count, tokenizer and licence would look authoritative and would be predicting from the quantities that were measured not to predict. Shipping it would be worse than shipping nothing.

The rule

Composition is an experiment, not a construction. Evaluate the merge against every constituent on the same suite, and report per-task results — a merge that loses to its own best expert is a common outcome, not an anomaly.

Every figure in this lesson is sourced or computed. Where a claim would not verify against its source, it was removed rather than qualified.