MoE COMPOSER / TIER 1 AND TIER 3

Merging is cheap. Winning is not.

Three different things get called "combining models into a mixture of experts". One takes minutes and usually loses to a model you already had. One works and costs a pretraining run. One needs no GPUs at all and is the least crowded of the three.

This page compiles a real mergekit-moe configuration against catalogue-exact revisions, reports what the metadata mechanically establishes, and refuses to score what it cannot measure. Snapshot 2026-09-05.

TierMethodCost to tryWhat it needsWhat the published evidence says
1Weight-time mergemergekit-moeMinutes on a CPU, or one 8 GB GPUDense donors of one architecture, at one scaleThe canonical showcase merge scores below one of its own experts. Treat a gain as the surprise, not the default.
2UpcyclingSparse upcycling · Branch-Train-MiX · LLaMA-MoE · OLMoEHundreds of billions of training tokensA pretraining cluster and a data pipelineWorks. Whether the head start lasts is unsettled: in OLMoE's own ablation a from-scratch MoE had caught up before 600B tokens, on a comparison the authors qualify themselves.
3Inference-time routingModel routing · mixture-of-agentsNo GPUs. The endpoints you already haveA query set from your own traffic, and the discipline to calibrate on itEvery router in the RouterArena evaluation falls short of the oracle. That is a gap to work in, not a solved market.

THE BASE RATE

The showcase merge loses to one of its own experts.

Beyonder-4x7B-v3 is the model the frankenMoE technique is usually demonstrated with. Its own model card publishes the Nous benchmark suite results, and AlphaMonarch-7B — one of the four 7B models merged into it — comes out ahead on the average. Nobody quotes that row. It is on the same card as the row everybody quotes.

ModelAverageAGIEvalGPT4AllTruthfulQABigbench
mlabonne/Beyonder-4x7B-v3the four-expert merge61.9145.8576.6774.9850.12
mlabonne/AlphaMonarch-7Bone of its own four experts62.7445.3777.0178.3950.2
Differencemerge minus best expert-0.830.48-0.34-3.41-0.08

WHAT THE LOSS COST

Merged artifact
24.2B24,153,690,112 parameters
Checkpoint
48.3 GBsummed per dtype from the safetensors index
Beaten by
one 7B expertalready inside the merge
Inherited licence
cc-by-nc-4.0non-commercial, from the merge chain

Scores transcribed from the model card, read 2026-09-05. Differences computed here. Artifact figures from the catalogue record.

WHAT WOULD HAVE TO BE TRUE FOR YOURS TO DIFFER

  1. The experts have to disagree usefully. A merge of four models fine-tuned from the same base on overlapping data has four copies of nearly the same MLP. Routing between them buys nothing and costs memory for all of them.
  2. Your traffic has to be mixed. The gain, where there is one, comes from queries that genuinely belong to different experts. A single-domain workload has a single best expert, and you already own it.
  3. The router has to be right often enough. Gates derived from a handful of prompts are a weak signal, and a wrong route is worse than no route.
  4. You have to measure against the experts, not a leaderboard. Beating some unrelated model proves nothing. The only comparison that settles it is the merge against each of its own constituents, on the same questions.

TIER 1 · COMPILE IT ANYWAY

Pick the members. Read what the metadata establishes.

Every source below is pinned to the commit revision this catalogue indexed, so the merge you run is the merge this page described. The checks are mechanical: each one names the API field it read, and none of them returns a verdict.

51 of 2,601 · 3/8 experts selected

BaseExpertRepositoryRevisionArchitectureParametersCheckpointLicenceTemplateDownloads
Qwen/Qwen3-0.6Bc1899deqwen3752M1.5 GBapache-2.0yes22.0M
Qwen/Qwen3-8Bb968826qwen38.2B16.4 GBapache-2.0yes13.5M
Qwen/Qwen3-Embedding-0.6B97b0c61qwen3596M1.2 GBapache-2.0yes7.1M
Qwen/Qwen3-4B1cfa9a7qwen34.0B8.0 GBapache-2.0yes6.3M
Qwen/Qwen3-32B9216db5qwen332.8B65.5 GBapache-2.0yes5.1M
Qwen/Qwen3-4B-Instruct-2507cdbee75qwen34.0B8.0 GBapache-2.0yes3.6M
Qwen/Qwen3-1.7B70d244cqwen32.0B4.1 GBapache-2.0yes3.5M
Qwen/Qwen3-Embedding-4B5cf2132qwen34.0B8.0 GBapache-2.0yes2.8M
RadixArk/Kimi-K3-DSpark3c5bac3qwen32.2B4.5 GBundeclaredunknown2.7M
Qwen/Qwen3-Reranker-4B22e6836qwen34.0B8.0 GBapache-2.0yes2.5M
Qwen/Qwen3-Embedding-8B1d8ad4cqwen37.6B15.1 GBapache-2.0yes2.2M
Qwen/Qwen3-4B-Base906bfd4qwen34.0B8.0 GBapache-2.0yes2.0M
Qwen/Qwen3-1.7B-Baseea980cbqwen31.7B3.4 GBapache-2.0yes1.9M
trl-internal-testing/tiny-Qwen3ForCausalLM52b2e48qwen32M5 MBundeclaredunknown1.8M
Qwen/Qwen3-14B40c0698qwen314.8B29.5 GBapache-2.0yes1.8M
Qwen/Qwen3-Reranker-0.6Be61197eqwen3596M1.2 GBapache-2.0yes1.2M
deepseek-ai/DeepSeek-R1-0528-Qwen3-8B6e8885aqwen38.2B16.4 GBmityes1.0M
Qwen/Qwen3-0.6B-Baseda87bfbqwen3596M1.2 GBapache-2.0yes940K
Qwen/Qwen3-14B-Base0b0bd37qwen314.8B29.5 GBapache-2.0yes631K
trl-internal-testing/small-Qwen3ForCausalLMa3f682eqwen339M78 MBundeclaredunknown579K
Qwen/Qwen3-8B-Base49e3418qwen38.2B16.4 GBapache-2.0yes420K
Qwen/Qwen3-4B-Thinking-2507768f209qwen34.0B8.0 GBapache-2.0yes372K
incoai/Qwen3.8-27B-DFlash2dedf8dfqwen31.9B3.8 GBapache-2.0unknown284K
voyageai/voyage-4-nano67fabc9qwen3346M693 MBapache-2.0yes280K
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-8bite524281qwen38.2B—mityes273K
RadixArk/Qwen3.8-27B-DSparkb9a5dbdqwen31.9B3.7 GBotherunknown268K
z-lab/Qwen3.8-27B-DFlash250307d4qwen31.9B3.8 GBapache-2.0unknown242K
trl-internal-testing/tiny-Qwen3ForCausalLM-Instruct-25076f2ac43qwen32M5 MBundeclaredunknown234K
z-lab/Qwen3.6-35B-A3B-DFlashf181eecqwen3386M772 MBapache-2.0unknown228K
Qwen/Qwen3Guard-Gen-0.6Bfada3b2qwen3752M1.5 GBapache-2.0yes160K
typhoon-ai/typhoon2.5-qwen3-4bce0a741qwen34.0B8.0 GBapache-2.0unknown157K
unsloth/Qwen3-Embedding-4B8edd3ddqwen34.0B8.0 GBapache-2.0yes156K
z-lab/Qwen3.6-27B-DFlash0919688qwen31.7B3.5 GBmitunknown147K
llamafactory/tiny-random-qwen381d6f5fqwen32M5 MBapache-2.0yes122K
Qwen/Qwen3-Reranker-8B77d193cqwen38.2B16.4 GBapache-2.0yes121K
Qwen/Qwen3Guard-Gen-4B6ec4282qwen34.4B8.8 GBapache-2.0yes106K
mlx-community/Qwen3-0.6B-8bit11de968qwen3596M—apache-2.0yes102K
OpenOneRec/OneRec-1.7Bcc4d0b5qwen32.1B4.3 GBapache-2.0yes96K
Goedel-LM/Goedel-Prover-V2-32B851bf85qwen332.8B65.5 GBapache-2.0yes75K
lmstudio-community/Qwen3-14B-MLX-8bit86a7e1cqwen314.8B—apache-2.0yes67K

Gate

hidden — Derives the gate from the hidden-state representations of the positive and negative prompts. The documented default and the best of the three. Runs every prompt through the base model. --load-in-8bit or --load-in-4bit reduces the VRAM this needs.

Routing prompts

The gate is derived from these. One line per prompt. Under gate_mode: hidden, every expert needs at least one positive prompt or mergekit has nothing to build a gate from.

Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218
Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3

Mechanical findings

Each row names the API field it read. There is no overall score and no "compatible" verdict, because the only published predictor of merge collapse is representational distance, which no metadata field carries.

SeverityCheckFindingRead from
Warningparameter-scale-spreadParameter counts span 14× across the members, from 595,776,512 to 8,190,735,360. mergekit-moe requires the shared tensors to have identical shapes, and a differing total is evidence they do not.safetensors.total
Warninglicence-chainEvery member declares apache-2.0, so the merged artifact inherits it. mergekit attaches no licence to its output; you have to declare one, and it cannot be looser than this.cardData.license on every member

Nothing here blocks the merge. Every one of these findings still costs you something later. Licence: every member declares apache-2.0, and the output inherits it.

Merged size

LOWER BOUND16.4 GB

Largest single member's parameter count written at bfloat16 at 2 bytes per parameter. The merged artifact holds every member's FFN stack, so it cannot be smaller than the biggest one.

UPPER BOUND19.1 GB

Every member's parameter count summed and written at bfloat16 at 2 bytes per parameter. This is the size if no weight were shared, which over-counts: the shared tensors are stored once.

An exact figure is shared_weights + experts x ffn_weights + router. Splitting a checkpoint into those two parts requires the layer dimensions in the repository's full config.json, which the Hub's config summary omits. This project will not publish a point estimate it cannot derive.

To compute it exactly this page would need, from each member's full config.json: config.json → hidden_size · config.json → intermediate_size · config.json → num_hidden_layers · config.json → vocab_size. The Hub's config summary does not carry them, so no point estimate is offered.

mergekit-moe configuration

# mergekit-moe configuration generated by MOEModels.ai
# Sources are pinned to the commit revisions indexed in this catalogue snapshot.
base_model: Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
gate_mode: hidden
dtype: bfloat16
experts_per_token: 2
experts:
  - source_model: Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
    # positive_prompts: required by gate_mode "hidden" and not yet written for this expert
  - source_model: Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218
    # positive_prompts: required by gate_mode "hidden" and not yet written for this expert
  - source_model: Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3
    # positive_prompts: required by gate_mode "hidden" and not yet written for this expert

Command

mergekit-moe config.yaml ./out --copy-tokenizer
  • Qwen/Qwen3-0.6B has no positive prompts. gate_mode "hidden" derives the router from them, so mergekit cannot build a gate without at least one.
  • Qwen/Qwen3-8B has no positive prompts. gate_mode "hidden" derives the router from them, so mergekit cannot build a gate without at least one.
  • Qwen/Qwen3-Embedding-0.6B has no positive prompts. gate_mode "hidden" derives the router from them, so mergekit cannot build a gate without at least one.
  • gate_mode hidden runs every prompt through the base model. Add --load-in-8bit or --load-in-4bit if the base does not fit in the VRAM you have.

Sources are pinned with mergekit's path@revision syntax to the commits this snapshot indexed on 2026-09-05, so the merge is reproducible against the exact artifacts described above rather than against whatever the default branches hold later. mergekit-moe documentation ↗

The evaluation that would settle it

Not "how does the merge score". Whether it beats each of the models already inside it, on the same questions, on one pinned harness.

ComparisonReference, pinnedQuestion it answers
merged-vs-Qwen/Qwen3-0.6BQwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47caDoes the merged artifact beat Qwen3-0.6B by at least 3 points, on the same questions?
merged-vs-Qwen/Qwen3-8BQwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218Does the merged artifact beat Qwen3-8B by at least 3 points, on the same questions?
merged-vs-Qwen/Qwen3-Embedding-0.6BQwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3Does the merged artifact beat Qwen3-Embedding-0.6B by at least 3 points, on the same questions?
2,181QUESTIONS REQUIRED
n = (z_{alpha/2} + z_{beta})^2 · 2p(1-p)(1-rho) / delta^2

zα/2 = 1.959964 · zβ = 0.841621 · paired variance 0.2500. One deterministic answer per question, per arm. Resampling answers reduces the count; an unpaired comparison, or a harness that does not record per-question outcomes, raises it sharply.

Reference benchmarkQuestionsSmallest resolvable differenceResolves 3 points?
GPQA Diamondsource1989.95 ptsno
GPQA (main)source4486.62 ptsno
GPQA Extendedsource5465.99 ptsno
Recommended eval floorsource1,0004.43 ptsno
  1. Pin the harnessOne harness at one commit, one prompt template, one decoding configuration, for every arm. A merged model and its experts evaluated on different harness versions are not comparable, and the difference between harness versions is routinely larger than the effect being chased.
  2. Run the arms on identical questions4 arms — the merged artifact and 3 references — over the same question set, in the same order, at temperature 0.
  3. Record per-question outcomesStore the score for every question and every arm, not just the aggregate. The paired analysis that makes this affordable is impossible to run afterwards from summary numbers alone.
  4. Report the paired differencePublish the difference, its standard error and the per-question correlation between arms — not two averages side by side. Two averages with no error bar cannot distinguish a real gain from sampling noise.
  5. Compare against the base rateThe claim is only interesting if the merged artifact beats its best expert. mlabonne/Beyonder-4x7B-v3 did not: 61.91 against 62.74 for mlabonne/AlphaMonarch-7B, one of its own four experts.
  • Detecting a 3-point difference at 80% power and alpha 0.05 needs 2,181 questions under these assumptions.
  • GPQA Diamond has 198 questions. At that size the smallest difference the comparison can resolve is about 9.9 points, so a 3-point result on it is structurally unresolvable — no amount of careful running fixes a sample size.
  • These are sample sizes, not permission. A benchmark large enough to resolve the effect still says nothing about whether the effect transfers to the work you actually do.

Tier 3 — the same models, composed at inference time

No merge, no GPUs, no derivative artifact and therefore no licence chain on an output. The weights stay where they are; a policy decides which endpoint answers. This is the only tier of the three that a browser can carry end to end, and it is measurably under-served.

OrderEndpoint, pinnedRoleParametersOrdered by
1Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3first-attempt596Msafetensors.total, ascending
2Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47caescalation752Msafetensors.total, ascending
3Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218escalation8.2Bsafetensors.total, ascending
# Inference-time routing policy generated by MOEModels.ai
# A policy document, not a vendor configuration. No router product reads this format.
# Weights are never modified: every candidate is served from its own endpoint.
kind: routing_policy
version: 1
escalation:
  - order: 1
    model: Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3
    role: first-attempt
    parameters: 595776512
    accept_when: "TODO: calibrate — see calibration below"
  - order: 2
    model: Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
    role: escalation
    parameters: 751632384
    accept_when: "TODO: calibrate — see calibration below"
  - order: 3
    model: Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218
    role: escalation
    parameters: 8190735360
    accept_when: "TODO: calibrate — see calibration below"
  1. Collect a query set from your own traffic. A router calibrated on someone else's distribution is a router calibrated on someone else's problem.
  2. Score every candidate on every query. The escalation order above is by size; whether size predicts quality on your queries is a measurement, not an assumption.
  3. Compute the oracle: the accuracy you would reach by always picking the best candidate per query. That is the ceiling your policy is competing against, and no published router reaches it.
  4. Set the accept threshold from the measured accuracy-versus-cost curve, then re-measure. A threshold copied from a blog post is a guess wearing a number.

All twelve evaluated routers fall short of the oracle, chiefly by failing to notice when a smaller, cheaper model would have been sufficient. Of the 12 evaluated, the commercial entrant NotDiamond placed 12th on the paper's composite score, which prices accuracy against cost: it reaches 68% accuracy at $9.34 per thousand queries. Routing is the tier a browser can honestly ship end to end; it is not a solved problem, and buying one does not make it solved. RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers ↗

WHY THERE IS NO COMPATIBILITY SCORE

The predictor that works is not in the metadata.

A "will this merge work?" score would be the most clickable thing on this page. It would also be invented. The published work on task-level merging collapse tested two candidate predictors across five merging methods.

Merging methodHidden-state distanceParameter sign conflict
TIESp = 0.006p = 0.379
LAp = 0.001p = 0.460
DAREp = 0.145p = 0.882
SLERPp = 0.001p = 0.408
TAp = 0.006p = 0.659

Representational distance between the members predicts collapse in four of the five methods. Parameter-space sign conflict predicts it in none of them. Every field this catalogue holds — architecture, parameter count, licence, dtype, expert topology — lives on the parameter side of that result. The one signal that works requires running both models over a shared input set and comparing hidden states, which is a measurement, not a lookup.

So the composer reports findings and stops. A number here would be a confident-looking function of the predictor that was tested and failed.

An Empirical Study and Theoretical Explanation on Task-Level Model-Merging Collapse, read 2026-09-05. p-values as published, for pairwise merging.

TIER 2 · NOT OFFERED

Upcycling works. It costs a pretraining run.

Sparse upcycling, Branch-Train-MiX, LLaMA-MoE and OLMoE all convert dense checkpoints into genuinely trained mixtures. None of them is a merge, and none of them finishes in a browser. This page describes the path rather than pretending to offer it.

LLaMA-MoE, continued pre-training
200Btokens of continued pre-training after splitting LLaMA's FFNs into experts.source
OLMoE, upcycled versus from scratch
500Btokens at which a from-scratch MoE had caught the upcycled one; by 600B it was ahead. The authors qualify it themselves: the checkpoint they upcycled was a 1.0B-parameter model already trained on 2T tokens, and an upcycled MoE inherits hyperparameters from the dense model it came from. Read it as a caution against assuming upcycling always wins, not as a finding that it loses.source
OLMoE, total pre-training
5.133Ttokens. In that single comparison the upcycling head start was worth roughly a tenth of the budget.

The honest reading: upcycling buys time, and how much of a ceiling it buys is not settled by one ablation. What is settled is the entry price. Every method in this tier is measured in hundreds of billions of tokens, which is why this page describes it and does not offer it.

SOURCES

Catalogue figures are derived mechanically from the Hugging Face Hub API in the snapshot generated 2026-09-05. Evidence class hub_derived. How evidence is classified

12 routers evaluated in RouterArena: 3 commercial and 9 open source.