MoE COMPOSER / TIER 1 AND TIER 3
Merging is cheap. Winning is not.
Three different things get called "combining models into a mixture of experts". One takes minutes and usually loses to a model you already had. One works and costs a pretraining run. One needs no GPUs at all and is the least crowded of the three.
This page compiles a real mergekit-moe configuration against catalogue-exact revisions, reports what the metadata mechanically establishes, and refuses to score what it cannot measure. Snapshot 2026-09-05.
| Tier | Method | Cost to try | What it needs | What the published evidence says |
|---|---|---|---|---|
| 1Weight-time merge | mergekit-moe | Minutes on a CPU, or one 8 GB GPU | Dense donors of one architecture, at one scale | The canonical showcase merge scores below one of its own experts. Treat a gain as the surprise, not the default. |
| 2Upcycling | Sparse upcycling · Branch-Train-MiX · LLaMA-MoE · OLMoE | Hundreds of billions of training tokens | A pretraining cluster and a data pipeline | Works. Whether the head start lasts is unsettled: in OLMoE's own ablation a from-scratch MoE had caught up before 600B tokens, on a comparison the authors qualify themselves. |
| 3Inference-time routing | Model routing · mixture-of-agents | No GPUs. The endpoints you already have | A query set from your own traffic, and the discipline to calibrate on it | Every router in the RouterArena evaluation falls short of the oracle. That is a gap to work in, not a solved market. |
THE BASE RATE
The showcase merge loses to one of its own experts.
Beyonder-4x7B-v3 is the model the frankenMoE technique is usually demonstrated with. Its own model card publishes the Nous benchmark suite results, and AlphaMonarch-7B — one of the four 7B models merged into it — comes out ahead on the average. Nobody quotes that row. It is on the same card as the row everybody quotes.
| Model | Average | AGIEval | GPT4All | TruthfulQA | Bigbench |
|---|---|---|---|---|---|
| mlabonne/Beyonder-4x7B-v3the four-expert merge | 61.91 | 45.85 | 76.67 | 74.98 | 50.12 |
| mlabonne/AlphaMonarch-7Bone of its own four experts | 62.74 | 45.37 | 77.01 | 78.39 | 50.2 |
| Differencemerge minus best expert | -0.83 | 0.48 | -0.34 | -3.41 | -0.08 |
WHAT THE LOSS COST
- Merged artifact
- 24.2B24,153,690,112 parameters
- Checkpoint
- 48.3 GBsummed per dtype from the safetensors index
- Beaten by
- one 7B expertalready inside the merge
- Inherited licence
- cc-by-nc-4.0non-commercial, from the merge chain
Scores transcribed from the model card, read 2026-09-05. Differences computed here. Artifact figures from the catalogue record.
WHAT WOULD HAVE TO BE TRUE FOR YOURS TO DIFFER
- The experts have to disagree usefully. A merge of four models fine-tuned from the same base on overlapping data has four copies of nearly the same MLP. Routing between them buys nothing and costs memory for all of them.
- Your traffic has to be mixed. The gain, where there is one, comes from queries that genuinely belong to different experts. A single-domain workload has a single best expert, and you already own it.
- The router has to be right often enough. Gates derived from a handful of prompts are a weak signal, and a wrong route is worse than no route.
- You have to measure against the experts, not a leaderboard. Beating some unrelated model proves nothing. The only comparison that settles it is the merge against each of its own constituents, on the same questions.
TIER 1 · COMPILE IT ANYWAY
Pick the members. Read what the metadata establishes.
Every source below is pinned to the commit revision this catalogue indexed, so the merge you run is the merge this page described. The checks are mechanical: each one names the API field it read, and none of them returns a verdict.
51 of 2,601 · 3/8 experts selected
| Base | Expert | Repository | Revision | Architecture | Parameters | Checkpoint | Licence | Template | Downloads |
|---|---|---|---|---|---|---|---|---|---|
| Qwen/Qwen3-0.6B | c1899de | qwen3 | 752M | 1.5 GB | apache-2.0 | yes | 22.0M | ||
| Qwen/Qwen3-8B | b968826 | qwen3 | 8.2B | 16.4 GB | apache-2.0 | yes | 13.5M | ||
| Qwen/Qwen3-Embedding-0.6B | 97b0c61 | qwen3 | 596M | 1.2 GB | apache-2.0 | yes | 7.1M | ||
| Qwen/Qwen3-4B | 1cfa9a7 | qwen3 | 4.0B | 8.0 GB | apache-2.0 | yes | 6.3M | ||
| Qwen/Qwen3-32B | 9216db5 | qwen3 | 32.8B | 65.5 GB | apache-2.0 | yes | 5.1M | ||
| Qwen/Qwen3-4B-Instruct-2507 | cdbee75 | qwen3 | 4.0B | 8.0 GB | apache-2.0 | yes | 3.6M | ||
| Qwen/Qwen3-1.7B | 70d244c | qwen3 | 2.0B | 4.1 GB | apache-2.0 | yes | 3.5M | ||
| Qwen/Qwen3-Embedding-4B | 5cf2132 | qwen3 | 4.0B | 8.0 GB | apache-2.0 | yes | 2.8M | ||
| RadixArk/Kimi-K3-DSpark | 3c5bac3 | qwen3 | 2.2B | 4.5 GB | undeclared | unknown | 2.7M | ||
| Qwen/Qwen3-Reranker-4B | 22e6836 | qwen3 | 4.0B | 8.0 GB | apache-2.0 | yes | 2.5M | ||
| Qwen/Qwen3-Embedding-8B | 1d8ad4c | qwen3 | 7.6B | 15.1 GB | apache-2.0 | yes | 2.2M | ||
| Qwen/Qwen3-4B-Base | 906bfd4 | qwen3 | 4.0B | 8.0 GB | apache-2.0 | yes | 2.0M | ||
| Qwen/Qwen3-1.7B-Base | ea980cb | qwen3 | 1.7B | 3.4 GB | apache-2.0 | yes | 1.9M | ||
| trl-internal-testing/tiny-Qwen3ForCausalLM | 52b2e48 | qwen3 | 2M | 5 MB | undeclared | unknown | 1.8M | ||
| Qwen/Qwen3-14B | 40c0698 | qwen3 | 14.8B | 29.5 GB | apache-2.0 | yes | 1.8M | ||
| Qwen/Qwen3-Reranker-0.6B | e61197e | qwen3 | 596M | 1.2 GB | apache-2.0 | yes | 1.2M | ||
| deepseek-ai/DeepSeek-R1-0528-Qwen3-8B | 6e8885a | qwen3 | 8.2B | 16.4 GB | mit | yes | 1.0M | ||
| Qwen/Qwen3-0.6B-Base | da87bfb | qwen3 | 596M | 1.2 GB | apache-2.0 | yes | 940K | ||
| Qwen/Qwen3-14B-Base | 0b0bd37 | qwen3 | 14.8B | 29.5 GB | apache-2.0 | yes | 631K | ||
| trl-internal-testing/small-Qwen3ForCausalLM | a3f682e | qwen3 | 39M | 78 MB | undeclared | unknown | 579K | ||
| Qwen/Qwen3-8B-Base | 49e3418 | qwen3 | 8.2B | 16.4 GB | apache-2.0 | yes | 420K | ||
| Qwen/Qwen3-4B-Thinking-2507 | 768f209 | qwen3 | 4.0B | 8.0 GB | apache-2.0 | yes | 372K | ||
| incoai/Qwen3.8-27B-DFlash2 | dedf8df | qwen3 | 1.9B | 3.8 GB | apache-2.0 | unknown | 284K | ||
| voyageai/voyage-4-nano | 67fabc9 | qwen3 | 346M | 693 MB | apache-2.0 | yes | 280K | ||
| lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-8bit | e524281 | qwen3 | 8.2B | — | mit | yes | 273K | ||
| RadixArk/Qwen3.8-27B-DSpark | b9a5dbd | qwen3 | 1.9B | 3.7 GB | other | unknown | 268K | ||
| z-lab/Qwen3.8-27B-DFlash2 | 50307d4 | qwen3 | 1.9B | 3.8 GB | apache-2.0 | unknown | 242K | ||
| trl-internal-testing/tiny-Qwen3ForCausalLM-Instruct-2507 | 6f2ac43 | qwen3 | 2M | 5 MB | undeclared | unknown | 234K | ||
| z-lab/Qwen3.6-35B-A3B-DFlash | f181eec | qwen3 | 386M | 772 MB | apache-2.0 | unknown | 228K | ||
| Qwen/Qwen3Guard-Gen-0.6B | fada3b2 | qwen3 | 752M | 1.5 GB | apache-2.0 | yes | 160K | ||
| typhoon-ai/typhoon2.5-qwen3-4b | ce0a741 | qwen3 | 4.0B | 8.0 GB | apache-2.0 | unknown | 157K | ||
| unsloth/Qwen3-Embedding-4B | 8edd3dd | qwen3 | 4.0B | 8.0 GB | apache-2.0 | yes | 156K | ||
| z-lab/Qwen3.6-27B-DFlash | 0919688 | qwen3 | 1.7B | 3.5 GB | mit | unknown | 147K | ||
| llamafactory/tiny-random-qwen3 | 81d6f5f | qwen3 | 2M | 5 MB | apache-2.0 | yes | 122K | ||
| Qwen/Qwen3-Reranker-8B | 77d193c | qwen3 | 8.2B | 16.4 GB | apache-2.0 | yes | 121K | ||
| Qwen/Qwen3Guard-Gen-4B | 6ec4282 | qwen3 | 4.4B | 8.8 GB | apache-2.0 | yes | 106K | ||
| mlx-community/Qwen3-0.6B-8bit | 11de968 | qwen3 | 596M | — | apache-2.0 | yes | 102K | ||
| OpenOneRec/OneRec-1.7B | cc4d0b5 | qwen3 | 2.1B | 4.3 GB | apache-2.0 | yes | 96K | ||
| Goedel-LM/Goedel-Prover-V2-32B | 851bf85 | qwen3 | 32.8B | 65.5 GB | apache-2.0 | yes | 75K | ||
| lmstudio-community/Qwen3-14B-MLX-8bit | 86a7e1c | qwen3 | 14.8B | — | apache-2.0 | yes | 67K |
Gate
hidden — Derives the gate from the hidden-state representations of the positive and negative prompts. The documented default and the best of the three. Runs every prompt through the base model. --load-in-8bit or --load-in-4bit reduces the VRAM this needs.
Routing prompts
The gate is derived from these. One line per prompt. Under gate_mode: hidden, every expert needs at least one positive prompt or mergekit has nothing to build a gate from.
Mechanical findings
Each row names the API field it read. There is no overall score and no "compatible" verdict, because the only published predictor of merge collapse is representational distance, which no metadata field carries.
| Severity | Check | Finding | Read from |
|---|---|---|---|
| Warning | parameter-scale-spread | Parameter counts span 14× across the members, from 595,776,512 to 8,190,735,360. mergekit-moe requires the shared tensors to have identical shapes, and a differing total is evidence they do not. | safetensors.total |
| Warning | licence-chain | Every member declares apache-2.0, so the merged artifact inherits it. mergekit attaches no licence to its output; you have to declare one, and it cannot be looser than this. | cardData.license on every member |
Nothing here blocks the merge. Every one of these findings still costs you something later. Licence: every member declares apache-2.0, and the output inherits it.
Merged size
Largest single member's parameter count written at bfloat16 at 2 bytes per parameter. The merged artifact holds every member's FFN stack, so it cannot be smaller than the biggest one.
Every member's parameter count summed and written at bfloat16 at 2 bytes per parameter. This is the size if no weight were shared, which over-counts: the shared tensors are stored once.
An exact figure is shared_weights + experts x ffn_weights + router. Splitting a checkpoint into those two parts requires the layer dimensions in the repository's full config.json, which the Hub's config summary omits. This project will not publish a point estimate it cannot derive.
To compute it exactly this page would need, from each member's full config.json: config.json → hidden_size · config.json → intermediate_size · config.json → num_hidden_layers · config.json → vocab_size. The Hub's config summary does not carry them, so no point estimate is offered.
mergekit-moe configuration
# mergekit-moe configuration generated by MOEModels.ai
# Sources are pinned to the commit revisions indexed in this catalogue snapshot.
base_model: Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
gate_mode: hidden
dtype: bfloat16
experts_per_token: 2
experts:
- source_model: Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
# positive_prompts: required by gate_mode "hidden" and not yet written for this expert
- source_model: Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218
# positive_prompts: required by gate_mode "hidden" and not yet written for this expert
- source_model: Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3
# positive_prompts: required by gate_mode "hidden" and not yet written for this expert
Command
mergekit-moe config.yaml ./out --copy-tokenizer- Qwen/Qwen3-0.6B has no positive prompts. gate_mode "hidden" derives the router from them, so mergekit cannot build a gate without at least one.
- Qwen/Qwen3-8B has no positive prompts. gate_mode "hidden" derives the router from them, so mergekit cannot build a gate without at least one.
- Qwen/Qwen3-Embedding-0.6B has no positive prompts. gate_mode "hidden" derives the router from them, so mergekit cannot build a gate without at least one.
- gate_mode hidden runs every prompt through the base model. Add --load-in-8bit or --load-in-4bit if the base does not fit in the VRAM you have.
Sources are pinned with mergekit's path@revision syntax to the commits this snapshot indexed on 2026-09-05, so the merge is reproducible against the exact artifacts described above rather than against whatever the default branches hold later. mergekit-moe documentation ↗
The evaluation that would settle it
Not "how does the merge score". Whether it beats each of the models already inside it, on the same questions, on one pinned harness.
| Comparison | Reference, pinned | Question it answers |
|---|---|---|
| merged-vs-Qwen/Qwen3-0.6B | Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca | Does the merged artifact beat Qwen3-0.6B by at least 3 points, on the same questions? |
| merged-vs-Qwen/Qwen3-8B | Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218 | Does the merged artifact beat Qwen3-8B by at least 3 points, on the same questions? |
| merged-vs-Qwen/Qwen3-Embedding-0.6B | Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3 | Does the merged artifact beat Qwen3-Embedding-0.6B by at least 3 points, on the same questions? |
n = (z_{alpha/2} + z_{beta})^2 · 2p(1-p)(1-rho) / delta^2zα/2 = 1.959964 · zβ = 0.841621 · paired variance 0.2500. One deterministic answer per question, per arm. Resampling answers reduces the count; an unpaired comparison, or a harness that does not record per-question outcomes, raises it sharply.
| Reference benchmark | Questions | Smallest resolvable difference | Resolves 3 points? |
|---|---|---|---|
| GPQA Diamondsource | 198 | 9.95 pts | no |
| GPQA (main)source | 448 | 6.62 pts | no |
| GPQA Extendedsource | 546 | 5.99 pts | no |
| Recommended eval floorsource | 1,000 | 4.43 pts | no |
- Pin the harnessOne harness at one commit, one prompt template, one decoding configuration, for every arm. A merged model and its experts evaluated on different harness versions are not comparable, and the difference between harness versions is routinely larger than the effect being chased.
- Run the arms on identical questions4 arms — the merged artifact and 3 references — over the same question set, in the same order, at temperature 0.
- Record per-question outcomesStore the score for every question and every arm, not just the aggregate. The paired analysis that makes this affordable is impossible to run afterwards from summary numbers alone.
- Report the paired differencePublish the difference, its standard error and the per-question correlation between arms — not two averages side by side. Two averages with no error bar cannot distinguish a real gain from sampling noise.
- Compare against the base rateThe claim is only interesting if the merged artifact beats its best expert. mlabonne/Beyonder-4x7B-v3 did not: 61.91 against 62.74 for mlabonne/AlphaMonarch-7B, one of its own four experts.
- Detecting a 3-point difference at 80% power and alpha 0.05 needs 2,181 questions under these assumptions.
- GPQA Diamond has 198 questions. At that size the smallest difference the comparison can resolve is about 9.9 points, so a 3-point result on it is structurally unresolvable — no amount of careful running fixes a sample size.
- These are sample sizes, not permission. A benchmark large enough to resolve the effect still says nothing about whether the effect transfers to the work you actually do.
Tier 3 — the same models, composed at inference time
No merge, no GPUs, no derivative artifact and therefore no licence chain on an output. The weights stay where they are; a policy decides which endpoint answers. This is the only tier of the three that a browser can carry end to end, and it is measurably under-served.
| Order | Endpoint, pinned | Role | Parameters | Ordered by |
|---|---|---|---|---|
| 1 | Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3 | first-attempt | 596M | safetensors.total, ascending |
| 2 | Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca | escalation | 752M | safetensors.total, ascending |
| 3 | Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218 | escalation | 8.2B | safetensors.total, ascending |
# Inference-time routing policy generated by MOEModels.ai
# A policy document, not a vendor configuration. No router product reads this format.
# Weights are never modified: every candidate is served from its own endpoint.
kind: routing_policy
version: 1
escalation:
- order: 1
model: Qwen/Qwen3-Embedding-0.6B@97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3
role: first-attempt
parameters: 595776512
accept_when: "TODO: calibrate — see calibration below"
- order: 2
model: Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
role: escalation
parameters: 751632384
accept_when: "TODO: calibrate — see calibration below"
- order: 3
model: Qwen/Qwen3-8B@b968826d9c46dd6066d109eabc6255188de91218
role: escalation
parameters: 8190735360
accept_when: "TODO: calibrate — see calibration below"
- Collect a query set from your own traffic. A router calibrated on someone else's distribution is a router calibrated on someone else's problem.
- Score every candidate on every query. The escalation order above is by size; whether size predicts quality on your queries is a measurement, not an assumption.
- Compute the oracle: the accuracy you would reach by always picking the best candidate per query. That is the ceiling your policy is competing against, and no published router reaches it.
- Set the accept threshold from the measured accuracy-versus-cost curve, then re-measure. A threshold copied from a blog post is a guess wearing a number.
All twelve evaluated routers fall short of the oracle, chiefly by failing to notice when a smaller, cheaper model would have been sufficient. Of the 12 evaluated, the commercial entrant NotDiamond placed 12th on the paper's composite score, which prices accuracy against cost: it reaches 68% accuracy at $9.34 per thousand queries. Routing is the tier a browser can honestly ship end to end; it is not a solved problem, and buying one does not make it solved. RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers ↗
WHY THERE IS NO COMPATIBILITY SCORE
The predictor that works is not in the metadata.
A "will this merge work?" score would be the most clickable thing on this page. It would also be invented. The published work on task-level merging collapse tested two candidate predictors across five merging methods.
| Merging method | Hidden-state distance | Parameter sign conflict |
|---|---|---|
| TIES | p = 0.006 | p = 0.379 |
| LA | p = 0.001 | p = 0.460 |
| DARE | p = 0.145 | p = 0.882 |
| SLERP | p = 0.001 | p = 0.408 |
| TA | p = 0.006 | p = 0.659 |
Representational distance between the members predicts collapse in four of the five methods. Parameter-space sign conflict predicts it in none of them. Every field this catalogue holds — architecture, parameter count, licence, dtype, expert topology — lives on the parameter side of that result. The one signal that works requires running both models over a shared input set and comparing hidden states, which is a measurement, not a lookup.
So the composer reports findings and stops. A number here would be a confident-looking function of the predictor that was tested and failed.
An Empirical Study and Theoretical Explanation on Task-Level Model-Merging Collapse, read 2026-09-05. p-values as published, for pairwise merging.
TIER 2 · NOT OFFERED
Upcycling works. It costs a pretraining run.
Sparse upcycling, Branch-Train-MiX, LLaMA-MoE and OLMoE all convert dense checkpoints into genuinely trained mixtures. None of them is a merge, and none of them finishes in a browser. This page describes the path rather than pretending to offer it.
- LLaMA-MoE, continued pre-training
- 200Btokens of continued pre-training after splitting LLaMA's FFNs into experts.source
- OLMoE, upcycled versus from scratch
- 500Btokens at which a from-scratch MoE had caught the upcycled one; by 600B it was ahead. The authors qualify it themselves: the checkpoint they upcycled was a 1.0B-parameter model already trained on 2T tokens, and an upcycled MoE inherits hyperparameters from the dense model it came from. Read it as a caution against assuming upcycling always wins, not as a finding that it loses.source
- OLMoE, total pre-training
- 5.133Ttokens. In that single comparison the upcycling head start was worth roughly a tenth of the budget.
The honest reading: upcycling buys time, and how much of a ceiling it buys is not settled by one ablation. What is settled is the entry price. Every method in this tier is measured in hundreds of billions of tokens, which is why this page describes it and does not offer it.
SOURCES
- arcee-ai/mergekitread 2026-09-05
- mergekit-moe documentationread 2026-09-05
- mlabonne/Beyonder-4x7B-v3 model cardread 2026-09-05
- OLMoE: Open Mixture-of-Experts Language Modelsread 2026-09-05
- LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-trainingread 2026-09-05
- RouterArena: An Open Platform for Comprehensive Comparison of LLM Routersread 2026-09-05
- An Empirical Study and Theoretical Explanation on Task-Level Model-Merging Collapseread 2026-09-05
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluationsread 2026-09-05
- GPQA: A Graduate-Level Google-Proof Q&A Benchmarkread 2026-09-05
Catalogue figures are derived mechanically from the Hugging Face Hub API in the snapshot generated 2026-09-05. Evidence class hub_derived. How evidence is classified
12 routers evaluated in RouterArena: 3 commercial and 9 open source.