OPEN DATA / CC BY 4.0

Take the data with you.

Three datasets, published as files. You may redistribute them and reproduce them in structured, tabular and machine-readable form, commercially, provided you attribute them. That grant is deliberate: an index nobody may republish is a destination, not infrastructure.

Snapshot 2026-09-05 · schema 1.0.0 · evidence class hub_derived · derived from Hugging Face Hub metadata, never measured here.

3PUBLISHED DATASETS
2,601CATALOGUE RECORDS
739PUBLISHERS
2,318EXACT PARAMETER COUNTS
CC BY 4.0LICENCE

WHAT IS PUBLISHED

Three datasets, three evidence classes.

The classes are never merged. A mechanically derived value, a value a human read in a primary document, and a value this project measured are different kinds of claim, and a table that mixes them is a table that cannot be checked.

How evidence is classified

Hub-derived mixture-of-experts catalogue

EVIDENCE CLASS hub_derived

RECORDS
2,601 repositories from 739 publishers
VERSION
1.0.0
UPDATED
2026-09-05

Artifact-exact records for expert-routed model repositories: commit SHA, architecture, routed experts and experts per token, total parameters, checkpoint bytes summed per dtype, quantisation method, declared licence and access state. Each field in the full record names the API field it was computed from; each field the API does not establish is undetermined with a reason.

Coverage. Expert-routed repositories found by walking the Hub’s text-generation listing in descending download order.

Cadence. A scheduled ingest runs weekly and opens a pull request with the new snapshot. The published file changes when that pull request is merged, so read generatedAt rather than assuming the data is a week old at most.

Class. Computed mechanically from Hugging Face Hub API responses. Not measured by this project and not human-reviewed.

DownloadFormatContents
Catalogue JSON/api/v1/catalogapplication/jsonFlat projection of every record, with the licence, citation and null semantics in the envelope.
Catalogue CSV/api/v1/catalog?format=csvtext/csvThe same projection as a single file, one row per repository. This is the bulk download.
Single record with provenance/api/v1/catalog/Qwen%2FQwen3-0.6Bapplication/jsonOne repository — Qwen/Qwen3-0.6B in this snapshot — with every field's derivedFrom path, or the reason it is undetermined.
JSON Schema/schemas/hub-catalog-v1.jsonapplication/schema+jsonThe manifest contract, including the derived/undetermined value wrapper.

Reviewed model and hardware registry

EVIDENCE CLASS sourced

RECORDS
5 models, 4 accelerators, 20 compatibility records, 21 sources
VERSION
1.0.0
UPDATED
2026-08-02

Exact-artifact model records with pinned repository and revision, accelerator specifications, compatibility statements and the methodology constants the fit calculation uses. Every fact carries the source it came from; unknowns are recorded as unknown rather than filled in.

Coverage. Models and accelerators this project has reviewed in depth.

Cadence. Amended when a reviewer records a new sourced fact.

Class. A human read a primary document — model card, technical report, vendor specification — and recorded the claim against it.

DownloadFormatContents
Registry JSON/api/v1/registryapplication/jsonThe whole registry, exactly as the CLI and SDK read it.
OpenAPI 3.1 document/api/v1/openapiapplication/jsonThe machine-readable contract for every public endpoint, including the catalogue.

Evaluation evidence registry

EVIDENCE CLASS sourced

RECORDS
15 owner-reported claims, 0 normalised runs, 9 sources
VERSION
1.0.0
UPDATED
2026-08-03

Benchmark claims published by model owners, each bound to its source and to whether the claim names an exact artifact snapshot or only a model name, alongside the run and adapter contract a measured run must satisfy before publication.

Coverage. Claims and runs for models the registry covers.

Cadence. Amended when a claim is recorded or a run is normalised.

Class. Owner-reported claims are quoted with their source. No run executed by this project is published in this snapshot. When one is, it is kept in a separate class and becomes comparison-eligible only after its artifact, runtime, hardware and raw evidence pass review.

DownloadFormatContents
Evaluations JSON/api/v1/evaluationsapplication/jsonClaims, runs, adapters and sources, kept in separate evidence classes.
JSON Schema/schemas/evaluations-v1.jsonapplication/schema+jsonThe evaluation contract, including the artifact-association rule.

QUERYING IT

Four parameters, no key.

The catalogue endpoint sends Access-Control-Allow-Origin: *, needs no token, and carries its licence, attribution and citation in the response envelope, so a fetched copy stays traceable to its source.

ParameterValuesEffect
formatjson (default), csvcsv returns the same projection as one file, with a Content-Disposition filename. This is the bulk download.
moetrue, falseFilters on whether the published config declares expert routing.
archa config.model_type value, e.g. deepseek_v3Exact match on the architecture family the publisher declared.
limit1 to 10000First n records after filtering. Records are ordered by the Hub download count, which orders a listing and nothing else.
# the whole catalogue as one file
curl -L "https://moemodels.ai/api/v1/catalog?format=csv" -o moemodels-hub-catalog-2026-09-05.csv

# expert-routed repositories of one architecture family
curl -L "https://moemodels.ai/api/v1/catalog?moe=true&arch=deepseek_v3&limit=50"

# one record with every field's provenance intact
curl -L "https://moemodels.ai/api/v1/catalog/Qwen%2FQwen3-0.6B"

A null in the JSON and an empty cell in the CSV both mean the Hub did not establish that value for that repository. Neither means zero. The per-record endpoint carries the reason.

LICENCE / CC-BY-4.0

Redistribution is permitted. That is the point.

MOEModels’ own derived data — the catalogue, the reviewed registry and the evaluation records — is published under the Creative Commons Attribution 4.0 International licence.

What you may do

Copy the raw files. Redistribute them. Reproduce the data in structured, tabular and machine-readable form. Build a product on it, put it in a paper, load it into a model’s context, or use it to build an index that competes with this one. Commercial use is included. The single condition is attribution.

REQUIRED ATTRIBUTION

MOEModels.ai (LockedIn Labs), Hub-derived mixture-of-experts catalogue, snapshot 2026-09-05, CC BY 4.0.

What the grant does not cover

Model weights
Every repository in the catalogue carries its own licence, recorded in the licence field of its record and read from the publisher’s own card metadata. Some are permissive, some are bespoke vendor terms, some declare nothing at all. Nothing granted here touches them, and an absent licence is not a grant.
Hugging Face’s underlying metadata
The catalogue is derived from the Hub API. The grant covers MOEModels’ derivation — the selection, the per-field provenance, the per-dtype byte arithmetic and the record structure — not the upstream service’s own terms, which continue to govern the API responses the derivation reads.
Third-party material a record quotes or links to
Evaluation records reproduce owner-reported claims with a link to their source. The compiled record is ours to license. The model card, technical report or vendor page it points at is not.
Software and brand
The open protocol, CLI, SDK, MCP server and schemas are Apache-2.0 in the public repository. The MOEModels and LockedIn Labs names and marks are not licensed by the data grant.

CITATION

Cite the snapshot, not the site.

The catalogue is regenerated. A citation that names only the domain will not resolve to the numbers you read, so cite the snapshot date and the schema version — both are carried in every API response as generatedAt and schemaVersion.

CITE AS

MOEModels.ai (LockedIn Labs). Hub-derived mixture-of-experts model catalogue, snapshot 2026-09-05. Derived from Hugging Face Hub metadata. Evidence class hub_derived. CC BY 4.0. https://moemodels.ai/data

BIBTEX

@dataset{moemodels_hub_catalogue_2026,
  title     = {Hub-derived mixture-of-experts model catalogue},
  author    = {{MOEModels.ai} and {LockedIn Labs}},
  year      = {2026},
  month     = {sep},
  version   = {1.0.0},
  publisher = {LockedIn Labs},
  license   = {CC BY 4.0},
  url       = {https://moemodels.ai/data},
  note      = {Snapshot 2026-09-05, 2601 repositories. Derived from Hugging Face Hub metadata; evidence class hub_derived, not measured.},
  urldate   = {2026-09-05}
}

WHY IT IS LICENSED THIS WAY

A number you may not republish is not evidence.

Model-comparison data is often distributed under terms that forbid redistributing the raw files and forbid reproducing the data in structured, tabular or machine-readable form. Terms like those settle what the data can become. A researcher cannot put the table in a paper. A team cannot check a vendor claim against it in CI. An assistant answering a question cannot reproduce the figure and name where it came from, which means the figure circulates without its source or does not circulate at all.

This project takes the opposite position, for a self-interested reason as much as a principled one. The catalogue is only useful if it travels. Work that carries a figure and says where it came from is the objective, and attribution is a cheaper price for that than the enforcement it replaces.

Attribution is the only condition. Read the caveats below before drawing a conclusion the data does not support — most of the ways to misread this dataset are listed there.

CAVEATS

Read these before you publish a chart.

Each of these is a way the data is routinely misread. They are published here rather than in a footnote because a caveat nobody sees is a caveat nobody applies.

01Download counts are not a popularity measure
Hugging Face documents that it counts downloads server-side by watching a set of query files, and that every HTTP request to one of those files — GET and HEAD alike — is counted. It also documents that GGUF files are all counted individually, which double-counts a user who clones a whole repository. The field is published here because it is what the Hub reports, and it is useful for ordering a listing. Ranking models by it, or reading it as adoption, publishes a number most readers will misinterpret.
02A gated repository still publishes its metadata
Repositories whose access is gated_automatic or gated_manual expose configuration, safetensors index and card metadata to anyone, which is why their sizes and expert counts appear here in full. The weights themselves are not served without an accepted licence and an authenticated request. Treat presence in the catalogue as evidence about the artifact, never as evidence that you can obtain it.
03A snapshot is a point in time
Every record names the commit SHA it describes. A repository that is re-uploaded, requantised, relicensed or withdrawn after 2026-09-05 will disagree with this file, and the file is the one that is out of date. A weekly ingest proposes a new snapshot and a person merges it, so the published file is as old as the last merge, not as old as the last run; read generatedAt rather than assuming freshness.
04Coverage is download-ranked, not exhaustive
The ingestion walks the Hub’s text-generation listing in descending download order and keeps the expert-routed repositories it finds. It is not a census. A newly published or rarely downloaded MoE repository can be absent, and absence from this file says nothing about the model.
05Derived is not measured
Parameter counts and checkpoint bytes are computed from the publisher’s own safetensors index, so they are exact about the artifact and say nothing about behaviour. There is no quality score, no throughput, no latency, no cost and no active-parameter count in this dataset, because none of those can be derived from a repository manifest.
06Upstream values are preserved, not corrected
The Hub occasionally publishes a value that cannot be true — as of this snapshot, one repository reports a negative total storage. The snapshot keeps what the API returned, with the field it came from, so the defect stays visible and auditable. The catalogue API refuses to republish an impossible size: a negative repository size is reported as null rather than charted as a negative number of bytes.
07A null is not a zero
Where the Hub API does not establish a value, the field is null in the JSON projection and empty in the CSV, and the full record carries the reason. Filling those with a plausible default would make the file easier to chart and impossible to trust.

Hugging Face documents its download counting at https://huggingface.co/docs/hub/models-download-stats.

PROVENANCE

The catalogue is derived from https://huggingface.co/api/models in the snapshot generated 2026-09-05T21:15:59.232Z, under Hugging Face’s terms of service. Evidence class hub_derived: computed mechanically from publisher metadata, never measured by this project. Corrections are treated as defects — open an issue with the exact source locator.