Moonshot AI
Kimi K3
A native multimodal 2.8T-class model for long-horizon coding, knowledge work, and million-token contexts.
THE VERIFIABLE MODEL INDEX
An artifact-exact catalogue of open models, an evidence standard that travels with the result, and the tools to derive, size and measure your own. Every figure names where it came from — and every figure we cannot establish says so.
OPEN DATA CC BY 4.0 · bulk JSON and CSV · redistributable under attribution
Independent model intelligence from LockedIn Labs
Qwen3-30B-A3B · sourced registry facts + calculated static fit · topology illustration, not measured runtime data
ONE SYSTEM / MULTIPLE ENTRY POINTS
PLATFORM V0.1
The website is the visual workbench. The protocol, runner and developer surfaces make the same decision system portable.
Every expert-routed repository at a pinned revision, with parameter counts and checkpoint bytes summed from the actual tensor index.
Configure a post-training run and take away a pinned, content-addressed recipe you execute on your own compute — plus the receipt it should produce.
Author a merge, see the compatibility findings and the licence chain, and get the evaluation plan that would actually prove it helped.
Pack compatible trials under one digest, verify summaries and signatures locally, and keep every reproducibility gap attached.
How to read a benchmark, how many questions it takes to tell two models apart, and why the same model scores differently in two harnesses.
Bulk JSON and CSV, a stable API, and a citation block. Redistribute and reproduce it in machine-readable form under attribution.
THE DEPLOYMENT GAP
01 / 05
Sparse activation reduces compute. It does not erase model weights, KV cache, runtime buffers, replication, or all-to-all traffic.
MOEModels separates can load, can run, can scale, and makes economic sense— because they are four different answers.
Every specification names the field or document it came from.
Every derived value exposes its exact inputs and formula.
Every value the source does not establish is shown as unknown, with the reason.
THE ARTIFACT LAYER
02 / 05
The same model name spans a BF16 release, an FP8 release, a community requantisation and a fine-tune — different memory, different behaviour, different licence. Every record here is addressed by its commit.
Open the catalogueMoonshot AI
A native multimodal 2.8T-class model for long-horizon coding, knowledge work, and million-token contexts.
DeepSeek
A frontier-scale model combining fine-grained sparse experts with a million-token context for agentic work.
Z.ai
An open flagship designed for long engineering trajectories across a sourced million-token context.
A compact multimodal MoE with 25.2B total and 3.8B active parameters, designed for accessible inference.
The four records above are the human-reviewed subset, sourced from pinned manifests, configurations, model cards and technical releases. The catalogue behind them is derived mechanically from publisher metadata and is labelled as such. Neither is a benchmark ranking.
MOE FIT CHECK
03 / 05
Select a pinned checkpoint, hardware target, topology, and declared reserve. See what is proven—and what remains unknown.
Exact manifest bytes + integer math. No generated throughput, capacity, or price.
Deterministic registry engine
Test an exact, pinned artifact against advertised accelerator memory—then keep every unsupported conclusion visibly unknown.
02 · Fit decision
Kimi K3 · 896 / top-16 experts · 1M context · 8 × NVIDIA H200 SXM · 13% reserve
The checkpoint alone requires at least 13 GPUs at this reserve, before runtime allocations.
A baseline pass does not prove loader, kernel, quantization, sharding, or expert-parallel support.
No measured workload profile is attached, so throughput, latency, KV demand, and skew are not projected.
No dated provider, region, or utilization record is selected, so the engine emits no invented cost.
03 · Explainable topology
RequestedThe topology you asked the engine to test against the checkpoint baseline.
usable = floor(advertised bytes × (10,000 − reserve bps) / 10,000); minimum GPUs = ceil(checkpoint tensor bytes / usable bytes). A failure is conclusive for this no-offload baseline; a pass is only a candidate for runtime validation.
Artifact evidence: Kimi K3 pinned tensor manifest ↗
BENCHMARK EVIDENCE
04 / 05
Evidence is tiered, and every number carries its tier. Mechanically derived metadata covers thousands of artifacts. Reviewed claims keep scope, settings and missing context attached. No controlled run has crossed the comparison gate yet, and until one does this page says so.
Enter the evidence labOPEN INFRASTRUCTURE
05 / 05
Workbench, CLI, runner, REST API, TypeScript SDK and MCP server share the registry, evidence classes and deterministic planning contract.
$ npm run moemodels -- fit \
moonshotai/Kimi-K3 \
nvidia/h200-sxm-141gb \
--gpus 8 --reserve-pct 13
Real offline CLI · same registry and integer engine as the hosted Fit Check
RESEARCH DESK
FIELD GUIDE
A practical guide to resident weights, KV cache, runtime overhead, and the four different meanings of “fits.”
SYSTEMS NOTE
How topology, all-to-all traffic, load imbalance, and hot experts reshape the economics of sparse inference.
BUYER BRIEF
A decision framework for comparing utilization, privacy, operational burden, and three-year total cost.
START WITH THE PROOF
Verify a deterministic Passport fixture, change exactly one byte, and see the content address reject it—all locally in your browser. This is the standard every measured number on this site has to meet.
Run the 90-second check