BENCHMARK EVIDENCE / V1

Scores with receipts.

First-party results are useful signals—not a leaderboard. Every score keeps its model scope, reported settings, source, and missing run context attached.

Comparable means same artifact and methodology fingerprint. Everything else remains evidence, not rank.

15OWNER-REPORTED CLAIMS
5EXACT MODEL RECORDS
0NORMALIZED RUNS
0STRICTLY COMPARABLE

REPORTED EVIDENCE STREAM

Context stays attached.

These values come from official model cards and technical reports. None is represented as independently reproduced, cryptographically bound to the evaluated weights, or safe for cross-model ranking.

Open the evidence API
15OF 15 CLAIMS
Pinned-card associatedOWNER REPORTED

Kimi K3

GPQA

Diamond

93.5
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=0.95
4 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • Shot count, sample count, prompt, and benchmark revision are not reported.
  • Inference engine, runtime configuration, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Kimi K3

Terminal-Bench

2.1

88.3
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=1harness=Kimi Code
5 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • The Kimi Code harness version and full harness configuration are not reported.
  • Task revision, trial count, and aggregation method are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Kimi K3

BrowseComp

Version not reported

91.2
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=1context_compaction_trigger_tokens=300000
5 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • The browsing tool implementation, harness version, and context-compaction algorithm are not reported.
  • Benchmark revision, trial count, and aggregation method are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

DeepSeek V4 Pro

GPQA

Diamond

90.1
Pass@1percentage points0-100think_max
reasoning_mode=Think Maxtemperature=1context_tokens=384000system_prompt_variant=Think Max special system prompt
4 context gaps attached
  • Top-p, exact prompt text for GPQA, shot count, and sample count are not reported.
  • Benchmark revision and evaluation harness version are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

DeepSeek V4 Pro

LiveCodeBench

v6

93.5
Pass@1-COTpercentage points0-100think_max
reasoning_mode=Think Maxsystem_prompt_variant=Think Max special system prompt
4 context gaps attached
  • Prompt, top-p, sample count, and maximum output length are not reported for this benchmark.
  • LiveCodeBench task revision and harness version are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

DeepSeek V4 Pro

SWE-bench

Verified

80.6
Resolvedpercentage points0-100think_max
reasoning_mode=Think Maxharness=DeepSeek internal evaluation frameworktooling=bash and file-edit toolsmax_interaction_steps=500context_tokens=512000
4 context gaps attached
  • The internal framework version and complete tool prompts are not published.
  • SWE-bench dataset commit, container images, trial count, and aggregation method are not reported.
  • Sandbox resources and inference hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

GLM-5.2

GPQA

Diamond

91.2
owner-reported scorereported score0-100 as presentedunspecified
temperature=1top_p=0.95max_new_tokens=163840
4 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • Reasoning effort, prompt, shot count, sample count, and judge configuration are not reported for GPQA.
  • Benchmark revision, harness version, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

GLM-5.2

SWE-bench Pro

Version not reported

62.1
owner-reported scorereported score0-100 as presentedunspecified
framework=OpenHandsprompt_variant=tailored instruction prompttemperature=1top_p=1max_new_tokens=32000context_tokens=400000
5 context gaps attached
  • The owner does not explicitly define the metric or label the score as a percentage.
  • OpenHands version and the tailored instruction prompt are not published.
  • Dataset commit, container images, task count, trials, and aggregation method are not reported.
  • Inference hardware is not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

GLM-5.2

MCP-Atlas

500-task public set

76.8
owner-reported scorereported score0-100 as presentedthinking
thinking_mode=truetask_count=500timeout_seconds=600judge_model=Gemini-3.0-Pro
4 context gaps attached
  • The owner does not explicitly define the metric or label the score as a percentage.
  • Exact reasoning effort, temperature, context limit, harness version, and tool configuration are not reported.
  • Sample count per task, aggregation method, and inference hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Gemma 4 26B A4B IT

AIME

2026 · no tools

88.3
accuracypercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=truetool_use=false
4 context gaps attached
  • Prompt, sampling parameters, sample count, and scoring implementation are not reported.
  • Benchmark task revision and evaluation harness are not reported.
  • Context limit, maximum output length, runtime, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Gemma 4 26B A4B IT

LiveCodeBench

v6

77.1
owner-reported percentagepercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=true
4 context gaps attached
  • The exact LiveCodeBench metric, prompt, sampling parameters, and sample count are not reported.
  • Benchmark task revision and evaluation harness are not reported.
  • Context limit, maximum output length, runtime, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Gemma 4 26B A4B IT

GPQA

Diamond

82.3
accuracypercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=true
4 context gaps attached
  • Prompt, shot count, sampling parameters, sample count, and scoring implementation are not reported.
  • Benchmark task revision and evaluation harness are not reported.
  • Context limit, maximum output length, runtime, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED

Qwen3-30B-A3B

GPQA

Diamond

65.8
averaged accuracypercentage points0-100thinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=32768samples_per_query=10
4 context gaps attached
  • The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
  • Exact prompt, benchmark revision, and evaluation harness are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED

Qwen3-30B-A3B

AIME

2025 · Parts I and II

70.9
averaged accuracypercentage points0-100thinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=38912question_count=30samples_per_query=64
4 context gaps attached
  • The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
  • Exact prompt text, scoring parser, and test-data revision are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED

Qwen3-30B-A3B

LiveCodeBench

v5 · 2024-10 through 2025-02

62.6
owner-reported scorereported score0-100 as presentedthinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=32768benchmark_window=2024-10 through 2025-02prompt_variant=official prompt with the program-only restriction removed
4 context gaps attached
  • The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
  • The report does not explicitly define the table's LiveCodeBench metric or sample count.
  • Exact task commit, harness version, inference engine, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.

THE RESOLUTION QUESTION

Most reported gaps are smaller than their own error bars.

Scoring a benchmark is a sequence of Bernoulli trials, so a reported accuracy carries a standard error whether or not anyone prints it. On a small instrument that error is larger than the gaps leaderboards rank by.

Closed-form arithmetic over the values you enter. No model scores are supplied, assumed or stored.

NOT RESOLVABLE AT THIS SIZE

A 3.00 point gap on GPQA Diamond carries a 95% margin of ±9.84 points. The interval on the difference runs from -6.84 to 12.84 points, which contains zero. The two scores are indistinguishable on this instrument.

Interval and sample-size results
Observed difference3.00 ppModel A minus Model B
Standard error, unpaired5.02 pp√(SE² A + SE² B), treating the runs as independent samples
Standard error, paired5.02 pp√(SE² A + SE² B − 2ρ·SE A·SE B) at ρ = 0.00
Design effect1.00×1 + (m − 1)·ICC at m = 1, ICC = 0.00
Effective questions198198 nominal ÷ design effect
Smallest resolvable gap9.84 ppAnything below this cannot be distinguished from noise here
Questions needed, unpaired4,361(zα/2 + zβ)² · 2p(1−p) ÷ δ² at 80% power
Questions needed, paired4,361Same, with variance reduced by (1 − ρ)
Questions needed, clustered4,361Paired requirement × design effect

GPQA Diamond is short by 4,163 questions for this comparison. Reporting the ordering anyway is reporting noise.

The hardest GPQA subset, and the one leaderboards quote most often. Source ↗

ρ is the correlation between the two models’ per-question outcomes. It is left at zero by default because it can only be measured from per-question outputs, which almost no leaderboard publishes; raising it without that measurement produces an interval narrower than the evidence supports. Clustering applies when questions arrive in groups — several per passage, per repository, per record — because those questions are not independent draws.

The default view is the case leaderboards publish constantly: a three-point gap on GPQA Diamond, which has 198 questions. The 95% margin on that difference is roughly ten points. The gap would have to more than triple before the benchmark could tell the two models apart, and no site reporting that ordering says so.

Two levers make comparisons cheaper rather than louder. Pairing — scoring both models on the same questions and comparing per-question differences — removes the variance the models share, and it is why a comparison that needs several thousand questions when treated as two independent means can need roughly a thousand when paired. Clustering pushes the other way: when questions arrive in groups, the effective sample is smaller than the question count.

Method after Miller, “Adding Error Bars to Evals” (Anthropic, 2024): arxiv.org/abs/2411.00640 ↗

DETERMINISTIC INGESTION

Turn raw output into a reviewable staging pack.

The v0.1 importer hashes aggregate bytes, preserves metric/filter pairs and explicit unknowns, and never executes model or task code. Staging output is not publication eligibility.

MOEMODELS / EVIDENCE IMPORT
$ moemodels ingest lm-eval results.json \
  --source-url https://example.org/results.json \
  --retrieved-at 2026-08-03 \
  --model qwen3-30b-a3b \
  --json
REFERENCE HARNESSlm-eval@0.4.12PIN6d642546f468

A score is a record,
not a cell.

The comparison gate requires a matching fingerprint. A valid but incomplete import remains useful—and visibly excluded from rankings until the consequential fields exist.

COMPARABILITY FINGERPRINTREQUIRED FIELDS / V1
01checkpoint revision02raw artifact sha25603harness commit04task hash + version05dataset revision06chat template hash07system prompt hash08reasoning settings09sampling + seeds10sample counts + limit11runtime + precision12hardware topology
PUBLIC RUN REGISTRY

No publication-eligible MoE runs yet.

That is deliberate. We will publish the first normalized result only when its raw artifact, exact model revision, harness identity, task configuration, sample counts, and known gaps survive validation.

Create measured evidence

CONTRIBUTION STANDARD

Bring the run.
Keep the context.

The open schema is built for reviewable results, independent reproductions, and corrections that preserve history.

Open DeployBench