BENCHMARK EVIDENCE / V1

Scores with receipts.

First-party results are useful signals—not a leaderboard. Every score keeps its model scope, reported settings, source, and missing run context attached.

Comparable means same artifact and methodology fingerprint. Everything else remains evidence, not rank.

15OWNER-REPORTED CLAIMS
5EXACT MODEL RECORDS
0NORMALIZED RUNS
0STRICTLY COMPARABLE

REPORTED EVIDENCE STREAM

Context stays attached.

These values come from official model cards and technical reports. None is represented as independently reproduced, cryptographically bound to the evaluated weights, or safe for cross-model ranking.

Open the evidence API
15OF 15 CLAIMS
Pinned-card associatedOWNER REPORTED

Kimi K3

GPQA

Diamond

93.5
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=0.95
4 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • Shot count, sample count, prompt, and benchmark revision are not reported.
  • Inference engine, runtime configuration, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Kimi K3

Terminal-Bench

2.1

88.3
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=1harness=Kimi Code
5 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • The Kimi Code harness version and full harness configuration are not reported.
  • Task revision, trial count, and aggregation method are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Kimi K3

BrowseComp

Version not reported

91.2
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=1context_compaction_trigger_tokens=300000
5 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • The browsing tool implementation, harness version, and context-compaction algorithm are not reported.
  • Benchmark revision, trial count, and aggregation method are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

DeepSeek V4 Pro

GPQA

Diamond

90.1
Pass@1percentage points0-100think_max
reasoning_mode=Think Maxtemperature=1context_tokens=384000system_prompt_variant=Think Max special system prompt
4 context gaps attached
  • Top-p, exact prompt text for GPQA, shot count, and sample count are not reported.
  • Benchmark revision and evaluation harness version are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

DeepSeek V4 Pro

LiveCodeBench

v6

93.5
Pass@1-COTpercentage points0-100think_max
reasoning_mode=Think Maxsystem_prompt_variant=Think Max special system prompt
4 context gaps attached
  • Prompt, top-p, sample count, and maximum output length are not reported for this benchmark.
  • LiveCodeBench task revision and harness version are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

DeepSeek V4 Pro

SWE-bench

Verified

80.6
Resolvedpercentage points0-100think_max
reasoning_mode=Think Maxharness=DeepSeek internal evaluation frameworktooling=bash and file-edit toolsmax_interaction_steps=500context_tokens=512000
4 context gaps attached
  • The internal framework version and complete tool prompts are not published.
  • SWE-bench dataset commit, container images, trial count, and aggregation method are not reported.
  • Sandbox resources and inference hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

GLM-5.2

GPQA

Diamond

91.2
owner-reported scorereported score0-100 as presentedunspecified
temperature=1top_p=0.95max_new_tokens=163840
4 context gaps attached
  • The owner does not explicitly label the score as a percentage or define the metric scale.
  • Reasoning effort, prompt, shot count, sample count, and judge configuration are not reported for GPQA.
  • Benchmark revision, harness version, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

GLM-5.2

SWE-bench Pro

Version not reported

62.1
owner-reported scorereported score0-100 as presentedunspecified
framework=OpenHandsprompt_variant=tailored instruction prompttemperature=1top_p=1max_new_tokens=32000context_tokens=400000
5 context gaps attached
  • The owner does not explicitly define the metric or label the score as a percentage.
  • OpenHands version and the tailored instruction prompt are not published.
  • Dataset commit, container images, task count, trials, and aggregation method are not reported.
  • Inference hardware is not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

GLM-5.2

MCP-Atlas

500-task public set

76.8
owner-reported scorereported score0-100 as presentedthinking
thinking_mode=truetask_count=500timeout_seconds=600judge_model=Gemini-3.0-Pro
4 context gaps attached
  • The owner does not explicitly define the metric or label the score as a percentage.
  • Exact reasoning effort, temperature, context limit, harness version, and tool configuration are not reported.
  • Sample count per task, aggregation method, and inference hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Gemma 4 26B A4B IT

AIME

2026 · no tools

88.3
accuracypercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=truetool_use=false
4 context gaps attached
  • Prompt, sampling parameters, sample count, and scoring implementation are not reported.
  • Benchmark task revision and evaluation harness are not reported.
  • Context limit, maximum output length, runtime, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Gemma 4 26B A4B IT

LiveCodeBench

v6

77.1
owner-reported percentagepercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=true
4 context gaps attached
  • The exact LiveCodeBench metric, prompt, sampling parameters, and sample count are not reported.
  • Benchmark task revision and evaluation harness are not reported.
  • Context limit, maximum output length, runtime, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED

Gemma 4 26B A4B IT

GPQA

Diamond

82.3
accuracypercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=true
4 context gaps attached
  • Prompt, shot count, sampling parameters, sample count, and scoring implementation are not reported.
  • Benchmark task revision and evaluation harness are not reported.
  • Context limit, maximum output length, runtime, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED

Qwen3-30B-A3B

GPQA

Diamond

65.8
averaged accuracypercentage points0-100thinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=32768samples_per_query=10
4 context gaps attached
  • The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
  • Exact prompt, benchmark revision, and evaluation harness are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED

Qwen3-30B-A3B

AIME

2025 · Parts I and II

70.9
averaged accuracypercentage points0-100thinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=38912question_count=30samples_per_query=64
4 context gaps attached
  • The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
  • Exact prompt text, scoring parser, and test-data revision are not reported.
  • Inference engine and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED

Qwen3-30B-A3B

LiveCodeBench

v5 · 2024-10 through 2025-02

62.6
owner-reported scorereported score0-100 as presentedthinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=32768benchmark_window=2024-10 through 2025-02prompt_variant=official prompt with the program-only restriction removed
4 context gaps attached
  • The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
  • The report does not explicitly define the table's LiveCodeBench metric or sample count.
  • Exact task commit, harness version, inference engine, and hardware are not reported.
  • The provider does not publish a digest of the tensor artifact used for the evaluation.

DETERMINISTIC INGESTION

Turn raw output into a reviewable staging pack.

The v0.1 importer hashes aggregate bytes, preserves metric/filter pairs and explicit unknowns, and never executes model or task code. Staging output is not publication eligibility.

MOEMODELS / EVIDENCE IMPORT
$ moemodels ingest lm-eval results.json \
  --source-url https://example.org/results.json \
  --retrieved-at 2026-08-03 \
  --model qwen3-30b-a3b \
  --json
REFERENCE HARNESSlm-eval@0.4.12PIN6d642546f468

A score is a record,
not a cell.

The comparison gate requires a matching fingerprint. A valid but incomplete import remains useful—and visibly excluded from rankings until the consequential fields exist.

COMPARABILITY FINGERPRINTREQUIRED FIELDS / V1
01checkpoint revision02raw artifact sha25603harness commit04task hash + version05dataset revision06chat template hash07system prompt hash08reasoning settings09sampling + seeds10sample counts + limit11runtime + precision12hardware topology
PUBLIC RUN REGISTRY

No publication-eligible MoE runs yet.

That is deliberate. We will publish the first normalized result only when its raw artifact, exact model revision, harness identity, task configuration, sample counts, and known gaps survive validation.

Create measured evidence

CONTRIBUTION STANDARD

Bring the run.
Keep the context.

The open schema is built for reviewable results, independent reproductions, and corrections that preserve history.

Open DeployBench