Pinned-card associatedOWNER REPORTED
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=0.95
4 context gaps attached
- The owner does not explicitly label the score as a percentage or define the metric scale.
- Shot count, sample count, prompt, and benchmark revision are not reported.
- Inference engine, runtime configuration, and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=1harness=Kimi Code
5 context gaps attached
- The owner does not explicitly label the score as a percentage or define the metric scale.
- The Kimi Code harness version and full harness configuration are not reported.
- Task revision, trial count, and aggregation method are not reported.
- Inference engine and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
Kimi K3
BrowseComp
Version not reported
91.2owner-reported scorereported score0-100 as presentedmax
reasoning_effort=maxtemperature=1top_p=1context_compaction_trigger_tokens=300000
5 context gaps attached
- The owner does not explicitly label the score as a percentage or define the metric scale.
- The browsing tool implementation, harness version, and context-compaction algorithm are not reported.
- Benchmark revision, trial count, and aggregation method are not reported.
- Inference engine and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
DeepSeek V4 Pro
GPQA
Diamond
90.1Pass@1percentage points0-100think_max
reasoning_mode=Think Maxtemperature=1context_tokens=384000system_prompt_variant=Think Max special system prompt
4 context gaps attached
- Top-p, exact prompt text for GPQA, shot count, and sample count are not reported.
- Benchmark revision and evaluation harness version are not reported.
- Inference engine and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
DeepSeek V4 Pro
LiveCodeBench
v6
93.5Pass@1-COTpercentage points0-100think_max
reasoning_mode=Think Maxsystem_prompt_variant=Think Max special system prompt
4 context gaps attached
- Prompt, top-p, sample count, and maximum output length are not reported for this benchmark.
- LiveCodeBench task revision and harness version are not reported.
- Inference engine and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
DeepSeek V4 Pro
SWE-bench
Verified
80.6Resolvedpercentage points0-100think_max
reasoning_mode=Think Maxharness=DeepSeek internal evaluation frameworktooling=bash and file-edit toolsmax_interaction_steps=500context_tokens=512000
4 context gaps attached
- The internal framework version and complete tool prompts are not published.
- SWE-bench dataset commit, container images, trial count, and aggregation method are not reported.
- Sandbox resources and inference hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
owner-reported scorereported score0-100 as presentedunspecified
temperature=1top_p=0.95max_new_tokens=163840
4 context gaps attached
- The owner does not explicitly label the score as a percentage or define the metric scale.
- Reasoning effort, prompt, shot count, sample count, and judge configuration are not reported for GPQA.
- Benchmark revision, harness version, and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
GLM-5.2
SWE-bench Pro
Version not reported
62.1owner-reported scorereported score0-100 as presentedunspecified
framework=OpenHandsprompt_variant=tailored instruction prompttemperature=1top_p=1max_new_tokens=32000context_tokens=400000
5 context gaps attached
- The owner does not explicitly define the metric or label the score as a percentage.
- OpenHands version and the tailored instruction prompt are not published.
- Dataset commit, container images, task count, trials, and aggregation method are not reported.
- Inference hardware is not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
GLM-5.2
MCP-Atlas
500-task public set
76.8owner-reported scorereported score0-100 as presentedthinking
thinking_mode=truetask_count=500timeout_seconds=600judge_model=Gemini-3.0-Pro
4 context gaps attached
- The owner does not explicitly define the metric or label the score as a percentage.
- Exact reasoning effort, temperature, context limit, harness version, and tool configuration are not reported.
- Sample count per task, aggregation method, and inference hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
Gemma 4 26B A4B IT
AIME
2026 · no tools
88.3accuracypercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=truetool_use=false
4 context gaps attached
- Prompt, sampling parameters, sample count, and scoring implementation are not reported.
- Benchmark task revision and evaluation harness are not reported.
- Context limit, maximum output length, runtime, and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
Gemma 4 26B A4B IT
LiveCodeBench
v6
77.1owner-reported percentagepercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=true
4 context gaps attached
- The exact LiveCodeBench metric, prompt, sampling parameters, and sample count are not reported.
- Benchmark task revision and evaluation harness are not reported.
- Context limit, maximum output length, runtime, and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Pinned-card associatedOWNER REPORTED
Gemma 4 26B A4B IT
GPQA
Diamond
82.3accuracypercentage points0-100thinking
model_stage=instruction-tunedthinking_mode=true
4 context gaps attached
- Prompt, shot count, sampling parameters, sample count, and scoring implementation are not reported.
- Benchmark task revision and evaluation harness are not reported.
- Context limit, maximum output length, runtime, and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED
averaged accuracypercentage points0-100thinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=32768samples_per_query=10
4 context gaps attached
- The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
- Exact prompt, benchmark revision, and evaluation harness are not reported.
- Inference engine and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED
Qwen3-30B-A3B
AIME
2025 · Parts I and II
70.9averaged accuracypercentage points0-100thinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=38912question_count=30samples_per_query=64
4 context gaps attached
- The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
- Exact prompt text, scoring parser, and test-data revision are not reported.
- Inference engine and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.
Model name onlyOWNER REPORTED
Qwen3-30B-A3B
LiveCodeBench
v5 · 2024-10 through 2025-02
62.6owner-reported scorereported score0-100 as presentedthinking
thinking_mode=truetemperature=0.6top_p=0.95top_k=20max_new_tokens=32768benchmark_window=2024-10 through 2025-02prompt_variant=official prompt with the program-only restriction removed
4 context gaps attached
- The technical report identifies the model by name but does not bind the result to artifact revision ad44e777bcd18fa416d9da3bd8f70d33ebb85d39.
- The report does not explicitly define the table's LiveCodeBench metric or sample count.
- Exact task commit, harness version, inference engine, and hardware are not reported.
- The provider does not publish a digest of the tensor artifact used for the evaluation.