MashupAbility//OpenSource

BENCHMARK LEDGER · JUL 20, 2026

FRONTIER EVALS · OPEN-WEIGHT RANKINGS

Rank the evidence, not the marketing.

The leading benchmarks in one ledger—then a separate, defensible view of how open-weight models rank when the harness, version, and system setup actually match.

8

benchmark families

6

capability lanes

0

blended mystery scores

3

ranking gates

BENCHMARK CATALOG · 01

What the frontier is measured on.

Benchmarks are different instruments. A preference arena, a coding-agent test, and an expert exam should never be flattened into the same number.

01Strong candidate

General capability

LiveBench

Maintained by LiveBench

Reasoning, coding, mathematics, data analysis, language, instruction following, and agentic coding.

Primary metric
Overall and category scores
Open-rank rule
Same dated release; use the open-weights filter and identical evaluation settings.

Best suited to a direct open-model ranking when one official snapshot covers the field.

Official benchmark ↗
02Scaffold-sensitive

Software engineering

SWE-bench Verified

Maintained by SWE-bench

Whether an agent can resolve real GitHub issues by producing patches that pass repository tests.

Primary metric
% issues resolved
Open-rank rule
Rank model + agent pairs only. Never collapse different scaffolds into a model-only score.

Useful only when the agent, tools, or test-time system are reported as part of the result.

Official benchmark ↗
03Agent + model

Computer use

Terminal-Bench 2.0

Maintained by Terminal-Bench

End-to-end execution of real tasks inside isolated terminal environments.

Primary metric
Task resolution rate
Open-rank rule
Compare the same benchmark version and report the agent scaffold beside the model.

Useful only when the agent, tools, or test-time system are reported as part of the result.

Official benchmark ↗
04System-sensitive

Abstract reasoning

ARC-AGI-2

Maintained by ARC Prize

Novel visual transformation rules that demand adaptation rather than recalled knowledge.

Primary metric
% tasks solved + cost
Open-rank rule
Separate base-model results from test-time search, tools, and multi-model systems.

Useful only when the agent, tools, or test-time system are reported as part of the result.

Official benchmark ↗
05Strong candidate

Expert knowledge

Humanity’s Last Exam

Maintained by CAIS + Scale AI

Difficult, expert-written questions spanning many academic and professional domains.

Primary metric
Accuracy / exact match
Open-rank rule
Match the public set, grading method, and tool-access policy before ordering models.

Best suited to a direct open-model ranking when one official snapshot covers the field.

Official benchmark ↗
06Strong candidate

Tool use

Berkeley Function Calling

Maintained by UC Berkeley

Function selection, argument construction, relevance detection, and multi-turn tool use.

Primary metric
Overall accuracy
Open-rank rule
Use one BFCL release and preserve model format, prompting, and executable/non-executable split.

Best suited to a direct open-model ranking when one official snapshot covers the field.

Official benchmark ↗
07Prompt-sensitive

Knowledge + reasoning

MMLU-Pro

Maintained by TIGER-Lab

Harder, ten-choice questions across broad academic disciplines with more reasoning required.

Primary metric
Accuracy
Open-rank rule
Accept only results with the same prompt, chain-of-thought policy, and evaluation implementation.

Important frontier signal, but extra normalization is required before ordering open models.

Official benchmark ↗
08Preference signal

Human preference

LMArena

Maintained by LMArena

Blind, pairwise user preference across general prompts and specialist arenas.

Primary metric
Arena rating
Open-rank rule
Keep preference ranks separate from objective benchmarks and verify open-weight model identity.

Important frontier signal, but extra normalization is required before ordering open models.

Official benchmark ↗

OPEN-WEIGHT RANKING BOARD · 02

Not ranked until the runs match.

RANKING STATUSEvidence collection

The current model ledger contains useful reported scores, but not yet a normalized independent comparison set.

Reported scores below are an intake queue—not a leaderboard. They stay unranked until the gate is cleared.

Evidence waiting to qualify for open-weight benchmark rankings
BenchmarkCurrent evidenceReported scoresCoverageRanking gate
SWE-bench VerifiedProvider model cardsDeepSeek V4 Pro 80.6 · MiniMax M3 80.5 · Inkling 77.6 · Mistral Medium 3.5 77.6 · Qwen3.6 27B 77.25 tracked modelsBlockedAgent scaffold and run settings are not normalized.
Terminal-Bench 2.1Provider model cardGLM-5.2 81.01 tracked modelCollectingNeeds matched results for at least two more open-weight models.
MMLU-ProProvider model cardLlama 4 Maverick 80.51 tracked modelCollectingNeeds a shared harness snapshot and broader open-model coverage.
LiveCodeBench v6Provider model cardNemotron 3 Ultra 89.01 tracked modelCollectingNeeds matched version, date range, and inference settings.
MCP AtlasProvider model cardKimi K2.7 Code 76.01 tracked modelCollectingNeeds independent coverage across the tracked open-weight set.

RANKING PROTOCOL · 03

Three gates before rank.

01

Match the test

Same benchmark release, task split, grading rules, date window, and prompt policy.

02

Expose the system

Record the model checkpoint, quantization, inference settings, tools, scaffold, and test-time compute.

03

Keep ranks local

Rank within each benchmark. No composite “intelligence” score unless its weighting is explicit and justified.