BENCHMARK CATALOG · 01
What the frontier is measured on.
Benchmarks are different instruments. A preference arena, a coding-agent test, and an expert exam should never be flattened into the same number.
01Strong candidate
General capability
LiveBench
Maintained by LiveBench
Reasoning, coding, mathematics, data analysis, language, instruction following, and agentic coding.
- Primary metric
- Overall and category scores
- Open-rank rule
- Same dated release; use the open-weights filter and identical evaluation settings.
Best suited to a direct open-model ranking when one official snapshot covers the field.
Official benchmark ↗02Scaffold-sensitive
Software engineering
SWE-bench Verified
Maintained by SWE-bench
Whether an agent can resolve real GitHub issues by producing patches that pass repository tests.
- Primary metric
- % issues resolved
- Open-rank rule
- Rank model + agent pairs only. Never collapse different scaffolds into a model-only score.
Useful only when the agent, tools, or test-time system are reported as part of the result.
Official benchmark ↗03Agent + model
Computer use
Terminal-Bench 2.0
Maintained by Terminal-Bench
End-to-end execution of real tasks inside isolated terminal environments.
- Primary metric
- Task resolution rate
- Open-rank rule
- Compare the same benchmark version and report the agent scaffold beside the model.
Useful only when the agent, tools, or test-time system are reported as part of the result.
Official benchmark ↗04System-sensitive
Abstract reasoning
ARC-AGI-2
Maintained by ARC Prize
Novel visual transformation rules that demand adaptation rather than recalled knowledge.
- Primary metric
- % tasks solved + cost
- Open-rank rule
- Separate base-model results from test-time search, tools, and multi-model systems.
Useful only when the agent, tools, or test-time system are reported as part of the result.
Official benchmark ↗05Strong candidate
Expert knowledge
Humanity’s Last Exam
Maintained by CAIS + Scale AI
Difficult, expert-written questions spanning many academic and professional domains.
- Primary metric
- Accuracy / exact match
- Open-rank rule
- Match the public set, grading method, and tool-access policy before ordering models.
Best suited to a direct open-model ranking when one official snapshot covers the field.
Official benchmark ↗06Strong candidate
Tool use
Berkeley Function Calling
Maintained by UC Berkeley
Function selection, argument construction, relevance detection, and multi-turn tool use.
- Primary metric
- Overall accuracy
- Open-rank rule
- Use one BFCL release and preserve model format, prompting, and executable/non-executable split.
Best suited to a direct open-model ranking when one official snapshot covers the field.
Official benchmark ↗07Prompt-sensitive
Knowledge + reasoning
MMLU-Pro
Maintained by TIGER-Lab
Harder, ten-choice questions across broad academic disciplines with more reasoning required.
- Primary metric
- Accuracy
- Open-rank rule
- Accept only results with the same prompt, chain-of-thought policy, and evaluation implementation.
Important frontier signal, but extra normalization is required before ordering open models.
Official benchmark ↗08Preference signal
Human preference
LMArena
Maintained by LMArena
Blind, pairwise user preference across general prompts and specialist arenas.
- Primary metric
- Arena rating
- Open-rank rule
- Keep preference ranks separate from objective benchmarks and verify open-weight model identity.
Important frontier signal, but extra normalization is required before ordering open models.
Official benchmark ↗