Models
Frontier & open-weight models compared by capability — columns are flagship models, rows are the headline benchmarks that measure each one. Every number is the published figure, restated with a link to the source it came from · as of 2026-06-14. Curated matrix is 116d old — a refresh is due; live leaderboard moves still flow into Model changes below.
We cite the benchmark authorities rather than re-rank them — Epoch AI, LMArena and Artificial Analysis. ChangeRadar's job is the changes: what gets released, repriced, deprecated, or quietly shifts behind a stable model id.
Vendor champions · best flagship per vendor · GPQA Diamond
Each vendor's most recent publicly-available flagship, ordered by score (highest first) — Google · OpenAI · Anthropic · xAI highlighted. The weekly market-watch surfaces new releases automatically; one tagged reported is the latest release shown with vendor-reported scores (linked to source) until we independently cite it. Score = GPQA Diamond; every number links to its source.
Compare two models
Agentic coding
| Benchmark | Claude Fable 5 | Gemini 3.1 Pro | Δ |
|---|---|---|---|
| SWE-bench Verified i% resolved (pass@1) | 95 | 80.6 | +14.4 |
| SWE-bench Pro i% resolved (pass@1) | 80.3 | 54.2 | +26.1 |
3 coverage gaps — only one model reports these
- Terminal-Bench — Gemini 3.1 Pro 68.5
- LiveCodeBench — Gemini 3.1 Pro 2887
- FrontierCode — Claude Fable 5 29.3
Tool use & agents
5 coverage gaps — only one model reports these
- TAU-bench — Gemini 3.1 Pro 99.3
- OSWorld — Claude Fable 5 85
- BrowseComp — Gemini 3.1 Pro 85.9
- GDPval-AA — Claude Fable 5 1932
- MCP Atlas — Gemini 3.1 Pro 69.2
Science & reasoning
| Benchmark | Claude Fable 5 | Gemini 3.1 Pro | Δ |
|---|---|---|---|
| GPQA Diamond i% accuracy | 92.6 | 94.3 | -1.7 |
2 coverage gaps — only one model reports these
- Humanity's Last Exam — Gemini 3.1 Pro 44.4
- ARC-AGI-2 — Gemini 3.1 Pro 77.1
General knowledge
1 coverage gap — only one model reports these
- MMMLU — Gemini 3.1 Pro 92.6
Multimodal
2 coverage gaps — only one model reports these
- MMMU — Gemini 3.1 Pro 80.5
- GDP.pdf — Claude Fable 5 29.8
Long context
1 coverage gap — only one model reports these
- MRCR — Gemini 3.1 Pro 84.9
Shared benchmarks first (with Δ when both report the same scale); one-sided coverage collapses below. Pick any two models — or a champion above — and the URL becomes shareable.
All models · benchmark matrix
reported columns (Claude Opus 5.5, GPT-5, Grok 4.7, Qwen3.8-Max, DeepSeek-V4.1-Flash, Mistral Large 4, Kimi K3, GLM-5.3) are auto-discovered by our weekly market-watch from each vendor's own reported numbers — not independently verified, and shown when a vendor ships a model newer than the hand-cited column beside it. Full claim sets are in Vendor-reported benchmarks below.
Agentic coding
Can the model fix real bugs, ship features, and operate a dev environment end-to-end as a coding agent — the single most-watched capability in 2026 vendor launches.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SWE-bench Verified i % resolved (pass@1) | 95 | — | 88.6 | — | — | 74.9 | 80.6 | — | — | 80.6 | — | 80.4 | — | 80.2 | 76.8 | — | — | — | 77.6 | — | — |
| SWE-bench Pro i % resolved (pass@1) | 80.3 | 89.9 | 69.2 | 77.8 | 58.6 | — | 54.2 | — | — | — | — | 60.6 | 67.7 | 58.6 | — | 62.1 | — | — | — | — | — |
| Terminal-Bench i % solved (pass@1) | — | 66.4 | 74.6 | 88 | 82.7 | — | 68.5 | — | 37.6 | 67.9 | 90.6 | 69.7 | 86.6 | 66.7 | 88.3 | 81 | 28.3 | — | — | 28.3 | — |
| LiveCodeBench i % pass@1 | — | — | — | — | — | — | 2887 | — | — | 93.5 | — | — | — | 89.6 | — | — | 80.5 | 43.4 | — | — | 29.7 |
Tool use & agents
Beyond writing code: can the model select and chain tools, follow policy, drive a computer/browser, and complete long-horizon multi-step tasks.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TAU-bench i % pass / pass^k | — | — | — | — | — | — | 99.3 | 98 | — | — | — | — | — | — | — | — | — | — | — | — | — |
| OSWorld i % success | 85 | 81.8 | 83.4 | — | 78.7 | — | — | — | — | — | — | — | 86.1 | 73.1 | — | — | — | — | — | — | — |
| BrowseComp i % accuracy | — | — | 84.3 | — | 90.1 | — | 85.9 | — | — | — | — | — | — | 83.2 | 91.2 | — | — | — | — | — | — |
Math
Competition and research-level mathematical reasoning, increasingly reported on uncontaminated/post-cutoff problem sets.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HMMT i % accuracy (pass@1) | — | — | — | — | — | — | — | — | — | — | — | 97.1 | — | 92.7 | — | — | — | — | — | — | — |
| MATH i % accuracy | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 61.2 | — | — | 89 |
Science & reasoning
Expert-level, Google-proof reasoning across the sciences and broad academia — the benchmarks vendors point to when claiming 'PhD-level' or 'frontier' reasoning.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPQA Diamond i % accuracy | 92.6 | — | 93.6 | — | 93.6 | 88.4 | 94.3 | 90.1 | — | 90.1 | 90.9 | 92.4 | 92.6 | 90.5 | 93.5 | 91.2 | — | 69.8 | — | — | 42.4 |
| Humanity's Last Exam i % accuracy | — | 67.7 | 57.9 | 64.5 | 57.2 | — | 44.4 | — | — | 37.7 | — | 41.4 | 43.6 | 54 | 43.5 | 40.5 | 62.5 | — | — | — | — |
| ARC-AGI-2 i % solved | — | — | — | — | 85 | — | 77.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
General knowledge
Broad multi-subject factual and reasoning coverage; the classic 'how much does it know' bucket, now reported via the harder Pro variant since base MMLU is saturated.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MMLU-Pro i % accuracy | — | — | — | — | — | — | — | — | — | 87.5 | — | — | — | — | — | — | — | 80.5 | — | — | 67.5 |
Multimodal
Vision + language understanding and visual reasoning — how well the model interprets images, diagrams, charts and figures.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MMMU i % accuracy | — | — | — | — | 83.2 | 84.2 | 80.5 | — | — | — | — | — | — | 79.4 | — | — | — | 73.4 | — | — | 64.9 |
Long context
Retrieval and reasoning quality as context length grows into the hundreds-of-thousands / millions of tokens — beyond simple needle-in-a-haystack.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MRCR i % accuracy | — | — | — | — | 74 | — | 84.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Human preference
Aggregate real-user preference from blind head-to-head comparisons — the closest thing to a 'do people actually like the answers' metric, and the one number vendors love to top.
| Benchmark | Fable 5 | Opus 5.5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5 reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.7 reported | DeepSeek-V4-Pro open | DeepSeek-V4.1-Flash reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | GLM-5.3 reported | Llama 4 Maverick open | Mistral Medium 3.5 open | Mistral Large 4 reported | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LMArena Elo i Elo | — | — | — | 1458 | 1474 | — | — | — | — | — | — | — | — | 1466 | — | — | — | 1417 | — | — | 1338 |
Showing 13 flagship models across the headline benchmarks; 53 models and
84 benchmarks tracked in total from 60 primary & aggregator sources.
Claude Mythos 5 numbers are limited (access-restricted preview).
Numbers are published facts, restated with a per-cell source link; vendor benchmark charts are linked to their source, not rehosted.
Sources include: 9to5Google (Gemini 3 Flash launch coverage) · AIFire (citing OpenAI GPT-5.2 release) · Artificial Analysis · BenchLM.ai (citing LMArena) · BinaryVerse AI (xAI official figures) · BuildFastWithAI (citing OpenAI GPT-5.5 launch) · Caylent · Codersera (reporting Moonshot's figures) · DataCamp · DataCamp (reproducing Meta's Llama 4 launch chart) · DeepSeek-AI (arXiv 2512.02556) · DeepSeek-AI (Hugging Face model card) · Google (official Gemini 2.5 launch blog) · Google (official Gemini 3 Flash launch blog) · Google (official Gemini 3 launch blog) · Google DeepMind (Gemini 3.1 Pro model card) · +36 more.
Vendor-reported benchmarks
Numbers as claimed by the vendor on their own model/system card — not independently verified and often measured with a favourable harness. We track each vendor's claims over time and link to the source; cross-check against the cited matrix above.
| SWE-bench Pro · max effort | 89.9% |
| SWE-bench Multilingual · max effort | 93.9% |
| SWE-bench Multimodal · max effort | 61.4% |
| Terminal-Bench 4.0 · xhigh effort; standard error ±2.6 pts | 66.4% |
| Terminal-Bench-Science 0.1 · max effort | 58.7% |
| Humanity's Last Exam · with tools, max effort | 67.7% |
| Humanity's Last Exam (no tools) · max effort | 64.4% |
| FrontierCode v1.1 (Main) · max effort | 54.4% |
| CursorBench 4.0 · max effort | 57.8% |
| OSWorld 2.0 (Partial Pass) · max effort | 81.8% |
| GDPval-AA v2.1 · max effort, 44 occupations | 1846 Elo |
| AA-Briefcase v1.1 · max effort | 1822 Elo |
| AIME 2025 · without tools | 94.6% |
| SWE-bench Verified · fixed subset of 477 verified tasks | 74.9% |
| GPQA Diamond · GPT-5 pro variant, without tools | 88.4% |
| MMMU | 84.2% |
| Aider Polyglot | 88% |
| HealthBench Hard | 46.2% |
| CursorBench 4.0 · xhigh effort | 46.3% |
| DeepSWE v1.1 · high-effort score | 71.0% |
| EEBench · electrical engineering | 64.0% |
| AA Briefcase v1.1 · xhigh effort | 1657 Elo |
| Terminal-Bench 4.0 · agentic coding, xhigh effort | 37.6% |
| Harvey Legal Agent Benchmark | 19.6% |
| HealthBench Professional | 56.7% |
| GDPval · xhigh effort | 1695 Elo |
| LatchBio Biosafety Benchmark · dual-use domains | 62.4% |
| HackerBench v0.3 · risky dual-use prompts allowed through | 3.3% |
| GPQA Diamond · No tools | 94.3% |
| ARC-AGI-2 · ARC Prize Verified | 77.1% |
| SWE-Bench Verified · Single attempt | 80.6% |
| SWE-Bench Pro (Public) · Single attempt | 54.2% |
| Terminal-Bench 2.0 · Terminus-2 harness | 68.5% |
| Humanity's Last Exam · No tools | 44.4% |
| Humanity's Last Exam · Search (blocklist) + Code | 51.4% |
| LiveCodeBench Pro | 2887 Elo |
| MATH | 95.1% |
| Terminal Bench 2.1 | 86.6 % |
| SWE-bench Pro | 67.7 % |
| GPQA Diamond | 92.6 % |
| PaperBench | 93.0 % |
| IFBench | 82.8 % |
| FrontierSWE | 73.5 % |
| DeepSWE 1.1 | 56.6 % |
| OSWorld-Verified · multimodal/agentic | 86.1 % |
| QwenSWEBench · internal benchmark | 80.7 % |
| NL2Repo-Bench | 55.9 % |
| MLS-Bench-Lite | 41.0 % |
| Humanity's Last Exam · text-only subset | 43.6 % |
| GPQA Diamond · Pass@1 | 90.9% |
| Terminal-Bench 2.1 | 90.6% |
| Codeforces | 3471 Rating |
| DeepSWE v1.1 | 74.2% |
| NL2Repo-Bench | 65.4% |
| CyberGym | 88.1% |
| HLE · without tools; 39.1% with tools | 36.8% |
| MathArena Apex | 65.6% |
| Terminal-Bench 3.0 | 30.0% |
| Terminal-Bench 4.0 | 31.2% |
| Agents' Last Exam | 31.8% |
| MMLU Pro · instruction-tuned, 0-shot | 80.5% |
| GPQA Diamond · instruction-tuned, 0-shot | 69.8% |
| MMMU · instruction-tuned, image reasoning, 0-shot | 73.4% |
| MathVista · instruction-tuned, 0-shot | 73.7% |
| ChartQA · instruction-tuned, 0-shot relaxed accuracy | 90.0% |
| DocVQA · instruction-tuned, 0-shot anls metric | 94.4% |
| LiveCodeBench · instruction-tuned, 0-shot pass@1, 10.01.2024-02.01.2025 | 43.4% |
| MATH · pre-trained, 4-shot em_maj1@1 | 61.2% |
| MBPP · pre-trained, 3-shot pass@1 | 77.6% |
| Multilingual MMLU · instruction-tuned | 84.6% |
| MTOB (half-book) · instruction-tuned, eng->kgv/kgv->eng long context | 54.0/46.4 chrF |
| DeepSWE v1.1 | 61.7% |
| SWE-Atlas-QnA | 59.4% |
| Terminal-Bench 4 | 28.3% |
| Coding Agent Index | 49.8% |
| AutomationBench · 657 business workflows across Gmail, Google Sheets, Slack, Salesforce | 59.9% |
| Cybench · 40 security challenges | 93% |
| CyberGym-E2E-AA · Reproduce and patch vulnerability test | 82% |
| Lakera B3 Attack Resistance | 93.3% |
| AA-Briefcase | 1393 Elo |
| Dense200 · Visual grounding benchmark | 42% |
| GPQA Diamond · max reasoning effort, temperature=1.0 | 93.5% |
| SWE-Bench Verified · max reasoning effort | 76.8% |
| AIME 2025 · max reasoning effort, from K2.5 card; K3 reports AIME 2026 at 96.4% | 96.1% |
| SWE-Bench Multilingual · from K2.5; K3 model card shows 76.7% | 73.0% |
| Terminal-Bench 2.1 · Kimi Code harness, max reasoning | 88.3% |
| FrontierSWE · max reasoning, Kimi Code harness | 81.2 dominance score |
| SWE-Marathon · max reasoning effort | 42.0% |
| BrowseComp · max reasoning, compaction at 300K | 91.2% |
| DeepSWE · Kimi Code harness; 67.3% with mini-SWE-agent | 67.5% |
| Humanity's Last Exam (HLE-Full) · max reasoning, no tools; 56.0% with tools | 43.5% |
| DeepSearchQA · max reasoning effort | 95.0 F1 |
| Program Bench · raw hidden-test pass rate, max reasoning | 77.8% |
| Terminal-Bench 3.0 · max reasoning effort, avg@3 | 28.3 pass@1 |
| DeepSWE v1.1 · max reasoning effort | 66.9 % |
| Agents' Last Exam · CLI, max reasoning effort | 28.5 % |
| Terminal-Bench 2.1 · temperature=1.0, top_p=1, 65536 max_new_tokens | 88.2 % |
| Humanity's Last Exam · with tools, 300K maximum context | 62.5 % |
| CyberGym · Pass@1, max reasoning effort, 1,507 tasks | 84.5 % |
| SWE-Marathon v1.1 · long-horizon coding | 42.5 % |
| FrontierSWE v2 · Max effort, Proximal evaluation | 78.1 % |
| LiveCodeBench · Vals AI run | 80.5 % |
| GDPval-AA v2 · Economic knowledge work evaluation | 1769 Elo |
Model changes
New releases, deprecations, and benchmark score moves we’ve recorded — newest first.
-
LMArena (text) scores changed
Arena Elo (text, overall) updated.
-
Mistral AI reported benchmarks updated
Mistral Large 4: 10 benchmark claims (via web search)
-
AI2 OLMo models changed
Model registry updated with new model additions and removals: MolmoWeb models removed, new Bolmo, Bwen, and Llama-3-Blama models added.
-
Anthropic models model list changed
New Anthropic model claude-haiku-5-5 is now available.
-
NVIDIA Nemotron models changed
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 model removed or no longer available as of 2026-06-03.
-
IBM Granite models changed
IBM Granite model ibm-granite/granite-4.0-micro-base with Apache 2.0 license was removed as of 2025-09-16.
-
Meta reported benchmarks updated
Llama 4 Maverick: 11 benchmark claims (via web search)
-
Google Gemma models changed
New Google Gemma model variant DiarizationLM-Gemma-4-E4B-v1 released with Apache 2.0 license.
-
Moonshot AI reported benchmarks updated
Kimi K3: 12 benchmark claims (via web search)
-
Zhipu AI (Z.ai) reported benchmarks updated
GLM-5.3: 10 benchmark claims (via web search)
-
Alibaba reported benchmarks updated
Qwen3.8-Max: 12 benchmark claims (via web search)
-
DeepSeek reported benchmarks updated
DeepSeek-V4.1-Flash: 11 benchmark claims (via web search)
-
xAI reported benchmarks updated
Grok 4.7: 10 benchmark claims (via web search)
-
Google reported benchmarks updated
Gemini 3.1 Pro: 9 benchmark claims (via web search)
-
OpenAI reported benchmarks updated
GPT-5: 6 benchmark claims (via web search)
-
Anthropic reported benchmarks updated
Claude Opus 5.5: 15 benchmark claims (via web search)
-
LMArena (text) scores changed
Arena Elo (text, overall) updated.
-
Microsoft Phi models changed
Removed Microsoft Phi model Dayhoff-170M-UR90-46000 with MIT license.
-
AI2 OLMo models changed
Model changed from MolmoWeb-4B-Native to AstaBrief_8B_SFT with updated license date from 2026-03-23 to 2026-09-10.
-
GPQA Diamond (Epoch) scores changed
GPQA Diamond accuracy (0–1) updated.
-
NVIDIA Nemotron models changed
NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16 model removed from availability (was licensed as other, scheduled expiration 2026-06-03).
-
LMArena (text) scores changed
Arena Elo (text, overall) updated.
-
GPQA Diamond (Epoch) scores changed
GPQA Diamond accuracy (0–1) updated.
-
Mistral AI reported benchmarks updated
Mistral Large 3: 7 benchmark claims (via web search)
-
Google Gemma models changed
Removed google/t5gemma-xl-xl-prefixlm-it model with gemma license from available models as of 2025-06-19.