Models

Frontier & open-weight models compared by capability — columns are flagship models, rows are the headline benchmarks that measure each one. Every number is the published figure, restated with a link to the source it came from · as of 2026-06-14. Curated matrix is 116d old — a refresh is due; live leaderboard moves still flow into Model changes below.

We cite the benchmark authorities rather than re-rank them — Epoch AI, LMArena and Artificial Analysis. ChangeRadar's job is the changes: what gets released, repriced, deprecated, or quietly shifts behind a stable model id.

Vendor champions · best flagship per vendor · GPQA Diamond

Google
Gemini 3.1 Pro
94.3GPQA Diamond
Moonshot AI reported
Kimi K3
93.5GPQA Diamond
Source ↗
Alibaba reported
Qwen3.8-Max
92.6GPQA Diamond
Source ↗
DeepSeek reported
DeepSeek-V4.1-Flash
90.9GPQA Diamond
Source ↗
Anthropic reported
Claude Opus 5.5
89.9SWE-bench Pro
Source ↗
OpenAI reported
GPT-5
88.4GPQA Diamond
Source ↗
Zhipu AI (Z.ai) reported
GLM-5.3
80.5LiveCodeBench
Source ↗
Meta open
Llama 4 Maverick
69.8GPQA Diamond
xAI reported
Grok 4.7
37.6Terminal-Bench
Source ↗
Mistral AI reported
Mistral Large 4
28.3Terminal-Bench
Source ↗

Each vendor's most recent publicly-available flagship, ordered by score (highest first) — Google · OpenAI · Anthropic · xAI highlighted. The weekly market-watch surfaces new releases automatically; one tagged reported is the latest release shown with vendor-reported scores (linked to source) until we independently cite it. Score = GPQA Diamond; every number links to its source.

Compare two models

2 wins · Claude Fable 5 1 wins · Gemini 3.1 Pro 0 ties

Agentic coding

BenchmarkClaude Fable 5Gemini 3.1 ProΔ
SWE-bench Verified i% resolved (pass@1) 95 80.6 +14.4
SWE-bench Pro i% resolved (pass@1) 80.3 54.2 +26.1
3 coverage gaps — only one model reports these
  • Terminal-Bench — Gemini 3.1 Pro 68.5
  • LiveCodeBench — Gemini 3.1 Pro 2887
  • FrontierCode — Claude Fable 5 29.3

Tool use & agents

5 coverage gaps — only one model reports these
  • TAU-bench — Gemini 3.1 Pro 99.3
  • OSWorld — Claude Fable 5 85
  • BrowseComp — Gemini 3.1 Pro 85.9
  • GDPval-AA — Claude Fable 5 1932
  • MCP Atlas — Gemini 3.1 Pro 69.2

Science & reasoning

BenchmarkClaude Fable 5Gemini 3.1 ProΔ
GPQA Diamond i% accuracy 92.6 94.3 -1.7
2 coverage gaps — only one model reports these
  • Humanity's Last Exam — Gemini 3.1 Pro 44.4
  • ARC-AGI-2 — Gemini 3.1 Pro 77.1

General knowledge

1 coverage gap — only one model reports these
  • MMMLU — Gemini 3.1 Pro 92.6

Multimodal

2 coverage gaps — only one model reports these
  • MMMU — Gemini 3.1 Pro 80.5
  • GDP.pdf — Claude Fable 5 29.8

Long context

1 coverage gap — only one model reports these
  • MRCR — Gemini 3.1 Pro 84.9

Shared benchmarks first (with Δ when both report the same scale); one-sided coverage collapses below. Pick any two models — or a champion above — and the URL becomes shareable.

All models · benchmark matrix

Vendor
Capability
21 of 21 columns shown

reported columns (Claude Opus 5.5, GPT-5, Grok 4.7, Qwen3.8-Max, DeepSeek-V4.1-Flash, Mistral Large 4, Kimi K3, GLM-5.3) are auto-discovered by our weekly market-watch from each vendor's own reported numbers — not independently verified, and shown when a vendor ships a model newer than the hand-cited column beside it. Full claim sets are in Vendor-reported benchmarks below.

Agentic coding

Can the model fix real bugs, ship features, and operate a dev environment end-to-end as a coding agent — the single most-watched capability in 2026 vendor launches.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
SWE-bench Verified i % resolved (pass@1) 95 — 88.6 — — 74.9 80.6 — — 80.6 — 80.4 — 80.2 76.8 — — — 77.6 — —
SWE-bench Pro i % resolved (pass@1) 80.3 89.9 69.2 77.8 58.6 — 54.2 — — — — 60.6 67.7 58.6 — 62.1 — — — — —
Terminal-Bench i % solved (pass@1) — 66.4 74.6 88 82.7 — 68.5 — 37.6 67.9 90.6 69.7 86.6 66.7 88.3 81 28.3 — — 28.3 —
LiveCodeBench i % pass@1 — — — — — — 2887 — — 93.5 — — — 89.6 — — 80.5 43.4 — — 29.7

Tool use & agents

Beyond writing code: can the model select and chain tools, follow policy, drive a computer/browser, and complete long-horizon multi-step tasks.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
TAU-bench i % pass / pass^k — — — — — — 99.3 98 — — — — — — — — — — — — —
OSWorld i % success 85 81.8 83.4 — 78.7 — — — — — — — 86.1 73.1 — — — — — — —
BrowseComp i % accuracy — — 84.3 — 90.1 — 85.9 — — — — — — 83.2 91.2 — — — — — —

Math

Competition and research-level mathematical reasoning, increasingly reported on uncontaminated/post-cutoff problem sets.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
HMMT i % accuracy (pass@1) — — — — — — — — — — — 97.1 — 92.7 — — — — — — —
MATH i % accuracy — — — — — — — — — — — — — — — — — 61.2 — — 89

Science & reasoning

Expert-level, Google-proof reasoning across the sciences and broad academia — the benchmarks vendors point to when claiming 'PhD-level' or 'frontier' reasoning.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
GPQA Diamond i % accuracy 92.6 — 93.6 — 93.6 88.4 94.3 90.1 — 90.1 90.9 92.4 92.6 90.5 93.5 91.2 — 69.8 — — 42.4
Humanity's Last Exam i % accuracy — 67.7 57.9 64.5 57.2 — 44.4 — — 37.7 — 41.4 43.6 54 43.5 40.5 62.5 — — — —
ARC-AGI-2 i % solved — — — — 85 — 77.1 — — — — — — — — — — — — — —

General knowledge

Broad multi-subject factual and reasoning coverage; the classic 'how much does it know' bucket, now reported via the harder Pro variant since base MMLU is saturated.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
MMLU-Pro i % accuracy — — — — — — — — — 87.5 — — — — — — — 80.5 — — 67.5

Multimodal

Vision + language understanding and visual reasoning — how well the model interprets images, diagrams, charts and figures.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
MMMU i % accuracy — — — — 83.2 84.2 80.5 — — — — — — 79.4 — — — 73.4 — — 64.9

Long context

Retrieval and reasoning quality as context length grows into the hundreds-of-thousands / millions of tokens — beyond simple needle-in-a-haystack.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
MRCR i % accuracy — — — — 74 — 84.9 — — — — — — — — — — — — — —

Human preference

Aggregate real-user preference from blind head-to-head comparisons — the closest thing to a 'do people actually like the answers' metric, and the one number vendors love to top.

Benchmark Fable 5 Opus 5.5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5 reported Gemini 3.1 Pro Grok 4.3 Grok 4.7 reported DeepSeek-V4-Pro open DeepSeek-V4.1-Flash reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open GLM-5.3 reported Llama 4 Maverick open Mistral Medium 3.5 open Mistral Large 4 reported Gemma 3 27B open
LMArena Elo i Elo — — — 1458 1474 — — — — — — — — 1466 — — — 1417 — — 1338

Showing 13 flagship models across the headline benchmarks; 53 models and 84 benchmarks tracked in total from 60 primary & aggregator sources. Claude Mythos 5 numbers are limited (access-restricted preview). Numbers are published facts, restated with a per-cell source link; vendor benchmark charts are linked to their source, not rehosted.
Sources include: 9to5Google (Gemini 3 Flash launch coverage) · AIFire (citing OpenAI GPT-5.2 release) · Artificial Analysis · BenchLM.ai (citing LMArena) · BinaryVerse AI (xAI official figures) · BuildFastWithAI (citing OpenAI GPT-5.5 launch) · Caylent · Codersera (reporting Moonshot's figures) · DataCamp · DataCamp (reproducing Meta's Llama 4 launch chart) · DeepSeek-AI (arXiv 2512.02556) · DeepSeek-AI (Hugging Face model card) · Google (official Gemini 2.5 launch blog) · Google (official Gemini 3 Flash launch blog) · Google (official Gemini 3 launch blog) · Google DeepMind (Gemini 3.1 Pro model card) · +36 more.

Vendor-reported benchmarks

Numbers as claimed by the vendor on their own model/system card — not independently verified and often measured with a favourable harness. We track each vendor's claims over time and link to the source; cross-check against the cited matrix above.

Claude Opus 5.5 Anthropic
SWE-bench Pro · max effort 89.9%
SWE-bench Multilingual · max effort 93.9%
SWE-bench Multimodal · max effort 61.4%
Terminal-Bench 4.0 · xhigh effort; standard error ±2.6 pts 66.4%
Terminal-Bench-Science 0.1 · max effort 58.7%
Humanity's Last Exam · with tools, max effort 67.7%
Humanity's Last Exam (no tools) · max effort 64.4%
FrontierCode v1.1 (Main) · max effort 54.4%
CursorBench 4.0 · max effort 57.8%
OSWorld 2.0 (Partial Pass) · max effort 81.8%
GDPval-AA v2.1 · max effort, 44 occupations 1846 Elo
AA-Briefcase v1.1 · max effort 1822 Elo
vendor card ↗
GPT-5 OpenAI
AIME 2025 · without tools 94.6%
SWE-bench Verified · fixed subset of 477 verified tasks 74.9%
GPQA Diamond · GPT-5 pro variant, without tools 88.4%
MMMU 84.2%
Aider Polyglot 88%
HealthBench Hard 46.2%
vendor card ↗
Grok 4.7 xAI
CursorBench 4.0 · xhigh effort 46.3%
DeepSWE v1.1 · high-effort score 71.0%
EEBench · electrical engineering 64.0%
AA Briefcase v1.1 · xhigh effort 1657 Elo
Terminal-Bench 4.0 · agentic coding, xhigh effort 37.6%
Harvey Legal Agent Benchmark 19.6%
HealthBench Professional 56.7%
GDPval · xhigh effort 1695 Elo
LatchBio Biosafety Benchmark · dual-use domains 62.4%
HackerBench v0.3 · risky dual-use prompts allowed through 3.3%
vendor card ↗
Gemini 3.1 Pro Google
GPQA Diamond · No tools 94.3%
ARC-AGI-2 · ARC Prize Verified 77.1%
SWE-Bench Verified · Single attempt 80.6%
SWE-Bench Pro (Public) · Single attempt 54.2%
Terminal-Bench 2.0 · Terminus-2 harness 68.5%
Humanity's Last Exam · No tools 44.4%
Humanity's Last Exam · Search (blocklist) + Code 51.4%
LiveCodeBench Pro 2887 Elo
MATH 95.1%
vendor card ↗
Qwen3.8-Max Alibaba
Terminal Bench 2.1 86.6 %
SWE-bench Pro 67.7 %
GPQA Diamond 92.6 %
PaperBench 93.0 %
IFBench 82.8 %
FrontierSWE 73.5 %
DeepSWE 1.1 56.6 %
OSWorld-Verified · multimodal/agentic 86.1 %
QwenSWEBench · internal benchmark 80.7 %
NL2Repo-Bench 55.9 %
MLS-Bench-Lite 41.0 %
Humanity's Last Exam · text-only subset 43.6 %
vendor card ↗
DeepSeek-V4.1-Flash DeepSeek
GPQA Diamond · Pass@1 90.9%
Terminal-Bench 2.1 90.6%
Codeforces 3471 Rating
DeepSWE v1.1 74.2%
NL2Repo-Bench 65.4%
CyberGym 88.1%
HLE · without tools; 39.1% with tools 36.8%
MathArena Apex 65.6%
Terminal-Bench 3.0 30.0%
Terminal-Bench 4.0 31.2%
Agents' Last Exam 31.8%
vendor card ↗
Llama 4 Maverick Meta
MMLU Pro · instruction-tuned, 0-shot 80.5%
GPQA Diamond · instruction-tuned, 0-shot 69.8%
MMMU · instruction-tuned, image reasoning, 0-shot 73.4%
MathVista · instruction-tuned, 0-shot 73.7%
ChartQA · instruction-tuned, 0-shot relaxed accuracy 90.0%
DocVQA · instruction-tuned, 0-shot anls metric 94.4%
LiveCodeBench · instruction-tuned, 0-shot pass@1, 10.01.2024-02.01.2025 43.4%
MATH · pre-trained, 4-shot em_maj1@1 61.2%
MBPP · pre-trained, 3-shot pass@1 77.6%
Multilingual MMLU · instruction-tuned 84.6%
MTOB (half-book) · instruction-tuned, eng->kgv/kgv->eng long context 54.0/46.4 chrF
vendor card ↗
Mistral Large 4 Mistral AI
DeepSWE v1.1 61.7%
SWE-Atlas-QnA 59.4%
Terminal-Bench 4 28.3%
Coding Agent Index 49.8%
AutomationBench · 657 business workflows across Gmail, Google Sheets, Slack, Salesforce 59.9%
Cybench · 40 security challenges 93%
CyberGym-E2E-AA · Reproduce and patch vulnerability test 82%
Lakera B3 Attack Resistance 93.3%
AA-Briefcase 1393 Elo
Dense200 · Visual grounding benchmark 42%
vendor card ↗
Kimi K3 Moonshot AI
GPQA Diamond · max reasoning effort, temperature=1.0 93.5%
SWE-Bench Verified · max reasoning effort 76.8%
AIME 2025 · max reasoning effort, from K2.5 card; K3 reports AIME 2026 at 96.4% 96.1%
SWE-Bench Multilingual · from K2.5; K3 model card shows 76.7% 73.0%
Terminal-Bench 2.1 · Kimi Code harness, max reasoning 88.3%
FrontierSWE · max reasoning, Kimi Code harness 81.2 dominance score
SWE-Marathon · max reasoning effort 42.0%
BrowseComp · max reasoning, compaction at 300K 91.2%
DeepSWE · Kimi Code harness; 67.3% with mini-SWE-agent 67.5%
Humanity's Last Exam (HLE-Full) · max reasoning, no tools; 56.0% with tools 43.5%
DeepSearchQA · max reasoning effort 95.0 F1
Program Bench · raw hidden-test pass rate, max reasoning 77.8%
vendor card ↗
GLM-5.3 Zhipu AI (Z.ai)
Terminal-Bench 3.0 · max reasoning effort, avg@3 28.3 pass@1
DeepSWE v1.1 · max reasoning effort 66.9 %
Agents' Last Exam · CLI, max reasoning effort 28.5 %
Terminal-Bench 2.1 · temperature=1.0, top_p=1, 65536 max_new_tokens 88.2 %
Humanity's Last Exam · with tools, 300K maximum context 62.5 %
CyberGym · Pass@1, max reasoning effort, 1,507 tasks 84.5 %
SWE-Marathon v1.1 · long-horizon coding 42.5 %
FrontierSWE v2 · Max effort, Proximal evaluation 78.1 %
LiveCodeBench · Vals AI run 80.5 %
GDPval-AA v2 · Economic knowledge work evaluation 1769 Elo
vendor card ↗

Model changes

New releases, deprecations, and benchmark score moves we’ve recorded — newest first.