Benchmark

Anthropic — reported benchmarks

Anthropic · source page ↗ · last checked Oct 4, 2026, 12:03 PM

Reported benchmarks · Claude Opus 5.5

captured Oct 4, 2026, 12:03 PM
BenchmarkScore
SWE-bench Pro89.9% · max effort
SWE-bench Multilingual93.9% · max effort
SWE-bench Multimodal61.4% · max effort
Terminal-Bench 4.066.4% · xhigh effort; standard error ±2.6 pts
Terminal-Bench-Science 0.158.7% · max effort
Humanity's Last Exam67.7% · with tools, max effort
Humanity's Last Exam (no tools)64.4% · max effort
FrontierCode v1.1 (Main)54.4% · max effort
CursorBench 4.057.8% · max effort
OSWorld 2.0 (Partial Pass)81.8% · max effort
GDPval-AA v2.11846 Elo · max effort, 44 occupations
AA-Briefcase v1.11822 Elo · max effort
AutomationBench40.0% · Zapier evaluation, no fallback models
HealthBench Professional65.6% · max effort
Chartography (with tools)89.0% · max effort

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

In the news · Anthropic

Importance-filtered press coverage (Google News) mentioning Anthropic. Headlines link to the original; verify before acting.

Change history