Anthropic — reported benchmarks
Anthropic · source page ↗ · last checked Oct 4, 2026, 12:03 PM
Reported benchmarks · Claude Opus 5.5
captured Oct 4, 2026, 12:03 PM| Benchmark | Score |
|---|---|
| SWE-bench Pro | 89.9% · max effort |
| SWE-bench Multilingual | 93.9% · max effort |
| SWE-bench Multimodal | 61.4% · max effort |
| Terminal-Bench 4.0 | 66.4% · xhigh effort; standard error ±2.6 pts |
| Terminal-Bench-Science 0.1 | 58.7% · max effort |
| Humanity's Last Exam | 67.7% · with tools, max effort |
| Humanity's Last Exam (no tools) | 64.4% · max effort |
| FrontierCode v1.1 (Main) | 54.4% · max effort |
| CursorBench 4.0 | 57.8% · max effort |
| OSWorld 2.0 (Partial Pass) | 81.8% · max effort |
| GDPval-AA v2.1 | 1846 Elo · max effort, 44 occupations |
| AA-Briefcase v1.1 | 1822 Elo · max effort |
| AutomationBench | 40.0% · Zapier evaluation, no fallback models |
| HealthBench Professional | 65.6% · max effort |
| Chartography (with tools) | 89.0% · max effort |
Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.
In the news · Anthropic
- Anthropic Wants to Ban Chatbot Abuse: What Smart People Are Saying
- Anthropic to Ban Abusive Treatment of AI Chatbot Claude
- Anthropic Now Offers A Free Vulnerability-Finding Service For Open-Source Software
- Anthropic's updated usage policy bans sustained cruelty toward its models from November 12
- Anthropic To Ban "Cruel Behavior" Against Claude AI
- Anthropic IPO Forces Question of How to Price Rogue AI Risk
- Anthropic offers free AI security scans to open-source maintainers
- Can you be mean to AI? Anthropic’s new rules explicitly ban cruelty to Claude
Importance-filtered press coverage (Google News) mentioning Anthropic. Headlines link to the original; verify before acting.
Change history
- Vendor claim Oct 4, 2026, 12:03 PM
Anthropic reported benchmarks updated
Claude Opus 5.5: 15 benchmark claims (via web search)
- Vendor claim Sep 27, 2026, 12:02 PM
Anthropic reported benchmarks updated
Claude Opus 5.5: 4 benchmark claims (via web search)
- Vendor claim Sep 20, 2026, 6:02 AM
Anthropic reported benchmarks updated
Claude Opus 4.8: 9 benchmark claims (via web search)
- Vendor claim Sep 13, 2026, 12:03 AM
Anthropic reported benchmarks updated
Claude Opus 5.1: 7 benchmark claims (via web search)
- Vendor claim Sep 5, 2026, 6:02 PM
Anthropic reported benchmarks updated
Claude Opus 5: 7 benchmark claims (via web search)
- Vendor claim Aug 29, 2026, 6:02 PM
Anthropic reported benchmarks updated
Claude Opus 5: 11 benchmark claims (via web search)
- Vendor claim Aug 22, 2026, 12:02 PM
Anthropic reported benchmarks updated
Claude Opus 5: 12 benchmark claims (via web search)
- Vendor claim Aug 15, 2026, 12:02 PM
Anthropic reported benchmarks updated
Claude Opus 5: 5 benchmark claims (via web search)
- Vendor claim Aug 8, 2026, 12:02 AM
Anthropic reported benchmarks updated
Claude Opus 5: 12 benchmark claims (via web search)
- Vendor claim Jul 31, 2026, 6:02 PM
Anthropic reported benchmarks updated
Claude Opus 5: 11 benchmark claims (via web search)
- Vendor claim Jul 24, 2026, 12:02 PM
Anthropic reported benchmarks updated
Claude Opus 4.8: 11 benchmark claims (via web search)
- Vendor claim Jul 17, 2026, 6:02 AM
Anthropic reported benchmarks updated
Claude Opus 4.8: 12 benchmark claims (via web search)
- Vendor claim Jul 10, 2026, 12:02 AM
Anthropic reported benchmarks updated
Claude Opus 4.8: 12 benchmark claims (via web search)
- Vendor claim Jul 2, 2026, 6:02 PM
Anthropic reported benchmarks updated
Claude Opus 4.8: 11 benchmark claims (via web search)
- Vendor claim Jun 25, 2026, 1:18 PM
Anthropic reported benchmarks updated
Claude Opus 4.8: 10 benchmark claims (via web search)