Benchmark

Moonshot Kimi — reported benchmarks

Moonshot AI · source page ↗ · last checked Oct 4, 2026, 12:05 PM

Reported benchmarks · Kimi K3

captured Oct 4, 2026, 12:05 PM
BenchmarkScore
GPQA Diamond93.5% · max reasoning effort, temperature=1.0
SWE-Bench Verified76.8% · max reasoning effort
AIME 202596.1% · max reasoning effort, from K2.5 card; K3 reports AIME 2026 at 96.4%
SWE-Bench Multilingual73.0% · from K2.5; K3 model card shows 76.7%
Terminal-Bench 2.188.3% · Kimi Code harness, max reasoning
FrontierSWE81.2 dominance score · max reasoning, Kimi Code harness
SWE-Marathon42.0% pass@1 · max reasoning effort
BrowseComp91.2% · max reasoning, compaction at 300K
DeepSWE67.5% pass@1 · Kimi Code harness; 67.3% with mini-SWE-agent
Humanity's Last Exam (HLE-Full)43.5% pass@1 · max reasoning, no tools; 56.0% with tools
DeepSearchQA95.0 F1 · max reasoning effort
Program Bench77.8% pass@1 · raw hidden-test pass rate, max reasoning

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

Change history