Benchmark

xAI — reported benchmarks

xAI · source page ↗ · last checked Aug 22, 2026, 12:03 PM

Reported benchmarks · Grok 4.6

captured Aug 22, 2026, 12:03 PM
BenchmarkScore
Artificial Analysis Intelligence Index61 composite · nine-benchmark aggregate
CursorBench v3.269.9% · high thinking effort
DeepSWE v1.165.9% · high thinking effort
FrontierCode v1.1 Extended61.3% · extended split
APEX-Agents57.5% · high thinking effort
APEX-SWE56.4% · high thinking effort
Terminal-Bench v3.026% · Grok Build harness
GDPval-AA v21753 Elo
AA-Briefcase1577 Elo · long-horizon analyst work
Harvey Legal Agent Benchmark15.8% · Vals scoring
LiveCodeBench88.4% · xhigh thinking effort
Next.js Evals92%

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

In the news · xAI

Importance-filtered press coverage (Google News) mentioning xAI. Headlines link to the original; verify before acting.

Change history