Benchmark

Meta Llama — reported benchmarks

Meta · source page ↗ · last checked Aug 23, 2026, 6:02 AM

Reported benchmarks · Llama 4 Maverick

captured Aug 23, 2026, 6:02 AM
BenchmarkScore
MMLU Pro80.5 % · 5-shot, CoT
GPQA Diamond69.8 % · 0-shot, CoT, multiple generations averaged
LiveCodeBench43.4 % · 10.01.2024-02.01.2025, multiple generations averaged
MMMU73.4 % · Image reasoning
MathVista73.7 % · Image reasoning
ChartQA90.0 % · Image understanding
DocVQA94.4 % · Image understanding, test set
Multilingual MMLU84.6 % · Multilingual
MTOB Half Book54.0 / 46.4 % · Long context, eng->kgv/kgv->eng
MTOB Full Book50.8 / 46.7 % · Long context, eng->kgv/kgv->eng

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

Change history