Benchmark

Zhipu GLM — reported benchmarks

Zhipu AI (Z.ai) · source page ↗ · last checked Oct 4, 2026, 12:05 PM

Reported benchmarks · GLM-5.3

captured Oct 4, 2026, 12:05 PM
BenchmarkScore
Terminal-Bench 3.028.3 pass@1 · max reasoning effort, avg@3
DeepSWE v1.166.9 % · max reasoning effort
Agents' Last Exam28.5 % · CLI, max reasoning effort
Terminal-Bench 2.188.2 % · temperature=1.0, top_p=1, 65536 max_new_tokens
Humanity's Last Exam62.5 % · with tools, 300K maximum context
CyberGym84.5 % · Pass@1, max reasoning effort, 1,507 tasks
SWE-Marathon v1.142.5 % · long-horizon coding
FrontierSWE v278.1 % · Max effort, Proximal evaluation
LiveCodeBench80.5 % · Vals AI run
GDPval-AA v21769 Elo · Economic knowledge work evaluation

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

Change history