GLM-5.3
Z.ai / 智谱 GLM · 2026-08(仅精确到月) · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GLM-5.3
GLM-5.3 以「Frontier Coding with Emergent Cyber Capabilities」为定位发布(API thinking 档 low/high/max、默认 max 且不再支持关闭思考;权重承诺发布两周内开源)。已收录 22 项评测集中于编码与网络安全攻防:亮点 Terminal-Bench 2.1 88.2、CyberGym 84.5。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: Terminal Bench 2.1 row · row: Terminal Bench 2.1 · quote_snippet: 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8
{
"harness": "Claude Code 2.1.207",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=65536",
"turn_limit": null,
"time_limit": "6h timeout",
"run_count": null,
"aggregation": null,
"judge": null
}Comparison columns in same row (other protocols/harnesses per their vendors): GLM-5.2 81.0, Kimi K3 88.3, DeepSeek-V4 Pro-0813 87.9, Qwen3.8-Max 86.6, Opus 4.8 85.0, Fable 5 w/ fallback 88.0, GPT-5.6 Sol 88.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: Terminal Bench 3.0 row · row: Terminal Bench 3.0 · quote_snippet: 28.3 | 4.6 | 17.4 | 21.1 | 33.7 | 34.6
{
"harness": "Claude Code 2.1.207",
"tools": "Tool Search disabled",
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "400K context, 128K max output",
"turn_limit": 600,
"time_limit": "10h per rollout",
"run_count": 3,
"aggregation": "avg@3 over three rollouts per task",
"judge": "task's official separate verifier"
}Each rollout in an isolated container from the task's official image. Comparison columns same row: GLM-5.2 4.6, Kimi K3 17.4, Opus 4.8 21.1, Fable 5 w/ fallback 33.7, GPT-5.6 Sol 34.6. Not comparable to Terminal-Bench 2.1 rows (different benchmark version and protocol).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: DeepSWE v1.1 row · row: DeepSWE v1.1 · quote_snippet: 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7
{
"harness": "mini-swe-agent",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 0.95,
"top_p": 1,
"token_budget": "400K context",
"turn_limit": null,
"time_limit": "timeout=6h",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepswe not yet in data/benchmarks/. Comparison columns same row: GLM-5.2 46.2, Kimi K3 67.5, DeepSeek-V4 Pro-0813 62.7, Qwen3.8-Max 56.6, Opus 4.8 58.0, Fable 5 w/ fallback 69.7, GPT-5.6 Sol 72.7. Kimi's own footnote reports 67.3 under mini-SWE-agent; 67.5 vs 67.3 discrepancy between the two vendor pages is unresolved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: NL2Repo row · row: NL2Repo · quote_snippet: 58.0 | 48.9 | 58.0 | 61.1 | 55.9 | 69.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=64k, 1M context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "rule-based checks plus LLM-based judgement (anti-hacking)"
}new-benchmark: nl2repo not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 48.9, Kimi K3 58.0, DeepSeek-V4 Pro-0813 61.1, Qwen3.8-Max 55.9, Opus 4.8 69.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: ProgramBench Almost Solved row · row: ProgramBench Almost Solved · quote_snippet: 19.0 | 9.5 | 17.5 | 10.5 | 15.5 | 33.0 | 23.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: program-bench not yet in data/benchmarks.json (same family as Kimi's Program Bench row). Comparison columns same row: GLM-5.2 9.5, Kimi K3 17.5, Qwen3.8-Max 10.5, Opus 4.8 15.5, Fable 5 w/ fallback 33.0, GPT-5.6 Sol 23.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: FrontierSWE row · row: FrontierSWE · quote_snippet: The evaluation was conducted by Proximal with 1M context length, max effort level
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "1M context, 128K max output",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "dominance score",
"judge": null
}new-benchmark: frontierswe not yet in data/benchmarks/. Evaluation conducted by Proximal per footnote; dominance score reported as of 2026/08/14. Comparison columns same row: GLM-5.2 67.5, Opus 4.8 66.5, Fable 5 w/ fallback 88.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: SWE-Marathon v1.1 row · row: SWE-Marathon v1.1 · quote_snippet: 42.5 | 19.4 | 48.1 | 48.8 | 33.1 | 42.5
{
"harness": "Claude Code 2.1.207",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 0.95,
"token_budget": "max_new_tokens=128000, 1M context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "LLM-based inspection replacing removed anti-cheat pattern checks on strip-clone"
}new-benchmark: swe-marathon not yet in data/benchmarks.json. Footnote discloses protocol deviations: anti-cheat checks removed on strip-clone (LLM inspection instead), extra PyPI index added for parameter-golf and trimul-cuda Docker builds. Kimi's SWE-Marathon row uses an H20-recalibrated branch - variants differ across vendors.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Coding · table: PostTrainBench row · row: PostTrainBench · quote_snippet: weighted average over 3 runs; failed runs fall back to official zero-shot baseline
{
"harness": "Claude Code 2.1.207",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=128000, 1M context",
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "weighted average over 3 runs; failed runs fall back to official zero-shot base-model baseline",
"judge": "LLM agent inspects solutions for external API usage (replaces removed pattern-matching checks)"
}new-benchmark: posttrain-bench not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 31.7, Kimi K3 32.0, Opus 4.8 32.9, Fable 5 w/ fallback 41.8, GPT-5.6 Sol 36.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Cyber + Footnotes / CyberGym · table: CyberGym row · row: CyberGym · quote_snippet: GLM-5.3 scores 84.5%, up from GLM-5.2's 77.2% - the best result on the benchmark
{
"harness": "Claude Code 2.1.207",
"tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=128000",
"turn_limit": null,
"time_limit": "unlimited timeout per task",
"run_count": 1,
"aggregation": "single-run Pass@1 over 1,507 tasks",
"judge": null
}new-benchmark: cybergym not yet in data/benchmarks.json. Agent placed inside the task container; Git info removed. Comparison columns same row: GLM-5.2 77.2, Kimi K3 80.0, DeepSeek-V4 Pro-0813 83.3, Qwen3.8-Max 78.5, Opus 4.8 78.1, Fable 5 w/ fallback 83.8, GPT-5.6 Sol 83.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Cyber + Footnotes / ExploitGym · table: ExploitGym 2h / 6h row · row: ExploitGym 2h / 6h · quote_snippet: GLM-5.3 completes 105 tasks within two hours
{
"harness": "Claude Code 2.1.207",
"tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=128000",
"turn_limit": null,
"time_limit": "2h budget (API inference time rescaled by per-model TPS plus non-API overhead)",
"run_count": 1,
"aggregation": "single-run Pass@1 on 869 tasks",
"judge": null
}new-benchmark: exploitgym not yet in data/benchmarks.json. Time budgets normalized across models via per-model tokens-per-second (GLM-5.3 115 TPS, Kimi K3 40 TPS, Qwen3.8 Max 47 TPS per footnote). Comparison columns same row 2h values: GLM-5.2 29, Kimi K3 36, Qwen3.8-Max 14, Opus 4.8 80, Fable 5 w/ fallback 181, GPT-5.6 Sol 216.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Cyber + Footnotes / ExploitGym · table: ExploitGym 2h / 6h row · row: ExploitGym 2h / 6h · quote_snippet: and 130 within six hours
{
"harness": "Claude Code 2.1.207",
"tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=128000",
"turn_limit": null,
"time_limit": "6h budget (API inference time rescaled by per-model TPS plus non-API overhead)",
"run_count": 1,
"aggregation": "single-run Pass@1 on 869 tasks",
"judge": null
}new-benchmark: exploitgym not yet in data/benchmarks/. Comparison columns same row 6h values: GLM-5.2 39, Kimi K3 70, Qwen3.8-Max 26, Opus 4.8 120, Fable 5 w/ fallback 247, GPT-5.6 Sol 293.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Cyber + Footnotes / ExploitBench · table: ExploitBench row · row: ExploitBench · quote_snippet: GLM-5.3 reaches 54.4%, more than doubling GLM-5.2's 24.4%
{
"harness": "Claude Code 2.1.207",
"tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=128000",
"turn_limit": 300,
"time_limit": null,
"run_count": 3,
"aggregation": "average coverage score over 41 tasks across 3 revisions; per-task result is union of capabilities across revisions",
"judge": null
}new-benchmark: exploitbench not yet in data/benchmarks/. Comparison columns same row: GLM-5.2 24.4, Kimi K3 32.2, Qwen3.8-Max 28.8, Opus 4.8 40.0, Fable 5 w/ fallback 78.0, GPT-5.6 Sol 76.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Agentic + Footnotes / Toolathlon Verified · table: Toolathlon Verified row · row: Toolathlon Verified · quote_snippet: official evaluation service; pass@1 averaged over 3 independent runs
{
"harness": "official evaluation service",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "pass@1 averaged over 3 independent runs",
"judge": null
}new-benchmark: toolathlon not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 59.9, Kimi K3 76.5, DeepSeek-V4 Pro-0813 74.1, Qwen3.8-Max 72.5, Opus 4.8 76.2, Fable 5 w/ fallback 74.7, GPT-5.6 Sol 74.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Agentic + Footnotes / AutomationBench · table: AutomationBench v1.0.6 row · row: AutomationBench v1.0.6 · quote_snippet: AutomationBench v1.0.6, incorporating the fix ... in PR #13
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: automationbench not yet in data/benchmarks.json. Version pin v1.0.6 with zapier/AutomationBench PR #13 null-type fix. Comparison columns same row: GLM-5.2 26.2, Kimi K3 46.7, DeepSeek-V4 Pro-0813 43.2, Qwen3.8-Max 39.8, Opus 4.8 41.0, Fable 5 w/ fallback 46.2, GPT-5.6 Sol 45.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Agentic + Footnotes / Agent's Last Exam (CLI) · table: Agents' Last Exam ALE-CLI row · row: Agents' Last Exam ALE-CLI · quote_snippet: official evaluation protocol with the Claude Code harness ... 105 tasks
{
"harness": "Claude Code",
"tools": "Tool Search disabled",
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "1M context, 64K maximum output",
"turn_limit": null,
"time_limit": "default 4h, task-specific up to 8h",
"run_count": null,
"aggregation": null,
"judge": "official ALE evaluators"
}new-benchmark: agents-last-exam not yet in data/benchmarks.json. 105 tasks, each in isolated Docker container using Task Card resources. Comparison columns same row: GLM-5.2 23.8, Kimi K3 27.6, DeepSeek-V4 Pro-0813 25.7, Qwen3.8-Max 27.0, Opus 4.8 25.7, Fable 5 w/ fallback 23.8, GPT-5.6 Sol 28.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Agentic + Footnotes / HLE w/ tools · table: HLE w/ Tools row · row: HLE w/ Tools · quote_snippet: temperature=1.0 and top_p=0.95 ... maximum context length of 300,000 tokens
{
"harness": null,
"tools": "tools enabled; context management strategy",
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "max generation length 163,840 tokens; max context 300,000 tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "GPT-5.6-luna (medium)"
}Maps to existing benchmark hlehle (Humanity's Last Exam), tool-augmented variant with LLM judge. Comparison columns same row: GLM-5.2 54.7, Kimi K3 59.8, DeepSeek-V4 Pro-0813 60.0, Qwen3.8-Max 56.2, Opus 4.8 57.9, Fable 5 w/ fallback 63.9, GPT-5.6 Sol 64.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance across comparison models / Agentic + Footnotes / GDPval-AA v2 · table: GDPval-AA v2 row · row: GDPval-AA v2 · quote_snippet: GDPval-AA v2: Models are evaluated by Artificial Analysis.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gdpval-aa not yet in data/benchmarks.json. Footnote states all models in this row are evaluated by Artificial Analysis, so this is a third-party measurement transcribed in the vendor table, not a vendor-run score. Comparison columns same row: GLM-5.2 1508, Kimi K3 1682, DeepSeek-V4 Pro-0813 1590, Qwen3.8-Max 1739, Opus 4.8 1588, Fable 5 w/ fallback 1743, GPT-5.6 Sol 1730.
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Emergent Cyber Capability · row: CyberGym · quote_snippet: ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cybergym not yet in data/benchmarks/. Competitor score cited by Z.ai; protocol under which Mythos 5 was run is not disclosed on this page, so the comparison is not proven same-protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Emergent Cyber Capability · row: CyberGym · quote_snippet: ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cybergym not yet in data/benchmarks/. Competitor score cited by Z.ai; GPT-5.6 Sol run protocol not disclosed on this page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Emergent Cyber Capability · row: ExploitBench · quote_snippet: Mythos 5 and GPT-5.6 Sol score 78.0% and 76.5%, respectively
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitbench not yet in data/benchmarks/. Competitor score cited by Z.ai; protocol for the cited run not disclosed on this page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Emergent Cyber Capability · row: ExploitBench · quote_snippet: Mythos 5 and GPT-5.6 Sol score 78.0% and 76.5%, respectively
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitbench not yet in data/benchmarks/. Competitor score cited by Z.ai; protocol for the cited run not disclosed on this page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Emergent Cyber Capability · row: ExploitGym · quote_snippet: Mythos 5 remains well ahead at 181 and 247 tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "2h and 6h budgets",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: exploitgym not yet in data/benchmarks/. Competitor score cited by Z.ai for both budget variants (181 tasks at 2h, 247 at 6h); value field stores the 2h figure, display shows both.