← 模型目录

GLM-5.2

Z.ai / 智谱 GLM · 2026-06-16 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GLM-5.2

智谱将 GLM-5.2 定位为面向长周期任务构建的模型,提供稳固支撑长程工作的 1M 上下文与更强编码能力(High/Max 档位控制),权重以 MIT 协议完全开源。已收录评测聚焦智能体编码与工具调用、长程软件工程与数学推理等领域,AIME 2026 99.2、Terminal-Bench 2.1(Terminus-2)81。

输入模态
文本
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 40.5 模型 glm-5.2 · 版本 text-only subset (default) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HLE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max generation 163,840 tokens",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5.1 31.0, Qwen3.7-Max 41.4, MiniMax M3 37.0, DeepSeek-V4-Pro 37.7, Opus 4.8 49.8*, GPT-5.5 41.4*, Gemini 3.1 Pro 45.0 (* = full-set scores, not same subset).

打开官方来源

hlehle 54.7 模型 glm-5.2 · 版本 w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HLE w/ Tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "300,000-token max context, no context management",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5.1 52.3, Qwen3.7-Max 53.5, DeepSeek-V4-Pro 48.2, Opus 4.8 57.9*, GPT-5.5 52.2*, Gemini 3.1 Pro 51.4*.

打开官方来源

critpt 20.9 模型 glm-5.2 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: CritPt

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: critpt not yet in data/benchmarks/. Competitor cells: GLM-5.1 4.6, Qwen3.7-Max 13.4, M3 3.7, V4-Pro 12.9, Opus 4.8 20.9, GPT-5.5 27.1, Gemini 3.1 Pro 17.7.

打开官方来源

aime-26 99.2 模型 glm-5.2 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: AIME 2026

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "GPT-5.5 (medium) with mandated Explanation/Exact Answer/Confidence format"
}

new-benchmark: aime-26 (also introduced in kimi-k2-6.json this batch). Competitor cells: GLM-5.1 95.3, Qwen3.7-Max 97.0, V4-Pro 94.6, Opus 4.8 95.7, GPT-5.5 98.3, Gemini 3.1 Pro 98.2.

打开官方来源

hmmt25 94.4 模型 glm-5.2 · 版本 November 2025 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HMMT Nov. 2025

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

HMMT Nov 2025 mapped to hmmt-25 with edition variant. Competitor cells: GLM-5.1 94.0, Qwen3.7-Max 95.0, M3 84.4, V4-Pro 94.4, Opus 4.8 96.5, GPT-5.5 96.5, Gemini 3.1 Pro 94.8.

打开官方来源

hmmt-26 92.5 模型 glm-5.2 · 版本 February 2026 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HMMT Feb. 2026

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hmmt-26 not yet in data/benchmarks/. Competitor cells: GLM-5.1 82.6, Qwen3.7-Max 97.1, M3 84.4, V4-Pro 95.2, Opus 4.8 96.7, GPT-5.5 96.7, Gemini 3.1 Pro 87.3.

打开官方来源

imo-answerbench 91 模型 glm-5.2 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: IMOAnswerBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

imo-answerbench id already introduced by prior batches. Competitor cells: GLM-5.1 83.8, Qwen3.7-Max 90.0, V4-Pro 89.8, Opus 4.8 83.5, Gemini 3.1 Pro 81.0.

打开官方来源

gpqa 91.2 模型 glm-5.2 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: GPQA-Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5.1 86.2, Qwen3.7-Max 90.0, M3 93.0, V4-Pro 90.1, Opus 4.8 93.6, GPT-5.5 93.6, Gemini 3.1 Pro 94.3.

打开官方来源

swebench-pro 62.1 模型 glm-5.2 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: SWE-bench Pro

{
  "harness": "OpenHands with tailored instruction prompt",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens 32k, 400K context window",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose cross-check: '62.1 vs. 58.4 on SWE-bench Pro' over GLM-5.1. Competitor cells: GLM-5.1 58.4, Qwen3.7-Max 60.6, M3 59.0, V4-Pro 55.4, Opus 4.8 69.2, GPT-5.5 58.6, Gemini 3.1 Pro 54.2.

打开官方来源

nl2repo 48.9 模型 glm-5.2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: NL2Repo

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens 48k, 400k context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

nl2repo id already introduced by prior batches. Anti-hacking: rule-based + LLM-based interception. Competitor cells: GLM-5.1 42.7, Qwen3.7-Max 47.2, M3 42.1, V4-Pro 35.5, Opus 4.8 69.7, GPT-5.5 50.7, Gemini 3.1 Pro 33.4.

打开官方来源

deepswe 46.2 模型 glm-5.2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: DeepSWE

{
  "harness": "official pier evaluation framework + mini-swe-agent",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "400K context",
  "turn_limit": null,
  "time_limit": "2h timeout",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

deepswe id already introduced by prior batches. Competitor cells: GLM-5.1 18.0, Qwen3.7-Max 18.0, M3 20.0, V4-Pro 8.0, Opus 4.8 58.0, GPT-5.5 70.0, Gemini 3.1 Pro 10.0.

打开官方来源

program-bench 63.7 模型 glm-5.2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: ProgramBench

{
  "harness": "Claude-Code 2.1.156",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_tokens 64,000, max_turns 2000, 6h sample timeout, 400K context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

program-bench id already introduced by prior batches (200 instances). Competitor cells: GLM-5.1 50.9, V4-Pro 47.8, Opus 4.8 71.9, GPT-5.5 70.8, Gemini 3.1 Pro 39.5.

打开官方来源

terminalbench 81 模型 glm-5.2 · 版本 2.1, Terminus-2 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: Terminal Bench 2.1 Terminus-2

{
  "harness": "Terminus-2, parser=json",
  "tools": "resource caps 4 CPU / 8GB RAM",
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens 48k, 256K context, 500 max episodes",
  "turn_limit": null,
  "time_limit": "4h timeout",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose cross-check: '81.0 vs. 63.5 on Terminal-Bench 2.1' over GLM-5.1; 'within a few points of Claude Opus 4.8 (85.0)'. Competitor cells: GLM-5.1 63.5, Qwen3.7-Max 75.0, M3 65.0, V4-Pro 64.0, Opus 4.8 85.0, GPT-5.5 84.0, Gemini 3.1 Pro 74.0.

打开官方来源

terminalbench 82.7 (Claude Code) 模型 glm-5.2 · 版本 2.1, Claude Code (best reported harness) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: Terminal Bench 2.1 Best Reported Harness

{
  "harness": "Claude Code 2.1.167, max_new_tokens overridden to 128k via transparent proxy",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "wall-clock limits removed; per-task CPU/memory constraints preserved",
  "run_count": 5,
  "aggregation": "mean over 5 runs",
  "judge": null
}

Best-across-harness row; per-cell harness in parentheses: GLM-5.1 69.0 (Claude Code), Opus 4.8 78.9 (Claude Code), GPT-5.5 83.4 (Codex), Gemini 3.1 Pro 70.7 (Gemini CLI). Different harness per cell - not same-protocol.

打开官方来源

frontierswe 74.4 模型 glm-5.2 · 版本 dominance as of 2026/06/16 · 指标 dominance_score · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: FrontierSWE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "1M context, 128K max output",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

frontierswe id already introduced by prior batches. Footnote: evaluated by Proximal (third party). Prose: 'trails Opus 4.8 by only 1%, edging out GPT-5.5 by 1% and Opus 4.7 by 11%'. Competitor cells: GLM-5.1 30.5, V4-Pro 29.0, Opus 4.8 75.1, GPT-5.5 72.6, Gemini 3.1 Pro 39.6.

打开官方来源

posttrain-bench 34.3 模型 glm-5.2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: PostTrainBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "1M context, 128K max output",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

posttrain-bench id already introduced by prior batches. Footnote: evaluated by PostTrainBench (each agent given an H100 GPU). Prose: 'outperforms both Opus 4.7 and GPT-5.5, ranking second only to Opus 4.8'. Competitor cells: GLM-5.1 20.1, Opus 4.8 37.2, GPT-5.5 28.4, Gemini 3.1 Pro 21.6.

打开官方来源

swe-marathon 13 模型 glm-5.2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: SWE-Marathon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "1M context, 128K max output",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

swe-marathon id already introduced by prior batches. Footnote: evaluated by Abundant AI (third party). Prose: 'trailing Opus 4.8 by 13% while remaining second only to the Opus series'. Competitor cells: GLM-5.1 1.0, Opus 4.8 26.0, GPT-5.5 12.0, Gemini 3.1 Pro 4.0.

打开官方来源

mcp-atlas 76.8 模型 glm-5.2 · 版本 500-task public set · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: MCP-Atlas Public Set

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "10-minute timeout per task",
  "run_count": null,
  "aggregation": null,
  "judge": "Gemini-3.0-Pro"
}

mcp-atlas id already introduced by prior batches. Competitor cells: GLM-5.1 71.8, Qwen3.7-Max 76.4, M3 74.2, V4-Pro 73.6, Opus 4.8 77.8, GPT-5.5 75.3, Gemini 3.1 Pro 69.2.

打开官方来源

toolathlon 48.2 模型 glm-5.2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: Tool-Decathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "max_token 128K via official evaluation service",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tool-decathlon not yet in data/benchmarks/ (distinct from toolathlon). Competitor cells: GLM-5.1 40.7, Qwen3.7-Max -, V4-Pro 52.8, Opus 4.8 59.9, GPT-5.5 55.6, Gemini 3.1 Pro 48.8.

打开官方来源