GLM-5.2
Z.ai / 智谱 GLM · 2026-06-16 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GLM-5.2
智谱将 GLM-5.2 定位为面向长周期任务构建的模型,提供稳固支撑长程工作的 1M 上下文与更强编码能力(High/Max 档位控制),权重以 MIT 协议完全开源。已收录评测聚焦智能体编码与工具调用、长程软件工程与数学推理等领域,AIME 2026 99.2、Terminal-Bench 2.1(Terminus-2)81。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HLE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "max generation 163,840 tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5.1 31.0, Qwen3.7-Max 41.4, MiniMax M3 37.0, DeepSeek-V4-Pro 37.7, Opus 4.8 49.8*, GPT-5.5 41.4*, Gemini 3.1 Pro 45.0 (* = full-set scores, not same subset).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HLE w/ Tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "300,000-token max context, no context management",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5.1 52.3, Qwen3.7-Max 53.5, DeepSeek-V4-Pro 48.2, Opus 4.8 57.9*, GPT-5.5 52.2*, Gemini 3.1 Pro 51.4*.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: CritPt
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: critpt not yet in data/benchmarks/. Competitor cells: GLM-5.1 4.6, Qwen3.7-Max 13.4, M3 3.7, V4-Pro 12.9, Opus 4.8 20.9, GPT-5.5 27.1, Gemini 3.1 Pro 17.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: AIME 2026
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "GPT-5.5 (medium) with mandated Explanation/Exact Answer/Confidence format"
}new-benchmark: aime-26 (also introduced in kimi-k2-6.json this batch). Competitor cells: GLM-5.1 95.3, Qwen3.7-Max 97.0, V4-Pro 94.6, Opus 4.8 95.7, GPT-5.5 98.3, Gemini 3.1 Pro 98.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HMMT Nov. 2025
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}HMMT Nov 2025 mapped to hmmt-25 with edition variant. Competitor cells: GLM-5.1 94.0, Qwen3.7-Max 95.0, M3 84.4, V4-Pro 94.4, Opus 4.8 96.5, GPT-5.5 96.5, Gemini 3.1 Pro 94.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: HMMT Feb. 2026
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hmmt-26 not yet in data/benchmarks/. Competitor cells: GLM-5.1 82.6, Qwen3.7-Max 97.1, M3 84.4, V4-Pro 95.2, Opus 4.8 96.7, GPT-5.5 96.7, Gemini 3.1 Pro 87.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: IMOAnswerBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}imo-answerbench id already introduced by prior batches. Competitor cells: GLM-5.1 83.8, Qwen3.7-Max 90.0, V4-Pro 89.8, Opus 4.8 83.5, Gemini 3.1 Pro 81.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: GPQA-Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5.1 86.2, Qwen3.7-Max 90.0, M3 93.0, V4-Pro 90.1, Opus 4.8 93.6, GPT-5.5 93.6, Gemini 3.1 Pro 94.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: SWE-bench Pro
{
"harness": "OpenHands with tailored instruction prompt",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens 32k, 400K context window",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose cross-check: '62.1 vs. 58.4 on SWE-bench Pro' over GLM-5.1. Competitor cells: GLM-5.1 58.4, Qwen3.7-Max 60.6, M3 59.0, V4-Pro 55.4, Opus 4.8 69.2, GPT-5.5 58.6, Gemini 3.1 Pro 54.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: NL2Repo
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens 48k, 400k context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}nl2repo id already introduced by prior batches. Anti-hacking: rule-based + LLM-based interception. Competitor cells: GLM-5.1 42.7, Qwen3.7-Max 47.2, M3 42.1, V4-Pro 35.5, Opus 4.8 69.7, GPT-5.5 50.7, Gemini 3.1 Pro 33.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: DeepSWE
{
"harness": "official pier evaluation framework + mini-swe-agent",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "400K context",
"turn_limit": null,
"time_limit": "2h timeout",
"run_count": null,
"aggregation": null,
"judge": null
}deepswe id already introduced by prior batches. Competitor cells: GLM-5.1 18.0, Qwen3.7-Max 18.0, M3 20.0, V4-Pro 8.0, Opus 4.8 58.0, GPT-5.5 70.0, Gemini 3.1 Pro 10.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: ProgramBench
{
"harness": "Claude-Code 2.1.156",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 1,
"token_budget": "max_tokens 64,000, max_turns 2000, 6h sample timeout, 400K context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}program-bench id already introduced by prior batches (200 instances). Competitor cells: GLM-5.1 50.9, V4-Pro 47.8, Opus 4.8 71.9, GPT-5.5 70.8, Gemini 3.1 Pro 39.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: Terminal Bench 2.1 Terminus-2
{
"harness": "Terminus-2, parser=json",
"tools": "resource caps 4 CPU / 8GB RAM",
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens 48k, 256K context, 500 max episodes",
"turn_limit": null,
"time_limit": "4h timeout",
"run_count": null,
"aggregation": null,
"judge": null
}Prose cross-check: '81.0 vs. 63.5 on Terminal-Bench 2.1' over GLM-5.1; 'within a few points of Claude Opus 4.8 (85.0)'. Competitor cells: GLM-5.1 63.5, Qwen3.7-Max 75.0, M3 65.0, V4-Pro 64.0, Opus 4.8 85.0, GPT-5.5 84.0, Gemini 3.1 Pro 74.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: Terminal Bench 2.1 Best Reported Harness
{
"harness": "Claude Code 2.1.167, max_new_tokens overridden to 128k via transparent proxy",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": "wall-clock limits removed; per-task CPU/memory constraints preserved",
"run_count": 5,
"aggregation": "mean over 5 runs",
"judge": null
}Best-across-harness row; per-cell harness in parentheses: GLM-5.1 69.0 (Claude Code), Opus 4.8 78.9 (Claude Code), GPT-5.5 83.4 (Codex), Gemini 3.1 Pro 70.7 (Gemini CLI). Different harness per cell - not same-protocol.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: FrontierSWE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "1M context, 128K max output",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}frontierswe id already introduced by prior batches. Footnote: evaluated by Proximal (third party). Prose: 'trails Opus 4.8 by only 1%, edging out GPT-5.5 by 1% and Opus 4.7 by 11%'. Competitor cells: GLM-5.1 30.5, V4-Pro 29.0, Opus 4.8 75.1, GPT-5.5 72.6, Gemini 3.1 Pro 39.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: PostTrainBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "1M context, 128K max output",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}posttrain-bench id already introduced by prior batches. Footnote: evaluated by PostTrainBench (each agent given an H100 GPU). Prose: 'outperforms both Opus 4.7 and GPT-5.5, ranking second only to Opus 4.8'. Competitor cells: GLM-5.1 20.1, Opus 4.8 37.2, GPT-5.5 28.4, Gemini 3.1 Pro 21.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: SWE-Marathon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "1M context, 128K max output",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}swe-marathon id already introduced by prior batches. Footnote: evaluated by Abundant AI (third party). Prose: 'trailing Opus 4.8 by 13% while remaining second only to the Opus series'. Competitor cells: GLM-5.1 1.0, Opus 4.8 26.0, GPT-5.5 12.0, Gemini 3.1 Pro 4.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: MCP-Atlas Public Set
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "10-minute timeout per task",
"run_count": null,
"aggregation": null,
"judge": "Gemini-3.0-Pro"
}mcp-atlas id already introduced by prior batches. Competitor cells: GLM-5.1 71.8, Qwen3.7-Max 76.4, M3 74.2, V4-Pro 73.6, Opus 4.8 77.8, GPT-5.5 75.3, Gemini 3.1 Pro 69.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full Benchmark Table · table: DOM table under 'Full Benchmark Table' (19 rows x 8 models, machine-readable) · row: Tool-Decathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max_token 128K via official evaluation service",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tool-decathlon not yet in data/benchmarks/ (distinct from toolathlon). Competitor cells: GLM-5.1 40.7, Qwen3.7-Max -, V4-Pro 52.8, Opus 4.8 59.9, GPT-5.5 55.6, Gemini 3.1 Pro 48.8.