GLM-5.1
Z.ai / 智谱 GLM · 2026-04-07 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GLM-5.1
GLM-5.1 被 z.ai 定位为面向长程任务(Towards Long-Horizon Tasks)的代理工程旗舰,强调数百轮持续优化与长时程任务能力。评测集中在推理、编码与代理三类:SWE-Bench Pro 58.4、Terminal-Bench 2.0(Terminus-2)63.5、BrowseComp(w/ 上下文管理)79.3、HLE w/ Tools 52.3。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HLE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "max_new_tokens 163,840",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5 30.5, Qwen3.6-Plus 28.8, M2.7 28.0, DS-V3.2 25.1, K2.5 31.5, Opus 4.6 36.7, Gemini 3.1 Pro 45.0, GPT-5.4 39.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HLE w/ Tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "202,752 max context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5 50.4, Qwen3.6-Plus 50.6, DS-V3.2 40.8, K2.5 51.8, Opus 4.6 53.1*, Gemini 3.1 Pro 51.4*, GPT-5.4 52.1* (* = full set).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: AIME 2026
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-26. Competitor cells: GLM-5 95.4, Qwen3.6-Plus 95.1, M2.7 89.8, DS-V3.2 95.1, K2.5 94.5, Opus 4.6 95.6, Gemini 3.1 Pro 98.2, GPT-5.4 98.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HMMT Nov. 2025
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5 96.9, Qwen3.6-Plus 94.6, M2.7 81.0, DS-V3.2 90.2, K2.5 91.1, Opus 4.6 96.3, Gemini 3.1 Pro 94.8, GPT-5.4 95.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HMMT Feb. 2026
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hmmt-26. Competitor cells: GLM-5 82.8, Qwen3.6-Plus 87.8, M2.7 72.7, DS-V3.2 79.9, K2.5 81.3, Opus 4.6 84.3, Gemini 3.1 Pro 87.3, GPT-5.4 91.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: IMOAnswerBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Kimi K2.6's page cites this GLM-5.1 value as the source for its GPT-5.4/Opus 4.6 IMO-AnswerBench cells (cross-vendor citation chain). Competitor cells: GLM-5 82.5, Qwen3.6-Plus 83.8, M2.7 66.3, DS-V3.2 78.3, K2.5 81.8, Opus 4.6 75.3, Gemini 3.1 Pro 81.0, GPT-5.4 91.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: GPQA-Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5 86.0, Qwen3.6-Plus 90.4, M2.7 87.0, DS-V3.2 82.4, K2.5 87.6, Opus 4.6 91.3, Gemini 3.1 Pro 94.3, GPT-5.4 92.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: SWE-Bench Pro
{
"harness": "OpenHands with tailored instruction prompt",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "max_new_tokens 32768, 200K context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Headline claim: state-of-the-art at release. Competitor cells: GLM-5 55.1, Qwen3.6-Plus 56.6, M2.7 56.2, K2.5 53.8, Opus 4.6 57.3, Gemini 3.1 Pro 54.2, GPT-5.4 57.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: NL2Repo
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens 32768, 200k context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5 35.9, Qwen3.6-Plus 37.9, M2.7 39.8, K2.5 32.0, Opus 4.6 49.8, Gemini 3.1 Pro 33.4, GPT-5.4 41.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Terminal-Bench 2.0 Terminus-2
{
"harness": "Terminus-2",
"tools": "16 CPU / 32GB RAM caps",
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens 8192, 200K context",
"turn_limit": null,
"time_limit": "3h timeout",
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5 56.2, Qwen3.6-Plus 61.6, DS-V3.2 39.3, K2.5 50.8, Opus 4.6 65.4, Gemini 3.1 Pro 68.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Terminal-Bench 2.0 Best self-reported harness
{
"harness": "Claude Code 2.1.69 (think mode), max_new_tokens 128k via transparent proxy",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 runs",
"judge": null
}Per-cell harnesses differ: GLM-5 56.2 (CC), Qwen3.6-Plus -, M2.7 57.0 (CC), DS-V3.2 46.4 (CC), GPT-5.4 75.1 (Codex). Best-across-harness convention.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: CyberGym
{
"harness": "Claude Code 2.1.56 (think mode, no web tools)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens 32000",
"turn_limit": null,
"time_limit": "250-minute timeout per task",
"run_count": 1,
"aggregation": "single-run Pass@1 over 1,507 tasks",
"judge": null
}cybergym id already introduced by prior batches. Competitor harnesses differ (Gemini CLI / Codex CLI) and both sometimes refused on security grounds, lowering their scores - vendor-disclosed caveat. Cells: GLM-5 48.3, DS-V3.2 17.3, K2.5 41.3, Opus 4.6 66.6, Gemini 3.1 Pro 38.8, GPT-5.4 66.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "without CM: retain details from most recent 5 turns",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-5 62.0, DS-V3.2 51.4, K2.5 60.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: BrowseComp w/ Context Manage
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "discard-all strategy, same as GLM-5 and DeepSeek-V3.2",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Two BrowseComp protocols printed as separate rows - do not merge. Competitor cells: GLM-5 75.9, DS-V3.2 67.6, K2.5 74.9, Opus 4.6 84.0, Gemini 3.1 Pro 85.9, GPT-5.4 82.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: tau3-Bench
{
"harness": null,
"tools": "banking domain uses terminal-based agentic search retrieval",
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "user simulator GPT-5.2 (reasoning_effort low), 4 trials"
}new-benchmark: tau3-bench (tau-cubed, successor of tau2) not yet in data/benchmarks/. Extra user-simulator prompt added across domains. Competitor cells: GLM-5 69.2, Qwen3.6-Plus 70.7, M2.7 67.6, DS-V3.2 69.2, K2.5 66.0, Opus 4.6 72.4, Gemini 3.1 Pro 67.1, GPT-5.4 72.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: MCP-Atlas Public Set
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "10-minute timeout per task",
"run_count": null,
"aggregation": null,
"judge": "Gemini-3.0-Pro"
}Competitor cells: GLM-5 69.2, Qwen3.6-Plus 74.1, M2.7 48.8, DS-V3.2 62.2, K2.5 63.8, Opus 4.6 73.8, Gemini 3.1 Pro 69.2, GPT-5.4 67.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Tool-Decathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "max_token 128K, official evaluation service",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tool-decathlon. Competitor cells: GLM-5 38.0, Qwen3.6-Plus 39.8, M2.7 46.3, DS-V3.2 35.2, K2.5 27.8, Opus 4.6 47.2, Gemini 3.1 Pro 48.8, GPT-5.4 54.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Vending Bench 2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}vending-bench-2 id already introduced by prior batches. Unit is US dollars (final account balance over a simulated one-year vending business) - never aggregate with percent rows. Footnote: runs conducted independently by Andon Labs. Competitor cells: GLM-5 $4,432.12, Qwen3.6-Plus $5,114.87, DS-V3.2 $1,034.00, K2.5 $1,198.46, Opus 4.6 $8,017.59, Gemini 3.1 Pro $911.21, GPT-5.4 $6,144.18.