← 模型目录

GLM-5.1

Z.ai / 智谱 GLM · 2026-04-07 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GLM-5.1

GLM-5.1 被 z.ai 定位为面向长程任务(Towards Long-Horizon Tasks)的代理工程旗舰,强调数百轮持续优化与长时程任务能力。评测集中在推理、编码与代理三类:SWE-Bench Pro 58.4、Terminal-Bench 2.0(Terminus-2)63.5、BrowseComp(w/ 上下文管理)79.3、HLE w/ Tools 52.3。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 31 模型 glm-5.1 · 版本 text-only subset (default) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HLE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max_new_tokens 163,840",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5 30.5, Qwen3.6-Plus 28.8, M2.7 28.0, DS-V3.2 25.1, K2.5 31.5, Opus 4.6 36.7, Gemini 3.1 Pro 45.0, GPT-5.4 39.8.

打开官方来源

hlehle 52.3 模型 glm-5.1 · 版本 w/ tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HLE w/ Tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "202,752 max context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5 50.4, Qwen3.6-Plus 50.6, DS-V3.2 40.8, K2.5 51.8, Opus 4.6 53.1*, Gemini 3.1 Pro 51.4*, GPT-5.4 52.1* (* = full set).

打开官方来源

aime-26 95.3 模型 glm-5.1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: AIME 2026

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aime-26. Competitor cells: GLM-5 95.4, Qwen3.6-Plus 95.1, M2.7 89.8, DS-V3.2 95.1, K2.5 94.5, Opus 4.6 95.6, Gemini 3.1 Pro 98.2, GPT-5.4 98.7.

打开官方来源

hmmt25 94 模型 glm-5.1 · 版本 November 2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HMMT Nov. 2025

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5 96.9, Qwen3.6-Plus 94.6, M2.7 81.0, DS-V3.2 90.2, K2.5 91.1, Opus 4.6 96.3, Gemini 3.1 Pro 94.8, GPT-5.4 95.8.

打开官方来源

hmmt-26 82.6 模型 glm-5.1 · 版本 February 2026 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: HMMT Feb. 2026

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hmmt-26. Competitor cells: GLM-5 82.8, Qwen3.6-Plus 87.8, M2.7 72.7, DS-V3.2 79.9, K2.5 81.3, Opus 4.6 84.3, Gemini 3.1 Pro 87.3, GPT-5.4 91.8.

打开官方来源

imo-answerbench 83.8 模型 glm-5.1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: IMOAnswerBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Kimi K2.6's page cites this GLM-5.1 value as the source for its GPT-5.4/Opus 4.6 IMO-AnswerBench cells (cross-vendor citation chain). Competitor cells: GLM-5 82.5, Qwen3.6-Plus 83.8, M2.7 66.3, DS-V3.2 78.3, K2.5 81.8, Opus 4.6 75.3, Gemini 3.1 Pro 81.0, GPT-5.4 91.4.

打开官方来源

gpqa 86.2 模型 glm-5.1 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: GPQA-Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5 86.0, Qwen3.6-Plus 90.4, M2.7 87.0, DS-V3.2 82.4, K2.5 87.6, Opus 4.6 91.3, Gemini 3.1 Pro 94.3, GPT-5.4 92.0.

打开官方来源

swebench-pro 58.4 模型 glm-5.1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: SWE-Bench Pro

{
  "harness": "OpenHands with tailored instruction prompt",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max_new_tokens 32768, 200K context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Headline claim: state-of-the-art at release. Competitor cells: GLM-5 55.1, Qwen3.6-Plus 56.6, M2.7 56.2, K2.5 53.8, Opus 4.6 57.3, Gemini 3.1 Pro 54.2, GPT-5.4 57.7.

打开官方来源

nl2repo 42.7 模型 glm-5.1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: NL2Repo

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens 32768, 200k context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5 35.9, Qwen3.6-Plus 37.9, M2.7 39.8, K2.5 32.0, Opus 4.6 49.8, Gemini 3.1 Pro 33.4, GPT-5.4 41.3.

打开官方来源

terminalbench 63.5 模型 glm-5.1 · 版本 2.0, Terminus-2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Terminal-Bench 2.0 Terminus-2

{
  "harness": "Terminus-2",
  "tools": "16 CPU / 32GB RAM caps",
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens 8192, 200K context",
  "turn_limit": null,
  "time_limit": "3h timeout",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5 56.2, Qwen3.6-Plus 61.6, DS-V3.2 39.3, K2.5 50.8, Opus 4.6 65.4, Gemini 3.1 Pro 68.5.

打开官方来源

terminalbench 69.0 (Claude Code) 模型 glm-5.1 · 版本 2.0, Claude Code (best self-reported harness) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Terminal-Bench 2.0 Best self-reported harness

{
  "harness": "Claude Code 2.1.69 (think mode), max_new_tokens 128k via transparent proxy",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 runs",
  "judge": null
}

Per-cell harnesses differ: GLM-5 56.2 (CC), Qwen3.6-Plus -, M2.7 57.0 (CC), DS-V3.2 46.4 (CC), GPT-5.4 75.1 (Codex). Best-across-harness convention.

打开官方来源

cybergym 68.7 模型 glm-5.1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: CyberGym

{
  "harness": "Claude Code 2.1.56 (think mode, no web tools)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens 32000",
  "turn_limit": null,
  "time_limit": "250-minute timeout per task",
  "run_count": 1,
  "aggregation": "single-run Pass@1 over 1,507 tasks",
  "judge": null
}

cybergym id already introduced by prior batches. Competitor harnesses differ (Gemini CLI / Codex CLI) and both sometimes refused on security grounds, lowering their scores - vendor-disclosed caveat. Cells: GLM-5 48.3, DS-V3.2 17.3, K2.5 41.3, Opus 4.6 66.6, Gemini 3.1 Pro 38.8, GPT-5.4 66.3.

打开官方来源

browsecomp 68 模型 glm-5.1 · 版本 no context management · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "without CM: retain details from most recent 5 turns",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-5 62.0, DS-V3.2 51.4, K2.5 60.6.

打开官方来源

browsecomp 79.3 模型 glm-5.1 · 版本 w/ context management (discard-all) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: BrowseComp w/ Context Manage

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "discard-all strategy, same as GLM-5 and DeepSeek-V3.2",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Two BrowseComp protocols printed as separate rows - do not merge. Competitor cells: GLM-5 75.9, DS-V3.2 67.6, K2.5 74.9, Opus 4.6 84.0, Gemini 3.1 Pro 85.9, GPT-5.4 82.7.

打开官方来源

tau3-bench 70.6 模型 glm-5.1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: tau3-Bench

{
  "harness": null,
  "tools": "banking domain uses terminal-based agentic search retrieval",
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "user simulator GPT-5.2 (reasoning_effort low), 4 trials"
}

new-benchmark: tau3-bench (tau-cubed, successor of tau2) not yet in data/benchmarks/. Extra user-simulator prompt added across domains. Competitor cells: GLM-5 69.2, Qwen3.6-Plus 70.7, M2.7 67.6, DS-V3.2 69.2, K2.5 66.0, Opus 4.6 72.4, Gemini 3.1 Pro 67.1, GPT-5.4 72.9.

打开官方来源

mcp-atlas 71.8 模型 glm-5.1 · 版本 500-task public set · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: MCP-Atlas Public Set

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "10-minute timeout per task",
  "run_count": null,
  "aggregation": null,
  "judge": "Gemini-3.0-Pro"
}

Competitor cells: GLM-5 69.2, Qwen3.6-Plus 74.1, M2.7 48.8, DS-V3.2 62.2, K2.5 63.8, Opus 4.6 73.8, Gemini 3.1 Pro 69.2, GPT-5.4 67.2.

打开官方来源

toolathlon 40.7 模型 glm-5.1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Tool-Decathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "max_token 128K, official evaluation service",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tool-decathlon. Competitor cells: GLM-5 38.0, Qwen3.6-Plus 39.8, M2.7 46.3, DS-V3.2 35.2, K2.5 27.8, Opus 4.6 47.2, Gemini 3.1 Pro 48.8, GPT-5.4 54.6.

打开官方来源

vending-bench-2 $5,634.41 模型 glm-5.1 · 版本 未说明 · 指标 final_account_balance · 单位 usd 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 9 models, machine-readable) · row: Vending Bench 2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

vending-bench-2 id already introduced by prior batches. Unit is US dollars (final account balance over a simulated one-year vending business) - never aggregate with percent rows. Footnote: runs conducted independently by Andon Labs. Competitor cells: GLM-5 $4,432.12, Qwen3.6-Plus $5,114.87, DS-V3.2 $1,034.00, K2.5 $1,198.46, Opus 4.6 $8,017.59, Gemini 3.1 Pro $911.21, GPT-5.4 $6,144.18.

打开官方来源