← 模型目录

GLM-5

Z.ai / 智谱 GLM · 2026-02-11 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GLM-5

智谱以「From Vibe Coding to Agentic Engineering」为主题发布 GLM-5:744B 总参/40B 激活的 MoE(DSA 架构,28.5T 预训练 tokens),MIT 许可开源。评测覆盖推理、编码与通用智能体,亮点如 HMMT 2025.11 96.9%、τ²-Bench 89.7%。

输入模态
文本
上下文
200K
参数
744B-A40B MoE
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 30.5 模型 glm-5 · 版本 text-only subset (default), thinking · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Humanity's Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max_new_tokens 131,072",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 24.8, DS-V3.2 25.1, K2.5 31.5, Opus 4.5 28.4, Gemini 3.0 Pro 37.2, GPT-5.2 35.4.

打开官方来源

hlehle 50.4 模型 glm-5 · 版本 w/ tools, thinking · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Humanity's Last Exam w/ Tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "202,752 max context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 42.8, DS-V3.2 40.8, K2.5 51.8, Opus 4.5 43.4*, Gemini 3.0 Pro 45.8*, GPT-5.2 45.5* (* = full set).

打开官方来源

aime-26 92.7 模型 glm-5 · 版本 I · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: AIME 2026 I

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aime-26. AIME 2026 split into I (this row) - distinct from the combined 'AIME 2026' rows in GLM-5.2/K2.6 pages. Competitor cells: GLM-4.7 92.9, DS-V3.2 92.7, K2.5 92.5, Opus 4.5 93.3, Gemini 3.0 Pro 90.6.

打开官方来源

hmmt25 96.9 模型 glm-5 · 版本 November 2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: HMMT Nov. 2025

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 93.5, DS-V3.2 90.2, K2.5 91.1, Opus 4.5 91.7, Gemini 3.0 Pro 93.0, GPT-5.2 97.1.

打开官方来源

imo-answerbench 82.5 模型 glm-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: IMOAnswerBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 82.0, DS-V3.2 78.3, K2.5 81.8, Opus 4.5 78.5, Gemini 3.0 Pro 83.3, GPT-5.2 86.3.

打开官方来源

gpqa 86 模型 glm-5 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: GPQA-Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 85.7, DS-V3.2 82.4, K2.5 87.6, Opus 4.5 87.0, Gemini 3.0 Pro 91.9, GPT-5.2 92.4.

打开官方来源

swebench 77.8 模型 glm-5 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: SWE-bench Verified

{
  "harness": "OpenHands with tailored instruction prompt",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0.7,
  "top_p": 0.95,
  "token_budget": "max_new_tokens 16384, 200K context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 73.8, DS-V3.2 73.1, K2.5 76.8, Opus 4.5 80.9, Gemini 3.0 Pro 76.2, GPT-5.2 80.0.

打开官方来源

swebench-multilingual 73.3 模型 glm-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: SWE-bench Multilingual

{
  "harness": "OpenHands",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0.7,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-multilingual already introduced by prior batches. Competitor cells: GLM-4.7 66.7, DS-V3.2 70.2, K2.5 73.0, Opus 4.5 77.5, Gemini 3.0 Pro 65.0, GPT-5.2 72.0.

打开官方来源

terminalbench 56.2 / 60.7 dagger 模型 glm-5 · 版本 2.0, Terminus-2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Terminal-Bench 2.0 Terminus-2

{
  "harness": "Terminus-2",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0.7,
  "top_p": 1,
  "token_budget": "max_new_tokens 8192, 128K context",
  "turn_limit": null,
  "time_limit": "2h timeout",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Dagger = score on the vendor's 'verified' Terminal-Bench 2.0 dataset that fixes ambiguous instructions (HF: zai-org/terminal-bench-2-verified) - a dataset variant change, disclosed. Competitor cells: GLM-4.7 41.0, DS-V3.2 39.3, K2.5 50.8, Opus 4.5 59.3, Gemini 3.0 Pro 54.2, GPT-5.2 54.0.

打开官方来源

terminalbench 56.2 / 61.1 dagger 模型 glm-5 · 版本 2.0, Claude Code (incl. verified dataset variant) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Terminal-Bench 2.0 Claude Code

{
  "harness": "Claude Code 2.1.14 (think mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max_new_tokens 65536",
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "mean over 5 runs",
  "judge": null
}

Dagger = verified-dataset variant (61.1). Competitor cells: GLM-4.7 32.8, DS-V3.2 46.4, Opus 4.5 57.9.

打开官方来源

cybergym 43.2 模型 glm-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: CyberGym

{
  "harness": "Claude Code 2.1.18 (think mode, no web tools)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens 32000",
  "turn_limit": null,
  "time_limit": "250-minute timeout",
  "run_count": 1,
  "aggregation": "single-run Pass@1 over 1,507 tasks",
  "judge": null
}

Competitor cells: GLM-4.7 23.5, DS-V3.2 17.3, K2.5 41.3, Opus 4.5 50.6, Gemini 3.0 Pro 39.9.

打开官方来源

browsecomp 62 模型 glm-5 · 版本 no context management · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "retain most recent 5 turns",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 52.0, DS-V3.2 51.4, K2.5 60.6, Opus 4.5 37.0, Gemini 3.0 Pro 37.8.

打开官方来源

browsecomp 75.9 模型 glm-5 · 版本 w/ context management (discard-all) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: BrowseComp w/ Context Manage

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "discard-all strategy, same as DeepSeek-V3.2 and Kimi K2.5",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.7 67.5, DS-V3.2 67.6, K2.5 74.9, Opus 4.5 67.8, Gemini 3.0 Pro 59.2, GPT-5.2 65.8.

打开官方来源

browsecomp-zh 72.7 模型 glm-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: BrowseComp-Zh

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp-zh. Competitor cells: GLM-4.7 66.6, DS-V3.2 65.0, K2.5 62.3, Opus 4.5 62.4, Gemini 3.0 Pro 66.8, GPT-5.2 76.1.

打开官方来源

tau-bench 89.7 模型 glm-5 · 版本 2 (with domain prompt fixes) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: tau2-Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prompt adjustments in Retail/Telecom; Airline fixes from Claude Opus 4.5 system card - protocol deviations disclosed. Competitor cells: GLM-4.7 87.4, DS-V3.2 85.3, K2.5 80.2, Opus 4.5 91.6, Gemini 3.0 Pro 90.7, GPT-5.2 85.5.

打开官方来源

mcp-atlas 67.8 模型 glm-5 · 版本 500-task public set · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: MCP-Atlas Public Set

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "10-minute timeout per task",
  "run_count": null,
  "aggregation": null,
  "judge": "Gemini 3 Pro"
}

Competitor cells: GLM-4.7 52.0, DS-V3.2 62.2, K2.5 63.8, Opus 4.5 65.2, Gemini 3.0 Pro 66.6, GPT-5.2 68.0.

打开官方来源

toolathlon 39.2 模型 glm-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Tool-Decathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tool-decathlon. Competitor cells: GLM-4.7 23.8, DS-V3.2 35.2, K2.5 27.8, Opus 4.5 43.5, Gemini 3.0 Pro 36.4, GPT-5.2 46.3.

打开官方来源

vending-bench-2 $4,432.12 模型 glm-5 · 版本 未说明 · 指标 final_account_balance · 单位 usd 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Vending Bench 2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose: 'GLM-5 finishes with a final account balance of $4,432, approaching Claude Opus 4.5' and 'ranks #1 among open-source models'. Runs by Andon Labs. Competitor cells: GLM-4.7 $2,376.82, DS-V3.2 $1,034.00, K2.5 $1,198.46, Opus 4.5 $4,967.06, Gemini 3.0 Pro $5,478.16, GPT-5.2 $3,591.33.

打开官方来源