GLM-5
Z.ai / 智谱 GLM · 2026-02-11 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GLM-5
智谱以「From Vibe Coding to Agentic Engineering」为主题发布 GLM-5:744B 总参/40B 激活的 MoE(DSA 架构,28.5T 预训练 tokens),MIT 许可开源。评测覆盖推理、编码与通用智能体,亮点如 HMMT 2025.11 96.9%、τ²-Bench 89.7%。
- 输入模态
- 文本
- 上下文
- 200K
- 参数
- 744B-A40B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Humanity's Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "max_new_tokens 131,072",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 24.8, DS-V3.2 25.1, K2.5 31.5, Opus 4.5 28.4, Gemini 3.0 Pro 37.2, GPT-5.2 35.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Humanity's Last Exam w/ Tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "202,752 max context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 42.8, DS-V3.2 40.8, K2.5 51.8, Opus 4.5 43.4*, Gemini 3.0 Pro 45.8*, GPT-5.2 45.5* (* = full set).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: AIME 2026 I
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-26. AIME 2026 split into I (this row) - distinct from the combined 'AIME 2026' rows in GLM-5.2/K2.6 pages. Competitor cells: GLM-4.7 92.9, DS-V3.2 92.7, K2.5 92.5, Opus 4.5 93.3, Gemini 3.0 Pro 90.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: HMMT Nov. 2025
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 93.5, DS-V3.2 90.2, K2.5 91.1, Opus 4.5 91.7, Gemini 3.0 Pro 93.0, GPT-5.2 97.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: IMOAnswerBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 82.0, DS-V3.2 78.3, K2.5 81.8, Opus 4.5 78.5, Gemini 3.0 Pro 83.3, GPT-5.2 86.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: GPQA-Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 85.7, DS-V3.2 82.4, K2.5 87.6, Opus 4.5 87.0, Gemini 3.0 Pro 91.9, GPT-5.2 92.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: SWE-bench Verified
{
"harness": "OpenHands with tailored instruction prompt",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 0.7,
"top_p": 0.95,
"token_budget": "max_new_tokens 16384, 200K context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 73.8, DS-V3.2 73.1, K2.5 76.8, Opus 4.5 80.9, Gemini 3.0 Pro 76.2, GPT-5.2 80.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: SWE-bench Multilingual
{
"harness": "OpenHands",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 0.7,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual already introduced by prior batches. Competitor cells: GLM-4.7 66.7, DS-V3.2 70.2, K2.5 73.0, Opus 4.5 77.5, Gemini 3.0 Pro 65.0, GPT-5.2 72.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Terminal-Bench 2.0 Terminus-2
{
"harness": "Terminus-2",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 0.7,
"top_p": 1,
"token_budget": "max_new_tokens 8192, 128K context",
"turn_limit": null,
"time_limit": "2h timeout",
"run_count": null,
"aggregation": null,
"judge": null
}Dagger = score on the vendor's 'verified' Terminal-Bench 2.0 dataset that fixes ambiguous instructions (HF: zai-org/terminal-bench-2-verified) - a dataset variant change, disclosed. Competitor cells: GLM-4.7 41.0, DS-V3.2 39.3, K2.5 50.8, Opus 4.5 59.3, Gemini 3.0 Pro 54.2, GPT-5.2 54.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Terminal-Bench 2.0 Claude Code
{
"harness": "Claude Code 2.1.14 (think mode)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "max_new_tokens 65536",
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "mean over 5 runs",
"judge": null
}Dagger = verified-dataset variant (61.1). Competitor cells: GLM-4.7 32.8, DS-V3.2 46.4, Opus 4.5 57.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: CyberGym
{
"harness": "Claude Code 2.1.18 (think mode, no web tools)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens 32000",
"turn_limit": null,
"time_limit": "250-minute timeout",
"run_count": 1,
"aggregation": "single-run Pass@1 over 1,507 tasks",
"judge": null
}Competitor cells: GLM-4.7 23.5, DS-V3.2 17.3, K2.5 41.3, Opus 4.5 50.6, Gemini 3.0 Pro 39.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: BrowseComp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "retain most recent 5 turns",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 52.0, DS-V3.2 51.4, K2.5 60.6, Opus 4.5 37.0, Gemini 3.0 Pro 37.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: BrowseComp w/ Context Manage
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "discard-all strategy, same as DeepSeek-V3.2 and Kimi K2.5",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: GLM-4.7 67.5, DS-V3.2 67.6, K2.5 74.9, Opus 4.5 67.8, Gemini 3.0 Pro 59.2, GPT-5.2 65.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: BrowseComp-Zh
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp-zh. Competitor cells: GLM-4.7 66.6, DS-V3.2 65.0, K2.5 62.3, Opus 4.5 62.4, Gemini 3.0 Pro 66.8, GPT-5.2 76.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: tau2-Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prompt adjustments in Retail/Telecom; Airline fixes from Claude Opus 4.5 system card - protocol deviations disclosed. Competitor cells: GLM-4.7 87.4, DS-V3.2 85.3, K2.5 80.2, Opus 4.5 91.6, Gemini 3.0 Pro 90.7, GPT-5.2 85.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: MCP-Atlas Public Set
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "10-minute timeout per task",
"run_count": null,
"aggregation": null,
"judge": "Gemini 3 Pro"
}Competitor cells: GLM-4.7 52.0, DS-V3.2 62.2, K2.5 63.8, Opus 4.5 65.2, Gemini 3.0 Pro 66.6, GPT-5.2 68.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Tool-Decathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tool-decathlon. Competitor cells: GLM-4.7 23.8, DS-V3.2 35.2, K2.5 27.8, Opus 4.5 43.5, Gemini 3.0 Pro 36.4, GPT-5.2 46.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Full benchmark table (DOM) · table: DOM table (18 benchmark rows x 7 models, machine-readable); all columns Thinking mode · row: Vending Bench 2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose: 'GLM-5 finishes with a final account balance of $4,432, approaching Claude Opus 4.5' and 'ranks #1 among open-source models'. Runs by Andon Labs. Competitor cells: GLM-4.7 $2,376.82, DS-V3.2 $1,034.00, K2.5 $1,198.46, Opus 4.5 $4,967.06, Gemini 3.0 Pro $5,478.16, GPT-5.2 $3,591.33.