← 模型目录

GLM-4.7

Z.ai / 智谱 GLM · 2025-12-22 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GLM-4.7

发布文以“推进编码能力”为题,将 GLM-4.7 定位为编码与智能体任务导向的模型,支持交错/保留/回合级思考模式。17 项评测覆盖数理推理、代码智能体与通用智能体,亮点为 HMMT 2025 二月场 97.1 与 τ²-Bench 87.4(SWE-bench Verified 73.8)。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu-pro 84.3 模型 glm-4.7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: MMLU-Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max new tokens 131,072",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 83.2, K2 Thinking 84.6, DS-V3.2 85.0, Gemini 3.0 Pro 90.1, Sonnet 4.5 88.2, GPT-5 87.5, GPT-5.1 87.0.

打开官方来源

gpqa 85.7 模型 glm-4.7 · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: GPQA-Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 81.0, K2 Thinking 84.5, DS-V3.2 82.4, Gemini 3.0 Pro 91.9, Sonnet 4.5 83.4, GPT-5 85.7, GPT-5.1 88.1.

打开官方来源

hlehle 24.8 模型 glm-4.7 · 版本 text-only (default) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: HLE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max new tokens 131,072",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 17.2, K2 Thinking 23.9, DS-V3.2 25.1, Gemini 3.0 Pro 37.5, Sonnet 4.5 13.7, GPT-5 26.3, GPT-5.1 25.7. GLM-4.6's own page printed its HLE base at 17.2 - consistent.

打开官方来源

hlehle 42.8 模型 glm-4.7 · 版本 w/ tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: HLE w/ Tools · quote_snippet: achieving (42.8%, +12.4%) on the HLE benchmark compared to GLM-4.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose cross-check printed. Competitor cells: GLM-4.6 30.4, K2 Thinking 44.9, DS-V3.2 40.8, Gemini 3.0 Pro 45.8, Sonnet 4.5 32.0, GPT-5 35.2, GPT-5.1 42.7.

打开官方来源

aime-25 95.7 模型 glm-4.7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: AIME 2025

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 93.9, K2 Thinking 94.5, DS-V3.2 93.1, Gemini 3.0 Pro 95.0, Sonnet 4.5 87.0, GPT-5 94.6, GPT-5.1 94.0. GLM-4.6's page printed AIME 25 93.9 base / 98.6 w-tools - this row is presumably the no-tools condition.

打开官方来源

hmmt25 97.1 模型 glm-4.7 · 版本 February 2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: HMMT Feb. 2025

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Same contest as prior batches' hmmt-25 rows. Competitor cells: GLM-4.6 89.2, K2 Thinking 89.4, DS-V3.2 92.5, Gemini 3.0 Pro 97.5, Sonnet 4.5 79.2, GPT-5 88.3, GPT-5.1 96.3.

打开官方来源

hmmt25 93.5 模型 glm-4.7 · 版本 November 2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: HMMT Nov. 2025

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 87.7, K2 Thinking 89.2, DS-V3.2 90.2, Gemini 3.0 Pro 93.3, Sonnet 4.5 81.7, GPT-5 89.2.

打开官方来源

imo-answerbench 82 模型 glm-4.7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: IMOAnswerBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 73.5, K2 Thinking 78.6, DS-V3.2 78.3, Gemini 3.0 Pro 83.3, Sonnet 4.5 65.8, GPT-5 76.0.

打开官方来源

lcb 84.9 模型 glm-4.7 · 版本 v6 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: LiveCodeBench-v6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 82.8, K2 Thinking 83.1, DS-V3.2 83.3, Gemini 3.0 Pro 90.7, Sonnet 4.5 64.0, GPT-5 87.0, GPT-5.1 87.0.

打开官方来源

swebench 73.8 模型 glm-4.7 · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: SWE-bench Verified · quote_snippet: (73.8%, +5.8%) on SWE-bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0.7,
  "top_p": 1,
  "token_budget": "max new tokens 16,384",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose delta cross-check. Competitor cells: GLM-4.6 68.0, K2 Thinking 71.3, DS-V3.2 73.1, Gemini 3.0 Pro 76.2, Sonnet 4.5 77.2, GPT-5 74.9, GPT-5.1 76.3.

打开官方来源

swebench-multilingual 66.7 模型 glm-4.7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: SWE-bench Multilingual · quote_snippet: (66.7%, +12.9%) on SWE-bench Multilingual

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-multilingual. Competitor cells: GLM-4.6 53.8, K2 Thinking 61.1, DS-V3.2 70.2, Sonnet 4.5 68.0, GPT-5 55.3.

打开官方来源

terminalbench 33.3 模型 glm-4.7 · 版本 Hard · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: Terminal Bench Hard

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark variant: Terminal Bench Hard subset (id terminalbench exists; Hard recorded as variant). Competitor cells: GLM-4.6 23.6, K2 Thinking 30.6, DS-V3.2 35.4, Gemini 3.0 Pro 39.0, Sonnet 4.5 33.3, GPT-5 30.5, GPT-5.1 43.0.

打开官方来源

terminalbench 41 模型 glm-4.7 · 版本 2.0, Preserved Thinking mode · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: Terminal Bench 2.0 · quote_snippet: (41%, +16.5%) on Terminal Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Preserved Thinking mode enabled for multi-turn agentic tasks",
  "temperature": 0.7,
  "top_p": 1,
  "token_budget": "max new tokens 16,384",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose delta cross-check. Competitor cells: GLM-4.6 24.5, K2 Thinking 35.7, DS-V3.2 46.4, Gemini 3.0 Pro 54.2, Sonnet 4.5 42.8, GPT-5 35.2, GPT-5.1 47.6.

打开官方来源

browsecomp 52 模型 glm-4.7 · 版本 no context management · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: BrowseComp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 45.1, DS-V3.2 51.4, Sonnet 4.5 24.1, GPT-5 54.9, GPT-5.1 50.8. GLM-4.6's own chart printed 45.1 - consistent.

打开官方来源

browsecomp 67.5 模型 glm-4.7 · 版本 w/ context management · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: BrowseComp w/ Context Manage

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: GLM-4.6 57.5, K2 Thinking 60.2, DS-V3.2 67.6, Gemini 3.0 Pro 59.2.

打开官方来源

browsecomp-zh 66.6 模型 glm-4.7 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: BrowseComp-ZH

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp-zh. Competitor cells: GLM-4.6 49.5, K2 Thinking 62.3, DS-V3.2 65.0, Sonnet 4.5 42.4, GPT-5 63.0.

打开官方来源

tau-bench 87.4 模型 glm-4.7 · 版本 2, temp 0 + domain prompt fixes · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance table (DOM) · table: DOM table (17 benchmarks x 8 models, machine-readable) · row: tau2-Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Preserved Thinking mode",
  "temperature": 0,
  "top_p": null,
  "token_budget": "max new tokens 16,384",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Distinct sampling from the rest of the table (temp 0); Retail/Telecom prompt adjustments + Airline fixes from Claude Opus 4.5 report. Competitor cells: GLM-4.6 75.2, K2 Thinking 74.3, DS-V3.2 85.3, Gemini 3.0 Pro 90.7, Sonnet 4.5 87.2, GPT-5 82.4, GPT-5.1 82.7.

打开官方来源