← 模型目录

GLM-4.6

Z.ai / 智谱 GLM · 2025-09-30 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GLM-4.6

智谱将 GLM-4.6 定位为具备高级 Agentic、推理与编程能力的模型,上下文由 128K 扩展至 200K,推理中可交织工具调用。评测以编码与智能体为主,亮点如 CC-Bench 对 Claude Sonnet 4 胜率 48.6%、AIME 25 93.9%。

输入模态
文本
上下文
200K
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

cc-bench 48.6% 模型 glm-4.6 · 版本 extended from GLM-4.5 with more challenging tasks · 指标 win_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: CC-Bench paragraph after the coding benchmark image · row: CC-Bench · quote_snippet: GLM-4.6 improves over GLM-4.5 and reaches near parity with Claude Sonnet 4 (48.6% win rate)

{
  "harness": "human evaluators working with models inside isolated Docker containers",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "win rate vs Claude Sonnet 4 in 1-on-1 human evaluation",
  "judge": "human evaluators"
}

new-benchmark: cc-bench (GLM real-world multi-turn coding bench: front-end development, tool building, data analysis, testing, algorithm) not yet in data/benchmarks/. Win rate is against Claude Sonnet 4 as the reference, not an absolute accuracy. Token efficiency disclosed in same paragraph: GLM-4.6 finishes tasks with about 15% fewer tokens than GLM-4.5. Evaluation details and trajectories published at https://huggingface.co/datasets/zai-org/CC-Bench-trajectories.

打开官方来源

aime-25 93.9 模型 glm-4.6 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png ('We evaluated GLM-4.6 across eight public benchmarks') · table: benchmark comparison chart image coding_benchmark.png · row: AIME 25 · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context) · quote_snippet: We evaluated GLM-4.6 across eight public benchmarks covering agents, reasoning, and coding

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length (chart caption)",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed, per goal.md 12.5): AIME 25 GLM-4.6 93.9 base and 98.6 w/ Tools (two conditions printed in one cell); GLM-4.5 85.4, DeepSeek-V3.2-Exp 89.3, Claude Sonnet 4 74.3, Claude Sonnet 4.5 87.0. The base and w/-tools values are different protocols; split into variant rows when promoting to verified. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 93.9,同柱上方 w/ Tools 叠加值 98.6(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 85.4, DeepSeek-V3.2-Exp 89.3, Claude Sonnet 4 74.3, Claude Sonnet 4.5 87.0。

打开官方来源

gpqa 81 模型 glm-4.6 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: GPQA · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): GPQA GLM-4.6 81.0 base / 82.9 w/ Tools; GLM-4.5 79.9, DeepSeek-V3.2-Exp 79.9, Sonnet 4 77.7, Sonnet 4.5 83.4. Variant (Diamond vs main) not resolvable from chart. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 81,同柱上方 w/ Tools 叠加值 82.9(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 79.9, DeepSeek-V3.2-Exp 79.9, Claude Sonnet 4 77.7, Claude Sonnet 4.5 83.4。

打开官方来源

lcb 82.8 模型 glm-4.6 · 版本 v6 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: LiveCodeBench v6 · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): LiveCodeBench v6 GLM-4.6 82.8 base / 84.5 w/ Tools; GLM-4.5 63.3, DeepSeek-V3.2-Exp 70.1, Sonnet 4 48.9, Sonnet 4.5 57.7. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 82.8,同柱上方 w/ Tools 叠加值 84.5(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 63.3, DeepSeek-V3.2-Exp 70.1, Claude Sonnet 4 48.9, Claude Sonnet 4.5 57.7。

打开官方来源

hlehle 17.2 模型 glm-4.6 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: HLE · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): HLE GLM-4.6 17.2 base / 30.4 w/ Tools; GLM-4.5 14.4, DeepSeek-V3.2-Exp 19.8, Sonnet 4 9.6, Sonnet 4.5 17.3. Tool condition is a protocol split - do not merge base and w/-tools when promoting. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 17.2,同柱上方 w/ Tools 叠加值 30.4(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 14.4, DeepSeek-V3.2-Exp 19.8, Claude Sonnet 4 9.6, Claude Sonnet 4.5 17.3。

打开官方来源

browsecomp 45.1 模型 glm-4.6 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: BrowseComp · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): BrowseComp GLM-4.6 45.1; GLM-4.5 26.4, DeepSeek-V3.2-Exp 40.1, Sonnet 4 14.7, Sonnet 4.5 19.6. Single value (no w/-tools split on this row). 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 45.1。竞品:GLM-4.5 26.4, DeepSeek-V3.2-Exp 40.1, Claude Sonnet 4 14.7, Claude Sonnet 4.5 19.6。

打开官方来源

swebench 68 模型 glm-4.6 · 版本 Verified · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: SWE-bench Verified · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context) · quote_snippet: but still lags behind Claude Sonnet 4.5 in coding ability

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): SWE-bench Verified GLM-4.6 68.0; GLM-4.5 64.2, DeepSeek-V3.2-Exp 67.8, Sonnet 4 72.5, Sonnet 4.5 77.2. Prose cross-check: page explicitly concedes 'still lags behind Claude Sonnet 4.5 in coding ability'. Harness not printed on this page (GLM-5.3 pages disclose Claude Code harness; do not assume the same here). 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 68。竞品:GLM-4.5 64.2, DeepSeek-V3.2-Exp 67.8, Claude Sonnet 4 72.5, Claude Sonnet 4.5 77.2。

打开官方来源

terminalbench 40.5 模型 glm-4.6 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: Terminal-Bench · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): Terminal-Bench GLM-4.6 40.5; GLM-4.5 37.5, DeepSeek-V3.2-Exp 37.7, Sonnet 4 35.5, Sonnet 4.5 50.0. Terminal-Bench version not printed. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 40.5。竞品:GLM-4.5 37.5, DeepSeek-V3.2-Exp 37.7, Claude Sonnet 4 35.5, Claude Sonnet 4.5 50.0。

打开官方来源

tau-bench 75.9 模型 glm-4.6 · 版本 2, weighted · 指标 weighted_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: tau^2-Bench (Weighted) · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "128K context length",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "weighted",
  "judge": null
}

Vision-assisted read (unconfirmed): tau^2-Bench (Weighted) GLM-4.6 75.9; GLM-4.5 67.5, DeepSeek-V3.2-Exp 53.4, Sonnet 4 66.0, Sonnet 4.5 88.1. Weighted aggregation over domains differs from per-domain Tau2 rows (retail/airline/telecom) in kimi-k2.json. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 75.9。竞品:GLM-4.5 67.5, DeepSeek-V3.2-Exp 53.4, Claude Sonnet 4 66.0, Claude Sonnet 4.5 88.1。

打开官方来源