GLM-4.6
Z.ai / 智谱 GLM · 2025-09-30 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GLM-4.6
智谱将 GLM-4.6 定位为具备高级 Agentic、推理与编程能力的模型,上下文由 128K 扩展至 200K,推理中可交织工具调用。评测以编码与智能体为主,亮点如 CC-Bench 对 Claude Sonnet 4 胜率 48.6%、AIME 25 93.9%。
- 输入模态
- 文本
- 上下文
- 200K
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: CC-Bench paragraph after the coding benchmark image · row: CC-Bench · quote_snippet: GLM-4.6 improves over GLM-4.5 and reaches near parity with Claude Sonnet 4 (48.6% win rate)
{
"harness": "human evaluators working with models inside isolated Docker containers",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "win rate vs Claude Sonnet 4 in 1-on-1 human evaluation",
"judge": "human evaluators"
}new-benchmark: cc-bench (GLM real-world multi-turn coding bench: front-end development, tool building, data analysis, testing, algorithm) not yet in data/benchmarks/. Win rate is against Claude Sonnet 4 as the reference, not an absolute accuracy. Token efficiency disclosed in same paragraph: GLM-4.6 finishes tasks with about 15% fewer tokens than GLM-4.5. Evaluation details and trajectories published at https://huggingface.co/datasets/zai-org/CC-Bench-trajectories.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png ('We evaluated GLM-4.6 across eight public benchmarks') · table: benchmark comparison chart image coding_benchmark.png · row: AIME 25 · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context) · quote_snippet: We evaluated GLM-4.6 across eight public benchmarks covering agents, reasoning, and coding
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length (chart caption)",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed, per goal.md 12.5): AIME 25 GLM-4.6 93.9 base and 98.6 w/ Tools (two conditions printed in one cell); GLM-4.5 85.4, DeepSeek-V3.2-Exp 89.3, Claude Sonnet 4 74.3, Claude Sonnet 4.5 87.0. The base and w/-tools values are different protocols; split into variant rows when promoting to verified. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 93.9,同柱上方 w/ Tools 叠加值 98.6(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 85.4, DeepSeek-V3.2-Exp 89.3, Claude Sonnet 4 74.3, Claude Sonnet 4.5 87.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: GPQA · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): GPQA GLM-4.6 81.0 base / 82.9 w/ Tools; GLM-4.5 79.9, DeepSeek-V3.2-Exp 79.9, Sonnet 4 77.7, Sonnet 4.5 83.4. Variant (Diamond vs main) not resolvable from chart. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 81,同柱上方 w/ Tools 叠加值 82.9(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 79.9, DeepSeek-V3.2-Exp 79.9, Claude Sonnet 4 77.7, Claude Sonnet 4.5 83.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: LiveCodeBench v6 · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): LiveCodeBench v6 GLM-4.6 82.8 base / 84.5 w/ Tools; GLM-4.5 63.3, DeepSeek-V3.2-Exp 70.1, Sonnet 4 48.9, Sonnet 4.5 57.7. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 82.8,同柱上方 w/ Tools 叠加值 84.5(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 63.3, DeepSeek-V3.2-Exp 70.1, Claude Sonnet 4 48.9, Claude Sonnet 4.5 57.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: HLE · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): HLE GLM-4.6 17.2 base / 30.4 w/ Tools; GLM-4.5 14.4, DeepSeek-V3.2-Exp 19.8, Sonnet 4 9.6, Sonnet 4.5 17.3. Tool condition is a protocol split - do not merge base and w/-tools when promoting. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 17.2,同柱上方 w/ Tools 叠加值 30.4(工具条件单独口径,不与 base 合并)。竞品:GLM-4.5 14.4, DeepSeek-V3.2-Exp 19.8, Claude Sonnet 4 9.6, Claude Sonnet 4.5 17.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: BrowseComp · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): BrowseComp GLM-4.6 45.1; GLM-4.5 26.4, DeepSeek-V3.2-Exp 40.1, Sonnet 4 14.7, Sonnet 4.5 19.6. Single value (no w/-tools split on this row). 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 45.1。竞品:GLM-4.5 26.4, DeepSeek-V3.2-Exp 40.1, Claude Sonnet 4 14.7, Claude Sonnet 4.5 19.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: SWE-bench Verified · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context) · quote_snippet: but still lags behind Claude Sonnet 4.5 in coding ability
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): SWE-bench Verified GLM-4.6 68.0; GLM-4.5 64.2, DeepSeek-V3.2-Exp 67.8, Sonnet 4 72.5, Sonnet 4.5 77.2. Prose cross-check: page explicitly concedes 'still lags behind Claude Sonnet 4.5 in coding ability'. Harness not printed on this page (GLM-5.3 pages disclose Claude Code harness; do not assume the same here). 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 68。竞品:GLM-4.5 64.2, DeepSeek-V3.2-Exp 67.8, Claude Sonnet 4 72.5, Claude Sonnet 4.5 77.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: Terminal-Bench · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): Terminal-Bench GLM-4.6 40.5; GLM-4.5 37.5, DeepSeek-V3.2-Exp 37.7, Sonnet 4 35.5, Sonnet 4.5 50.0. Terminal-Bench version not printed. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 40.5。竞品:GLM-4.5 37.5, DeepSeek-V3.2-Exp 37.7, Claude Sonnet 4 35.5, Claude Sonnet 4.5 50.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: paragraph before coding_benchmark.png · table: benchmark comparison chart image coding_benchmark.png · row: tau^2-Bench (Weighted) · figure: images/02.png (archive of z-cdn.chatglm.cn z-blog/glm-4-6/coding_benchmark.png,8 分组柱状图,128K context)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "128K context length",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "weighted",
"judge": null
}Vision-assisted read (unconfirmed): tau^2-Bench (Weighted) GLM-4.6 75.9; GLM-4.5 67.5, DeepSeek-V3.2-Exp 53.4, Sonnet 4 66.0, Sonnet 4.5 88.1. Weighted aggregation over domains differs from per-domain Tau2 rows (retail/airline/telecom) in kimi-k2.json. 视觉转写自归档图 images/02.png(2026-09-01 复核,与先前读数一致)。GLM-4.6 柱为 base 值 75.9。竞品:GLM-4.5 67.5, DeepSeek-V3.2-Exp 53.4, Claude Sonnet 4 66.0, Claude Sonnet 4.5 88.1。