← 模型目录

GLM-5.3-Flash

Z.ai / 智谱 GLM · 2026-08-26 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GLM-5.3-Flash

GLM-5.3-Flash 以「Frontier Intelligence, Flash Cost」为定位发布,是 GLM-5 系列首个原生多模态模型(320B 总参/18B 激活,发布前以 ox-alpha 匿名测试)。已收录 16 项评测集中于终端编码、Agent 工具使用与视觉理解:亮点 Terminal-Bench 2.1 84.3、GDPval-AA v2 Elo 1773。

输入模态
文本 / 图像 / 视频
上下文
官方资料未说明
参数
320B-A18B MoE
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

terminalbench 84.3 模型 glm-5.3-flash · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Competitive Performance at Flash Cost / evaluation table / Coding · table: Terminal Bench 2.1 row · row: Terminal Bench 2.1 · quote_snippet: 84.3 | 81.0 | 83.9 | 85.0 | 87.4 | 85.8

{
  "harness": "Claude Code 2.1.207",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=65536",
  "turn_limit": null,
  "time_limit": "6h timeout",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Comparison columns same row: GLM-5.2 81.0, DeepSeek-V4-Vision-Exp 83.9, Opus 4.8 85.0, GPT-5.6 Terra 87.4, Gemini 3.7 Flash 85.8.

打开官方来源

deepswe 63.4 模型 glm-5.3-flash · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Competitive Performance at Flash Cost / evaluation table / Coding · table: DeepSWE v1.1 row · row: DeepSWE v1.1 · quote_snippet: 63.4 vs. 46.2 on DeepSWE v1.1

{
  "harness": "mini-swe-agent",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0.95,
  "top_p": 1,
  "token_budget": "400K context",
  "turn_limit": null,
  "time_limit": "timeout=6h",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepswe not yet in data/benchmarks/. In-text claim '63.4 vs. 46.2 on DeepSWE v1.1' matches the table. Comparison columns same row: GLM-5.2 46.2, DeepSeek-V4-Vision-Exp 59.3, Opus 4.8 58.0, GPT-5.6 Terra 69.6, Gemini 3.7 Flash 65.3.

打开官方来源

nl2repo 56.3 模型 glm-5.3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Competitive Performance at Flash Cost / evaluation table / Coding · table: NL2Repo row · row: NL2Repo · quote_snippet: 56.3 | 48.9 | 57.7 | 69.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=64k, 1M context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "rule-based checks plus LLM-based judgement (anti-hacking)"
}

new-benchmark: nl2repo not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 48.9, DeepSeek-V4-Vision-Exp 57.7, Opus 4.8 69.7.

打开官方来源

toolathlon 78.4 模型 glm-5.3-flash · 版本 Verified · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Agentic + Footnotes / Toolathlon Verified · table: Toolathlon Verified row · row: Toolathlon Verified · quote_snippet: pass@1 averaged over 3 independent runs

{
  "harness": "official evaluation service",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "pass@1 averaged over 3 independent runs",
  "judge": null
}

new-benchmark: toolathlon not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 59.9, DeepSeek-V4-Vision-Exp 75.9, Opus 4.8 76.2, GPT-5.6 Terra 74.9.

打开官方来源

automationbench 48.8 模型 glm-5.3-flash · 版本 v1.0.6 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Agentic + Footnotes / AutomationBench · table: AutomationBench v1.0.6 row · row: AutomationBench v1.0.6 · quote_snippet: 48.8 vs. 26.2 on AutomationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: automationbench not yet in data/benchmarks/. In-text claim '48.8 vs. 26.2 on AutomationBench' matches the table. Version pin v1.0.6 with zapier/AutomationBench PR #13 null-type fix. Comparison columns same row: GLM-5.2 26.2, DeepSeek-V4-Vision-Exp 38.8, Opus 4.8 41.0, GPT-5.6 Terra 37.2, Gemini 3.7 Flash 52.3.

打开官方来源

agents-last-exam 26.3 模型 glm-5.3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Agentic + Footnotes / Agent's Last Exam · table: Agents' Last Exam row · row: Agents' Last Exam · quote_snippet: official evaluation protocol with the Claude Code harness (reasoning effort=max)

{
  "harness": "Claude Code",
  "tools": "Tool Search disabled",
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "1M context, 64K maximum output",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "official ALE evaluators"
}

new-benchmark: agents-last-exam not yet in data/benchmarks.json. Variant on this release page is unlabeled; GLM-5.3 labels the equivalent row ALE-CLI. Comparison columns same row: GLM-5.2 20.4, DeepSeek-V4-Vision-Exp 27.3, Opus 4.8 27.0, GPT-5.6 Terra 28.0.

打开官方来源

hlehle 55.3 模型 glm-5.3-flash · 版本 w/ Tools (full set) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Agentic + Footnotes / HLE w/ tools (full set) · table: HLE w/ Tools row · row: HLE w/ Tools · quote_snippet: temperature=1.0 and top_p=0.95 ... maximum context length of 300,000 tokens

{
  "harness": null,
  "tools": "tools enabled; context management strategy",
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max generation length 163,840 tokens; max context 300,000 tokens",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "GPT-5.6-luna (medium)"
}

Maps to existing benchmark hlehle (Humanity's Last Exam), tool-augmented variant. Comparison columns same row: GLM-5.2 54.7, DeepSeek-V4-Vision-Exp 55.1, Opus 4.8 57.9.

打开官方来源

gdpval-aa 1773 模型 glm-5.3-flash · 版本 v2 · 指标 未说明 · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Agentic + Footnotes / GDPval-AA v2 · table: GDPval-AA v2 row · row: GDPval-AA v2 · quote_snippet: GDPval-AA v2: Models are evaluated by Artificial Analysis.

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: gdpval-aa not yet in data/benchmarks.json. Third-party measurement (Artificial Analysis) transcribed in vendor table per footnote. Comparison columns same row: GLM-5.2 1504, DeepSeek-V4-Vision-Exp 1675, Opus 4.8 1582, GPT-5.6 Terra 1571, Gemini 3.7 Flash 1527.

打开官方来源

officeqa-pro 62.4 模型 glm-5.3-flash · 版本 Treasury Bulletin corpus · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Vision + Footnotes / OfficeQA Pro · table: OfficeQA Pro row · row: OfficeQA Pro · quote_snippet: Treasury Bulletin PDF corpus without providing access to embedded text

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "maximum context length 512K tokens",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: officeqa-pro not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 57.9, Opus 4.8 48.9 (GLM-5.2, GPT-5.6 Terra, Gemini 3.7 Flash not run).

打开官方来源

charxiv-reasoning 89.4 模型 glm-5.3-flash · 版本 w/ Tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Vision + Footnotes / CharXiv Reasoning · table: CharXiv Reasoning w/ Tools row · row: CharXiv Reasoning w/ Tools · quote_snippet: temperature=1.0, top_p=0.95, and a maximum context length of 256K tokens

{
  "harness": null,
  "tools": "tools enabled",
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "maximum context length 256K tokens",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: charxiv-reasoning (CharXiv reasoning split) not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 80.4, Opus 4.8 89.9, GPT-5.6 Terra 88.0, Gemini 3.7 Flash 88.7.

打开官方来源

chartography 78.0 模型 glm-5.3-flash · 版本 w/ Tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Vision + Footnotes / Chartography · table: Chartography w/ Tools row · row: Chartography w/ Tools · quote_snippet: temperature=1.0, top_p=0.95, and a maximum context length of 256K tokens

{
  "harness": null,
  "tools": "tools enabled",
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "maximum context length 256K tokens",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: chartography not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 64.3, Opus 4.8 75.0, GPT-5.6 Terra 68.0, Gemini 3.7 Flash 65.0.

打开官方来源

babyvision 53.4 模型 glm-5.3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Vision + Footnotes / BabyVision · table: BabyVision row · row: BabyVision · quote_snippet: shorter side is at least 1.5K pixels, consistent with other baselines

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "maximum context length 164K tokens",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: babyvision not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 35.1, Opus 4.8 46.8, GPT-5.6 Terra 61.6, Gemini 3.7 Flash 70.9.

打开官方来源

mvbench 77.8 模型 glm-5.3-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Vision + Footnotes / MVBench and MMVU · table: MVbench row · row: MVbench · quote_snippet: default 1 fps frame-extraction strategy for models without native video input

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "maximum context length 256K tokens; 1 fps frame extraction when video input unsupported, uniform sampling at API max frame count",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mvbench not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 69.4, Opus 4.8 67.1, GPT-5.6 Terra 75.0, Gemini 3.7 Flash 82.2.

打开官方来源

mmvu 80.5 模型 glm-5.3-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: evaluation table / Vision + Footnotes / MVBench and MMVU · table: MMVU row · row: MMVU · quote_snippet: MMVU ... temperature=1.0, top_p=0.95, and a maximum context length of 256K tokens

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "maximum context length 256K tokens; 1 fps frame extraction fallback",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmvu not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 72.7, Opus 4.8 67.4, GPT-5.6 Terra 75.8, Gemini 3.7 Flash 82.3.

打开官方来源

aa-intelligence-index 57 模型 glm-5.3-flash · 版本 v4.1.1 · 指标 composite_intelligence_index · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Competitive Performance at Flash Cost · quote_snippet: Artificial Analysis Intelligence Index v4.1.1, scoring 57 at just $0.045 per task

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "$0.045 per task (discounted)",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "composite index",
  "judge": null
}

new-benchmark: aa-intelligence-index (Artificial Analysis Intelligence Index, composite of multiple benchmarks) not yet in data/benchmarks.json. Third-party index; score quoted by vendor.

打开官方来源

zai-code-bench 29.0 模型 glm-5.3-flash · 版本 v1.0, max effort · 指标 task_completion_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Competitive Performance at Flash Cost · figure: Chart image img_v3_0214u_eef3c372-bc89-44cd-9ec3-919ff63bac0g; numeric claim also printed in body text · quote_snippet: at max effort nearly matches Claude Opus 4.8 (29.0 vs. 29.5)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "end-to-end task completion rate",
  "judge": null
}

new-benchmark: zai-code-bench (Z.ai in-house coding evaluation, distinct from Kimi's KCB) not yet in data/benchmarks.json. Numeric claim printed in body text; underlying chart is an image.

打开官方来源