GLM-5.3-Flash
Z.ai / 智谱 GLM · 2026-08-26 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
GLM-5.3-Flash
GLM-5.3-Flash 以「Frontier Intelligence, Flash Cost」为定位发布,是 GLM-5 系列首个原生多模态模型(320B 总参/18B 激活,发布前以 ox-alpha 匿名测试)。已收录 16 项评测集中于终端编码、Agent 工具使用与视觉理解:亮点 Terminal-Bench 2.1 84.3、GDPval-AA v2 Elo 1773。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 官方资料未说明
- 参数
- 320B-A18B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Competitive Performance at Flash Cost / evaluation table / Coding · table: Terminal Bench 2.1 row · row: Terminal Bench 2.1 · quote_snippet: 84.3 | 81.0 | 83.9 | 85.0 | 87.4 | 85.8
{
"harness": "Claude Code 2.1.207",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=65536",
"turn_limit": null,
"time_limit": "6h timeout",
"run_count": null,
"aggregation": null,
"judge": null
}Comparison columns same row: GLM-5.2 81.0, DeepSeek-V4-Vision-Exp 83.9, Opus 4.8 85.0, GPT-5.6 Terra 87.4, Gemini 3.7 Flash 85.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Competitive Performance at Flash Cost / evaluation table / Coding · table: DeepSWE v1.1 row · row: DeepSWE v1.1 · quote_snippet: 63.4 vs. 46.2 on DeepSWE v1.1
{
"harness": "mini-swe-agent",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 0.95,
"top_p": 1,
"token_budget": "400K context",
"turn_limit": null,
"time_limit": "timeout=6h",
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepswe not yet in data/benchmarks/. In-text claim '63.4 vs. 46.2 on DeepSWE v1.1' matches the table. Comparison columns same row: GLM-5.2 46.2, DeepSeek-V4-Vision-Exp 59.3, Opus 4.8 58.0, GPT-5.6 Terra 69.6, Gemini 3.7 Flash 65.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Competitive Performance at Flash Cost / evaluation table / Coding · table: NL2Repo row · row: NL2Repo · quote_snippet: 56.3 | 48.9 | 57.7 | 69.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 1,
"token_budget": "max_new_tokens=64k, 1M context",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "rule-based checks plus LLM-based judgement (anti-hacking)"
}new-benchmark: nl2repo not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 48.9, DeepSeek-V4-Vision-Exp 57.7, Opus 4.8 69.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Agentic + Footnotes / Toolathlon Verified · table: Toolathlon Verified row · row: Toolathlon Verified · quote_snippet: pass@1 averaged over 3 independent runs
{
"harness": "official evaluation service",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "pass@1 averaged over 3 independent runs",
"judge": null
}new-benchmark: toolathlon not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 59.9, DeepSeek-V4-Vision-Exp 75.9, Opus 4.8 76.2, GPT-5.6 Terra 74.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Agentic + Footnotes / AutomationBench · table: AutomationBench v1.0.6 row · row: AutomationBench v1.0.6 · quote_snippet: 48.8 vs. 26.2 on AutomationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: automationbench not yet in data/benchmarks/. In-text claim '48.8 vs. 26.2 on AutomationBench' matches the table. Version pin v1.0.6 with zapier/AutomationBench PR #13 null-type fix. Comparison columns same row: GLM-5.2 26.2, DeepSeek-V4-Vision-Exp 38.8, Opus 4.8 41.0, GPT-5.6 Terra 37.2, Gemini 3.7 Flash 52.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Agentic + Footnotes / Agent's Last Exam · table: Agents' Last Exam row · row: Agents' Last Exam · quote_snippet: official evaluation protocol with the Claude Code harness (reasoning effort=max)
{
"harness": "Claude Code",
"tools": "Tool Search disabled",
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": "1M context, 64K maximum output",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "official ALE evaluators"
}new-benchmark: agents-last-exam not yet in data/benchmarks.json. Variant on this release page is unlabeled; GLM-5.3 labels the equivalent row ALE-CLI. Comparison columns same row: GLM-5.2 20.4, DeepSeek-V4-Vision-Exp 27.3, Opus 4.8 27.0, GPT-5.6 Terra 28.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Agentic + Footnotes / HLE w/ tools (full set) · table: HLE w/ Tools row · row: HLE w/ Tools · quote_snippet: temperature=1.0 and top_p=0.95 ... maximum context length of 300,000 tokens
{
"harness": null,
"tools": "tools enabled; context management strategy",
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "max generation length 163,840 tokens; max context 300,000 tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "GPT-5.6-luna (medium)"
}Maps to existing benchmark hlehle (Humanity's Last Exam), tool-augmented variant. Comparison columns same row: GLM-5.2 54.7, DeepSeek-V4-Vision-Exp 55.1, Opus 4.8 57.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Agentic + Footnotes / GDPval-AA v2 · table: GDPval-AA v2 row · row: GDPval-AA v2 · quote_snippet: GDPval-AA v2: Models are evaluated by Artificial Analysis.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gdpval-aa not yet in data/benchmarks.json. Third-party measurement (Artificial Analysis) transcribed in vendor table per footnote. Comparison columns same row: GLM-5.2 1504, DeepSeek-V4-Vision-Exp 1675, Opus 4.8 1582, GPT-5.6 Terra 1571, Gemini 3.7 Flash 1527.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Vision + Footnotes / OfficeQA Pro · table: OfficeQA Pro row · row: OfficeQA Pro · quote_snippet: Treasury Bulletin PDF corpus without providing access to embedded text
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "maximum context length 512K tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: officeqa-pro not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 57.9, Opus 4.8 48.9 (GLM-5.2, GPT-5.6 Terra, Gemini 3.7 Flash not run).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Vision + Footnotes / CharXiv Reasoning · table: CharXiv Reasoning w/ Tools row · row: CharXiv Reasoning w/ Tools · quote_snippet: temperature=1.0, top_p=0.95, and a maximum context length of 256K tokens
{
"harness": null,
"tools": "tools enabled",
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "maximum context length 256K tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: charxiv-reasoning (CharXiv reasoning split) not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 80.4, Opus 4.8 89.9, GPT-5.6 Terra 88.0, Gemini 3.7 Flash 88.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Vision + Footnotes / Chartography · table: Chartography w/ Tools row · row: Chartography w/ Tools · quote_snippet: temperature=1.0, top_p=0.95, and a maximum context length of 256K tokens
{
"harness": null,
"tools": "tools enabled",
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "maximum context length 256K tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: chartography not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 64.3, Opus 4.8 75.0, GPT-5.6 Terra 68.0, Gemini 3.7 Flash 65.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Vision + Footnotes / BabyVision · table: BabyVision row · row: BabyVision · quote_snippet: shorter side is at least 1.5K pixels, consistent with other baselines
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "maximum context length 164K tokens",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: babyvision not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 35.1, Opus 4.8 46.8, GPT-5.6 Terra 61.6, Gemini 3.7 Flash 70.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Vision + Footnotes / MVBench and MMVU · table: MVbench row · row: MVbench · quote_snippet: default 1 fps frame-extraction strategy for models without native video input
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "maximum context length 256K tokens; 1 fps frame extraction when video input unsupported, uniform sampling at API max frame count",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mvbench not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 69.4, Opus 4.8 67.1, GPT-5.6 Terra 75.0, Gemini 3.7 Flash 82.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: evaluation table / Vision + Footnotes / MVBench and MMVU · table: MMVU row · row: MMVU · quote_snippet: MMVU ... temperature=1.0, top_p=0.95, and a maximum context length of 256K tokens
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": 1,
"top_p": 0.95,
"token_budget": "maximum context length 256K tokens; 1 fps frame extraction fallback",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmvu not yet in data/benchmarks.json. Comparison columns same row: DeepSeek-V4-Vision-Exp 72.7, Opus 4.8 67.4, GPT-5.6 Terra 75.8, Gemini 3.7 Flash 82.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Competitive Performance at Flash Cost · quote_snippet: Artificial Analysis Intelligence Index v4.1.1, scoring 57 at just $0.045 per task
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "$0.045 per task (discounted)",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "composite index",
"judge": null
}new-benchmark: aa-intelligence-index (Artificial Analysis Intelligence Index, composite of multiple benchmarks) not yet in data/benchmarks.json. Third-party index; score quoted by vendor.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Competitive Performance at Flash Cost · figure: Chart image img_v3_0214u_eef3c372-bc89-44cd-9ec3-919ff63bac0g; numeric claim also printed in body text · quote_snippet: at max effort nearly matches Claude Opus 4.8 (29.0 vs. 29.5)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "end-to-end task completion rate",
"judge": null
}new-benchmark: zai-code-bench (Z.ai in-house coding evaluation, distinct from Kimi's KCB) not yet in data/benchmarks.json. Numeric claim printed in body text; underlying chart is an image.