← 模型目录

GLM-5.3

Z.ai / 智谱 GLM · 2026-08(仅精确到月) · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GLM-5.3

GLM-5.3 以「Frontier Coding with Emergent Cyber Capabilities」为定位发布(API thinking 档 low/high/max、默认 max 且不再支持关闭思考;权重承诺发布两周内开源)。已收录 22 项评测集中于编码与网络安全攻防:亮点 Terminal-Bench 2.1 88.2、CyberGym 84.5。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

terminalbench 88.2 模型 glm-5.3 · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: Terminal Bench 2.1 row · row: Terminal Bench 2.1 · quote_snippet: 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8

{
  "harness": "Claude Code 2.1.207",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=65536",
  "turn_limit": null,
  "time_limit": "6h timeout",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Comparison columns in same row (other protocols/harnesses per their vendors): GLM-5.2 81.0, Kimi K3 88.3, DeepSeek-V4 Pro-0813 87.9, Qwen3.8-Max 86.6, Opus 4.8 85.0, Fable 5 w/ fallback 88.0, GPT-5.6 Sol 88.8.

打开官方来源

terminalbench 28.3 模型 glm-5.3 · 版本 3.0 · 指标 avg@3 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: Terminal Bench 3.0 row · row: Terminal Bench 3.0 · quote_snippet: 28.3 | 4.6 | 17.4 | 21.1 | 33.7 | 34.6

{
  "harness": "Claude Code 2.1.207",
  "tools": "Tool Search disabled",
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "400K context, 128K max output",
  "turn_limit": 600,
  "time_limit": "10h per rollout",
  "run_count": 3,
  "aggregation": "avg@3 over three rollouts per task",
  "judge": "task's official separate verifier"
}

Each rollout in an isolated container from the task's official image. Comparison columns same row: GLM-5.2 4.6, Kimi K3 17.4, Opus 4.8 21.1, Fable 5 w/ fallback 33.7, GPT-5.6 Sol 34.6. Not comparable to Terminal-Bench 2.1 rows (different benchmark version and protocol).

打开官方来源

deepswe 66.9 模型 glm-5.3 · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: DeepSWE v1.1 row · row: DeepSWE v1.1 · quote_snippet: 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7

{
  "harness": "mini-swe-agent",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 0.95,
  "top_p": 1,
  "token_budget": "400K context",
  "turn_limit": null,
  "time_limit": "timeout=6h",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepswe not yet in data/benchmarks/. Comparison columns same row: GLM-5.2 46.2, Kimi K3 67.5, DeepSeek-V4 Pro-0813 62.7, Qwen3.8-Max 56.6, Opus 4.8 58.0, Fable 5 w/ fallback 69.7, GPT-5.6 Sol 72.7. Kimi's own footnote reports 67.3 under mini-SWE-agent; 67.5 vs 67.3 discrepancy between the two vendor pages is unresolved.

打开官方来源

nl2repo 58.0 模型 glm-5.3 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: NL2Repo row · row: NL2Repo · quote_snippet: 58.0 | 48.9 | 58.0 | 61.1 | 55.9 | 69.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=64k, 1M context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "rule-based checks plus LLM-based judgement (anti-hacking)"
}

new-benchmark: nl2repo not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 48.9, Kimi K3 58.0, DeepSeek-V4 Pro-0813 61.1, Qwen3.8-Max 55.9, Opus 4.8 69.7.

打开官方来源

program-bench 19.0 模型 glm-5.3 · 版本 Almost Solved · 指标 almost_solved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: ProgramBench Almost Solved row · row: ProgramBench Almost Solved · quote_snippet: 19.0 | 9.5 | 17.5 | 10.5 | 15.5 | 33.0 | 23.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: program-bench not yet in data/benchmarks.json (same family as Kimi's Program Bench row). Comparison columns same row: GLM-5.2 9.5, Kimi K3 17.5, Qwen3.8-Max 10.5, Opus 4.8 15.5, Fable 5 w/ fallback 33.0, GPT-5.6 Sol 23.0.

打开官方来源

frontierswe 78.1 模型 glm-5.3 · 版本 未说明 · 指标 dominance_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: FrontierSWE row · row: FrontierSWE · quote_snippet: The evaluation was conducted by Proximal with 1M context length, max effort level

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "1M context, 128K max output",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "dominance score",
  "judge": null
}

new-benchmark: frontierswe not yet in data/benchmarks/. Evaluation conducted by Proximal per footnote; dominance score reported as of 2026/08/14. Comparison columns same row: GLM-5.2 67.5, Opus 4.8 66.5, Fable 5 w/ fallback 88.2.

打开官方来源

swe-marathon 42.5 模型 glm-5.3 · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: SWE-Marathon v1.1 row · row: SWE-Marathon v1.1 · quote_snippet: 42.5 | 19.4 | 48.1 | 48.8 | 33.1 | 42.5

{
  "harness": "Claude Code 2.1.207",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max_new_tokens=128000, 1M context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "LLM-based inspection replacing removed anti-cheat pattern checks on strip-clone"
}

new-benchmark: swe-marathon not yet in data/benchmarks.json. Footnote discloses protocol deviations: anti-cheat checks removed on strip-clone (LLM inspection instead), extra PyPI index added for parameter-golf and trimul-cuda Docker builds. Kimi's SWE-Marathon row uses an H20-recalibrated branch - variants differ across vendors.

打开官方来源

posttrain-bench 39.8 模型 glm-5.3 · 版本 未说明 · 指标 weighted_average · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Coding · table: PostTrainBench row · row: PostTrainBench · quote_snippet: weighted average over 3 runs; failed runs fall back to official zero-shot baseline

{
  "harness": "Claude Code 2.1.207",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=128000, 1M context",
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "weighted average over 3 runs; failed runs fall back to official zero-shot base-model baseline",
  "judge": "LLM agent inspects solutions for external API usage (replaces removed pattern-matching checks)"
}

new-benchmark: posttrain-bench not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 31.7, Kimi K3 32.0, Opus 4.8 32.9, Fable 5 w/ fallback 41.8, GPT-5.6 Sol 36.2.

打开官方来源

cybergym 84.5 模型 glm-5.3 · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Cyber + Footnotes / CyberGym · table: CyberGym row · row: CyberGym · quote_snippet: GLM-5.3 scores 84.5%, up from GLM-5.2's 77.2% - the best result on the benchmark

{
  "harness": "Claude Code 2.1.207",
  "tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=128000",
  "turn_limit": null,
  "time_limit": "unlimited timeout per task",
  "run_count": 1,
  "aggregation": "single-run Pass@1 over 1,507 tasks",
  "judge": null
}

new-benchmark: cybergym not yet in data/benchmarks.json. Agent placed inside the task container; Git info removed. Comparison columns same row: GLM-5.2 77.2, Kimi K3 80.0, DeepSeek-V4 Pro-0813 83.3, Qwen3.8-Max 78.5, Opus 4.8 78.1, Fable 5 w/ fallback 83.8, GPT-5.6 Sol 83.6.

打开官方来源

exploitgym 105 模型 glm-5.3 · 版本 2h budget · 指标 tasks_completed_within_budget · 单位 tasks_completed 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Cyber + Footnotes / ExploitGym · table: ExploitGym 2h / 6h row · row: ExploitGym 2h / 6h · quote_snippet: GLM-5.3 completes 105 tasks within two hours

{
  "harness": "Claude Code 2.1.207",
  "tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=128000",
  "turn_limit": null,
  "time_limit": "2h budget (API inference time rescaled by per-model TPS plus non-API overhead)",
  "run_count": 1,
  "aggregation": "single-run Pass@1 on 869 tasks",
  "judge": null
}

new-benchmark: exploitgym not yet in data/benchmarks.json. Time budgets normalized across models via per-model tokens-per-second (GLM-5.3 115 TPS, Kimi K3 40 TPS, Qwen3.8 Max 47 TPS per footnote). Comparison columns same row 2h values: GLM-5.2 29, Kimi K3 36, Qwen3.8-Max 14, Opus 4.8 80, Fable 5 w/ fallback 181, GPT-5.6 Sol 216.

打开官方来源

exploitgym 130 模型 glm-5.3 · 版本 6h budget · 指标 tasks_completed_within_budget · 单位 tasks_completed 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Cyber + Footnotes / ExploitGym · table: ExploitGym 2h / 6h row · row: ExploitGym 2h / 6h · quote_snippet: and 130 within six hours

{
  "harness": "Claude Code 2.1.207",
  "tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=128000",
  "turn_limit": null,
  "time_limit": "6h budget (API inference time rescaled by per-model TPS plus non-API overhead)",
  "run_count": 1,
  "aggregation": "single-run Pass@1 on 869 tasks",
  "judge": null
}

new-benchmark: exploitgym not yet in data/benchmarks/. Comparison columns same row 6h values: GLM-5.2 39, Kimi K3 70, Qwen3.8-Max 26, Opus 4.8 120, Fable 5 w/ fallback 247, GPT-5.6 Sol 293.

打开官方来源

exploitbench 54.4 模型 glm-5.3 · 版本 未说明 · 指标 coverage_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Cyber + Footnotes / ExploitBench · table: ExploitBench row · row: ExploitBench · quote_snippet: GLM-5.3 reaches 54.4%, more than doubling GLM-5.2's 24.4%

{
  "harness": "Claude Code 2.1.207",
  "tools": "no web tools; domain whitelist (pypi.org, deb.debian.org)",
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 1,
  "token_budget": "max_new_tokens=128000",
  "turn_limit": 300,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "average coverage score over 41 tasks across 3 revisions; per-task result is union of capabilities across revisions",
  "judge": null
}

new-benchmark: exploitbench not yet in data/benchmarks/. Comparison columns same row: GLM-5.2 24.4, Kimi K3 32.2, Qwen3.8-Max 28.8, Opus 4.8 40.0, Fable 5 w/ fallback 78.0, GPT-5.6 Sol 76.5.

打开官方来源

toolathlon 73.0 模型 glm-5.3 · 版本 Verified · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Agentic + Footnotes / Toolathlon Verified · table: Toolathlon Verified row · row: Toolathlon Verified · quote_snippet: official evaluation service; pass@1 averaged over 3 independent runs

{
  "harness": "official evaluation service",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "pass@1 averaged over 3 independent runs",
  "judge": null
}

new-benchmark: toolathlon not yet in data/benchmarks.json. Comparison columns same row: GLM-5.2 59.9, Kimi K3 76.5, DeepSeek-V4 Pro-0813 74.1, Qwen3.8-Max 72.5, Opus 4.8 76.2, Fable 5 w/ fallback 74.7, GPT-5.6 Sol 74.9.

打开官方来源

automationbench 48.2 模型 glm-5.3 · 版本 v1.0.6 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Agentic + Footnotes / AutomationBench · table: AutomationBench v1.0.6 row · row: AutomationBench v1.0.6 · quote_snippet: AutomationBench v1.0.6, incorporating the fix ... in PR #13

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: automationbench not yet in data/benchmarks.json. Version pin v1.0.6 with zapier/AutomationBench PR #13 null-type fix. Comparison columns same row: GLM-5.2 26.2, Kimi K3 46.7, DeepSeek-V4 Pro-0813 43.2, Qwen3.8-Max 39.8, Opus 4.8 41.0, Fable 5 w/ fallback 46.2, GPT-5.6 Sol 45.8.

打开官方来源

agents-last-exam 28.5 模型 glm-5.3 · 版本 ALE-CLI · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Agentic + Footnotes / Agent's Last Exam (CLI) · table: Agents' Last Exam ALE-CLI row · row: Agents' Last Exam ALE-CLI · quote_snippet: official evaluation protocol with the Claude Code harness ... 105 tasks

{
  "harness": "Claude Code",
  "tools": "Tool Search disabled",
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": "1M context, 64K maximum output",
  "turn_limit": null,
  "time_limit": "default 4h, task-specific up to 8h",
  "run_count": null,
  "aggregation": null,
  "judge": "official ALE evaluators"
}

new-benchmark: agents-last-exam not yet in data/benchmarks.json. 105 tasks, each in isolated Docker container using Task Card resources. Comparison columns same row: GLM-5.2 23.8, Kimi K3 27.6, DeepSeek-V4 Pro-0813 25.7, Qwen3.8-Max 27.0, Opus 4.8 25.7, Fable 5 w/ fallback 23.8, GPT-5.6 Sol 28.6.

打开官方来源

hlehle 62.5 模型 glm-5.3 · 版本 w/ Tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Agentic + Footnotes / HLE w/ tools · table: HLE w/ Tools row · row: HLE w/ Tools · quote_snippet: temperature=1.0 and top_p=0.95 ... maximum context length of 300,000 tokens

{
  "harness": null,
  "tools": "tools enabled; context management strategy",
  "shots": null,
  "reasoning_effort": null,
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": "max generation length 163,840 tokens; max context 300,000 tokens",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "GPT-5.6-luna (medium)"
}

Maps to existing benchmark hlehle (Humanity's Last Exam), tool-augmented variant with LLM judge. Comparison columns same row: GLM-5.2 54.7, Kimi K3 59.8, DeepSeek-V4 Pro-0813 60.0, Qwen3.8-Max 56.2, Opus 4.8 57.9, Fable 5 w/ fallback 63.9, GPT-5.6 Sol 64.5.

打开官方来源

gdpval-aa 1769 模型 glm-5.3 · 版本 v2 · 指标 未说明 · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance across comparison models / Agentic + Footnotes / GDPval-AA v2 · table: GDPval-AA v2 row · row: GDPval-AA v2 · quote_snippet: GDPval-AA v2: Models are evaluated by Artificial Analysis.

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: gdpval-aa not yet in data/benchmarks.json. Footnote states all models in this row are evaluated by Artificial Analysis, so this is a third-party measurement transcribed in the vendor table, not a vendor-run score. Comparison columns same row: GLM-5.2 1508, Kimi K3 1682, DeepSeek-V4 Pro-0813 1590, Qwen3.8-Max 1739, Opus 4.8 1588, Fable 5 w/ fallback 1743, GPT-5.6 Sol 1730.

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

cybergym 83.8% 模型 claude-mythos-5 · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Emergent Cyber Capability · row: CyberGym · quote_snippet: ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cybergym not yet in data/benchmarks/. Competitor score cited by Z.ai; protocol under which Mythos 5 was run is not disclosed on this page, so the comparison is not proven same-protocol.

打开官方来源

cybergym 83.6% 模型 gpt-5-6-sol · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Emergent Cyber Capability · row: CyberGym · quote_snippet: ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cybergym not yet in data/benchmarks/. Competitor score cited by Z.ai; GPT-5.6 Sol run protocol not disclosed on this page.

打开官方来源

exploitbench 78.0% 模型 claude-mythos-5 · 版本 未说明 · 指标 coverage_score · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Emergent Cyber Capability · row: ExploitBench · quote_snippet: Mythos 5 and GPT-5.6 Sol score 78.0% and 76.5%, respectively

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitbench not yet in data/benchmarks/. Competitor score cited by Z.ai; protocol for the cited run not disclosed on this page.

打开官方来源

exploitbench 76.5% 模型 gpt-5-6-sol · 版本 未说明 · 指标 coverage_score · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Emergent Cyber Capability · row: ExploitBench · quote_snippet: Mythos 5 and GPT-5.6 Sol score 78.0% and 76.5%, respectively

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitbench not yet in data/benchmarks/. Competitor score cited by Z.ai; protocol for the cited run not disclosed on this page.

打开官方来源

exploitgym 181 / 247 模型 claude-mythos-5 · 版本 2h / 6h budgets · 指标 tasks_completed_within_budget · 单位 tasks_completed 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Emergent Cyber Capability · row: ExploitGym · quote_snippet: Mythos 5 remains well ahead at 181 and 247 tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "2h and 6h budgets",
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: exploitgym not yet in data/benchmarks/. Competitor score cited by Z.ai for both budget variants (181 tasks at 2h, 247 at 6h); value field stores the 2h figure, display shows both.

打开官方来源