← 模型目录

Grok 4.5

xAI / Grok · 2026-07-16 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Grok 4.5

xAI 发布 Grok 4.5,定位为面向编码、智能体任务与知识工作的最强模型,并与 Cursor 联合训练。评测集中于智能体编码,亮点如 DeepSWE 1.0 62.0%、Terminal-Bench 2.1 83.3%,并称平均输出 tokens 约为 Opus 4.8 (max) 的 1/4.2。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 2 / 输出 6

本变体的评测证据

deepswe 62.0% 模型 grok-4-5 · 版本 1.0 (Datacurve eval, run by AA with each provider's harness) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Real-world engineering excellence · figure: Benchmark bar charts; full numeric values present in the page DOM as chart accessibility text · quote_snippet: DeepSWE 1.0... Fable (max) 66.1%, GPT 5.5 (xhigh) 64.31%, Grok 4.5 62.0%

{
  "harness": "each model provider's own harness (eval created by Datacurve, run by AA)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id deepswe (registered batch 1). Harness caveat is explicit on the page: each provider's own harness ran the shared Datacurve eval, so this is not a same-harness comparison. Competitor bars (chart accessibility text): Fable max 66.1 / GPT 5.5 xhigh 64.31 / Opus 4.8 max 55.75 / Opus 4.7 max 40.12.

打开官方来源

deepswe 53% 模型 grok-4-5 · 版本 1.1 (mini-swe-agent harness, run by Datacurve) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: DeepSWE 1.1 (mini-swe-agent harness run by Datacurve): Fable (max) 70%, GPT 5.5 67%, Opus 4.8 59%, Grok 4.5 53%

{
  "harness": "mini-swe-agent harness (run by Datacurve)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Different harness than the 1.0 row (same-agent mini-swe-agent for all models — closer to same-protocol). Competitors: Fable 70 / GPT 5.5 67 / Opus 4.8 59 / GLM 5.2 44. Google's 3.7 Flash DeepSWE v1.1 65.3% is a vendor-self-run number on a different setup — align variant+harness before comparing.

打开官方来源

swe-marathon 29.0% 模型 grok-4-5 · 版本 未说明 · 指标 resolution_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: SWE Marathon resolution rate (pass@1): Grok 4.5 29.0%, Opus 4.8 (max) 26.0%, Fable (max) 24.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1",
  "judge": null
}

Reuses candidate id swe-marathon (registered batch 1; SWE-Marathon branch ambiguity flagged in batch-1 README item 5 — GLM official v1.1 vs Kimi H20 recalibration). Competitors: Opus 4.8 26.0 / Fable 24.0 / Opus 4.7 16.0.

打开官方来源

terminalbench 83.3% 模型 grok-4-5 · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: Terminal Bench 2.1: Fable (max) 84.3%, GPT 5.5 (xhigh) 83.4%, Grok 4.5 83.3%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark terminalbench, variant 2.1. Competitors: Fable 84.3 / GPT 5.5 83.4 / Opus 4.8 78.9 / Opus 4.7 78.9. Harness not stated for the Grok row (competitor harnesses unknown — comparability caution).

打开官方来源

swebench-pro 64.7% 模型 grok-4-5 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: SWE Bench Pro resolve rate: Fable (max) 80.4%, Opus 4.8 (max) 69.2%, Grok 4.5 64.7%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id swebench-pro (registered batch 2). Competitors: Fable 80.4 / Opus 4.8 69.2 / Opus 4.7 64.3 / GLM 5.2 62.1 / GPT 5.5 58.6. Google's model pages label this benchmark 'SWE-Bench Pro (Public)' — same-family id, align variant at migration.

打开官方来源

swebench-pro 15,954 output tokens avg per SWE Bench Pro task 模型 grok-4-5 · 版本 token efficiency (avg output tokens per task) · 指标 avg_output_tokens_per_task · 单位 tokens 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Faster than flash models · figure: Token efficiency chart; values restated in prose · quote_snippet: Grok 4.5 resolves tasks with 15,954 output tokens on average, about 4.2x fewer than Opus 4.8 (max) at 67,020

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "average across SWE Bench Pro tasks",
  "judge": null
}

Efficiency metric, not an accuracy score — unit is tokens (lower is better). Competitor value in prose: Opus 4.8 (max) 67,020 (4.2x more). Recorded because it is the page's headline cost/quality claim and precisely printed.

打开官方来源