Grok 4.5
xAI / Grok · 2026-07-16 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Grok 4.5
xAI 发布 Grok 4.5,定位为面向编码、智能体任务与知识工作的最强模型,并与 Cursor 联合训练。评测集中于智能体编码,亮点如 DeepSWE 1.0 62.0%、Terminal-Bench 2.1 83.3%,并称平均输出 tokens 约为 Opus 4.8 (max) 的 1/4.2。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 2 / 输出 6
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-world engineering excellence · figure: Benchmark bar charts; full numeric values present in the page DOM as chart accessibility text · quote_snippet: DeepSWE 1.0... Fable (max) 66.1%, GPT 5.5 (xhigh) 64.31%, Grok 4.5 62.0%
{
"harness": "each model provider's own harness (eval created by Datacurve, run by AA)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id deepswe (registered batch 1). Harness caveat is explicit on the page: each provider's own harness ran the shared Datacurve eval, so this is not a same-harness comparison. Competitor bars (chart accessibility text): Fable max 66.1 / GPT 5.5 xhigh 64.31 / Opus 4.8 max 55.75 / Opus 4.7 max 40.12.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: DeepSWE 1.1 (mini-swe-agent harness run by Datacurve): Fable (max) 70%, GPT 5.5 67%, Opus 4.8 59%, Grok 4.5 53%
{
"harness": "mini-swe-agent harness (run by Datacurve)",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Different harness than the 1.0 row (same-agent mini-swe-agent for all models — closer to same-protocol). Competitors: Fable 70 / GPT 5.5 67 / Opus 4.8 59 / GLM 5.2 44. Google's 3.7 Flash DeepSWE v1.1 65.3% is a vendor-self-run number on a different setup — align variant+harness before comparing.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: SWE Marathon resolution rate (pass@1): Grok 4.5 29.0%, Opus 4.8 (max) 26.0%, Fable (max) 24.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1",
"judge": null
}Reuses candidate id swe-marathon (registered batch 1; SWE-Marathon branch ambiguity flagged in batch-1 README item 5 — GLM official v1.1 vs Kimi H20 recalibration). Competitors: Opus 4.8 26.0 / Fable 24.0 / Opus 4.7 16.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: Terminal Bench 2.1: Fable (max) 84.3%, GPT 5.5 (xhigh) 83.4%, Grok 4.5 83.3%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark terminalbench, variant 2.1. Competitors: Fable 84.3 / GPT 5.5 83.4 / Opus 4.8 78.9 / Opus 4.7 78.9. Harness not stated for the Grok row (competitor harnesses unknown — comparability caution).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Real-world engineering excellence · figure: Benchmark bar charts; values in DOM chart accessibility text · quote_snippet: SWE Bench Pro resolve rate: Fable (max) 80.4%, Opus 4.8 (max) 69.2%, Grok 4.5 64.7%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id swebench-pro (registered batch 2). Competitors: Fable 80.4 / Opus 4.8 69.2 / Opus 4.7 64.3 / GLM 5.2 62.1 / GPT 5.5 58.6. Google's model pages label this benchmark 'SWE-Bench Pro (Public)' — same-family id, align variant at migration.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Faster than flash models · figure: Token efficiency chart; values restated in prose · quote_snippet: Grok 4.5 resolves tasks with 15,954 output tokens on average, about 4.2x fewer than Opus 4.8 (max) at 67,020
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "average across SWE Bench Pro tasks",
"judge": null
}Efficiency metric, not an accuracy score — unit is tokens (lower is better). Competitor value in prose: Opus 4.8 (max) 67,020 (4.2x more). Recorded because it is the page's headline cost/quality claim and precisely printed.