← 模型目录

Grok 4.6

xAI / Grok · 2026-08-12 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Grok 4.6

发布文称 Grok 4.6 在 Grok 4.5 基础上聚焦长时间运行的智能体与更具野心的交互和视觉任务。10 项评测覆盖智能体编码、知识与法律工作,亮点为 Artificial Analysis Intelligence Index 61(自述追平 GPT-5.6 Sol)与 GDPval-AA v2 1753 Elo。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 2 / 输出 6 · 另有 fast 变体,价格为标准价的 2 倍

本变体的评测证据

aa-intelligence-index 61 模型 grok-4-6 · 版本 未说明 · 指标 composite_intelligence_index · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals / Grok 4.6 High · table: AA Intelligence Index row · row: AA Intelligence Index · figure: Also rendered as bar chart; the DOM contains the numeric values · quote_snippet: AA Intelligence Index: 61 | 56 | 61 | 62

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "composite of nine benchmarks",
  "judge": null
}

new-benchmark: aa-intelligence-index not yet in data/benchmarks.json. Third-party composite index. Comparison columns same row: Grok 4.5 High 56, GPT-5.6 Sol Max 61, Fable 5 Max 62. Page states competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards; 'Best score per evaluation in bold. Third-party model scores are the best of self-reported or publicly available results.'

打开官方来源

gdpval-aa 1753 模型 grok-4-6 · 版本 v2 · 指标 未说明 · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: GDPVal-AA v2 row · row: GDPVal-AA v2 · quote_snippet: GDPVal-AA v2: 1753 | 1526 | 1728 | 1741

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: gdpval-aa not yet in data/benchmarks.json. AA-administered benchmark. Comparison columns same row: Grok 4.5 High 1526, GPT-5.6 Sol Max 1728, Fable 5 Max 1741.

打开官方来源

cursor-bench 69.9% 模型 grok-4-6 · 版本 v3.2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: CursorBench v3.2 row · row: CursorBench v3.2 · quote_snippet: CursorBench v3.2: 69.9% | 66.7% | 67.2% | 70.5%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cursor-bench not yet in data/benchmarks.json (same family as Anthropic's CursorBench 3.2 row; version notation v3.2 vs 3.2 must be unified before merging). Comparison columns same row: Grok 4.5 High 66.7%, GPT-5.6 Sol Max 67.2%, Fable 5 Max 70.5%.

打开官方来源

deepswe 65.9% 模型 grok-4-6 · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: DeepSWE v1.1 row · row: DeepSWE v1.1 · quote_snippet: DeepSWE v1.1: 65.9% | 54% | 73% | 70%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepswe not yet in data/benchmarks/. Comparison columns same row: Grok 4.5 High 54%, GPT-5.6 Sol Max 73%, Fable 5 Max 70%. xAI does not disclose which harness its own 65.9% used; GLM and Kimi report 66.9 and 67.5 under mini-swe-agent / Kimi Code respectively - protocols differ across these rows.

打开官方来源

frontier-code 61.3% 模型 grok-4-6 · 版本 v1.1 (Extended) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: FrontierCode v1.1 (Extended) row · row: FrontierCode v1.1 (Extended) · quote_snippet: FrontierCode v1.1 (Extended): 61.3% | 56.6% | 60.6% | 63.6%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: frontier-code not yet in data/benchmarks.json (distinct from frontier-bench and frontierswe). Comparison columns same row: Grok 4.5 High 56.6%, GPT-5.6 Sol Max 60.6%, Fable 5 Max 63.6%.

打开官方来源

apex-agents 57.5% 模型 grok-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: APEX-Agents row · row: APEX-Agents · quote_snippet: APEX-Agents: 57.5% | 47.1% | 56.7% | 59.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: apex-agents not yet in data/benchmarks.json. Kimi's page cites Artificial Analysis as origin for its APEX-Agents row. Comparison columns same row: Grok 4.5 High 47.1%, GPT-5.6 Sol Max 56.7%, Fable 5 Max 59.2%.

打开官方来源

terminalbench 26% 模型 grok-4-6 · 版本 3.0 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: Terminal-Bench v3.0 row · row: Terminal-Bench v3.0 · quote_snippet: Terminal-Bench v3.0: 26% | 15.7% | 34.6% | 34.1%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Comparison columns same row: Grok 4.5 High 15.7%, GPT-5.6 Sol Max 34.6%, Fable 5 Max 34.1%. GPT-5.6 Sol 34.6% matches GLM-5.3's Terminal Bench 3.0 value for Sol; run protocols (avg@3, rollouts) are not disclosed by xAI, so cross-vendor comparability is unconfirmed.

打开官方来源

apex-swe 56.4% 模型 grok-4-6 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: APEX-SWE row · row: APEX-SWE · quote_snippet: APEX-SWE: 56.4% | 53.6% | — | 58.8%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: apex-swe not yet in data/benchmarks.json. Comparison columns same row: Grok 4.5 High 53.6%, GPT-5.6 Sol Max not reported (em dash in source), Fable 5 Max 58.8%.

打开官方来源

aa-briefcase 1577 模型 grok-4-6 · 版本 未说明 · 指标 未说明 · 单位 points 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: AA-Briefcase row · row: AA-Briefcase · quote_snippet: AA-Briefcase: 1577 | 1313 | 1502 | 1574

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aa-briefcase not yet in data/benchmarks.json. Comparison columns same row: Grok 4.5 High 1313, GPT-5.6 Sol Max 1502, Fable 5 Max 1574.

打开官方来源

harvey-lab 15.8% 模型 grok-4-6 · 版本 Vals harness · 指标 未说明 · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evals · table: Harvey LAB (Vals) row · row: Harvey LAB (Vals) · quote_snippet: Harvey LAB (Vals): 15.8% | 12.9% | 2.5% | 11.3%

{
  "harness": "Vals",
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: harvey-lab (Harvey legal AI benchmark, run by Vals AI) not yet in data/benchmarks.json. Legal-domain vertical. Comparison columns same row: Grok 4.5 High 12.9%, GPT-5.6 Sol Max 2.5%, Fable 5 Max 11.3%.

打开官方来源