Grok 4.6
xAI / Grok · 2026-08-12 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Grok 4.6
发布文称 Grok 4.6 在 Grok 4.5 基础上聚焦长时间运行的智能体与更具野心的交互和视觉任务。10 项评测覆盖智能体编码、知识与法律工作,亮点为 Artificial Analysis Intelligence Index 61(自述追平 GPT-5.6 Sol)与 GDPval-AA v2 1753 Elo。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 2 / 输出 6 · 另有 fast 变体,价格为标准价的 2 倍
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals / Grok 4.6 High · table: AA Intelligence Index row · row: AA Intelligence Index · figure: Also rendered as bar chart; the DOM contains the numeric values · quote_snippet: AA Intelligence Index: 61 | 56 | 61 | 62
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "composite of nine benchmarks",
"judge": null
}new-benchmark: aa-intelligence-index not yet in data/benchmarks.json. Third-party composite index. Comparison columns same row: Grok 4.5 High 56, GPT-5.6 Sol Max 61, Fable 5 Max 62. Page states competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards; 'Best score per evaluation in bold. Third-party model scores are the best of self-reported or publicly available results.'
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: GDPVal-AA v2 row · row: GDPVal-AA v2 · quote_snippet: GDPVal-AA v2: 1753 | 1526 | 1728 | 1741
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gdpval-aa not yet in data/benchmarks.json. AA-administered benchmark. Comparison columns same row: Grok 4.5 High 1526, GPT-5.6 Sol Max 1728, Fable 5 Max 1741.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: CursorBench v3.2 row · row: CursorBench v3.2 · quote_snippet: CursorBench v3.2: 69.9% | 66.7% | 67.2% | 70.5%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cursor-bench not yet in data/benchmarks.json (same family as Anthropic's CursorBench 3.2 row; version notation v3.2 vs 3.2 must be unified before merging). Comparison columns same row: Grok 4.5 High 66.7%, GPT-5.6 Sol Max 67.2%, Fable 5 Max 70.5%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: DeepSWE v1.1 row · row: DeepSWE v1.1 · quote_snippet: DeepSWE v1.1: 65.9% | 54% | 73% | 70%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepswe not yet in data/benchmarks/. Comparison columns same row: Grok 4.5 High 54%, GPT-5.6 Sol Max 73%, Fable 5 Max 70%. xAI does not disclose which harness its own 65.9% used; GLM and Kimi report 66.9 and 67.5 under mini-swe-agent / Kimi Code respectively - protocols differ across these rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: FrontierCode v1.1 (Extended) row · row: FrontierCode v1.1 (Extended) · quote_snippet: FrontierCode v1.1 (Extended): 61.3% | 56.6% | 60.6% | 63.6%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: frontier-code not yet in data/benchmarks.json (distinct from frontier-bench and frontierswe). Comparison columns same row: Grok 4.5 High 56.6%, GPT-5.6 Sol Max 60.6%, Fable 5 Max 63.6%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: APEX-Agents row · row: APEX-Agents · quote_snippet: APEX-Agents: 57.5% | 47.1% | 56.7% | 59.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: apex-agents not yet in data/benchmarks.json. Kimi's page cites Artificial Analysis as origin for its APEX-Agents row. Comparison columns same row: Grok 4.5 High 47.1%, GPT-5.6 Sol Max 56.7%, Fable 5 Max 59.2%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: Terminal-Bench v3.0 row · row: Terminal-Bench v3.0 · quote_snippet: Terminal-Bench v3.0: 26% | 15.7% | 34.6% | 34.1%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Comparison columns same row: Grok 4.5 High 15.7%, GPT-5.6 Sol Max 34.6%, Fable 5 Max 34.1%. GPT-5.6 Sol 34.6% matches GLM-5.3's Terminal Bench 3.0 value for Sol; run protocols (avg@3, rollouts) are not disclosed by xAI, so cross-vendor comparability is unconfirmed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: APEX-SWE row · row: APEX-SWE · quote_snippet: APEX-SWE: 56.4% | 53.6% | — | 58.8%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: apex-swe not yet in data/benchmarks.json. Comparison columns same row: Grok 4.5 High 53.6%, GPT-5.6 Sol Max not reported (em dash in source), Fable 5 Max 58.8%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: AA-Briefcase row · row: AA-Briefcase · quote_snippet: AA-Briefcase: 1577 | 1313 | 1502 | 1574
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aa-briefcase not yet in data/benchmarks.json. Comparison columns same row: Grok 4.5 High 1313, GPT-5.6 Sol Max 1502, Fable 5 Max 1574.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evals · table: Harvey LAB (Vals) row · row: Harvey LAB (Vals) · quote_snippet: Harvey LAB (Vals): 15.8% | 12.9% | 2.5% | 11.3%
{
"harness": "Vals",
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: harvey-lab (Harvey legal AI benchmark, run by Vals AI) not yet in data/benchmarks.json. Legal-domain vertical. Comparison columns same row: Grok 4.5 High 12.9%, GPT-5.6 Sol Max 2.5%, Fable 5 Max 11.3%.