← 模型目录

Grok 4 / Grok 4 Heavy

xAI / Grok · 2025-07-09 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Grok 4

Grok 4 是 xAI 具备原生工具调用与实时搜索的模型,官方称其为 HLE text-only 子集首个突破 50%(50.7%)的模型。评测集中于数学、科学与编码:AIME'25 91.7%(无工具)、HMMT 2025 90%、GPQA 87.5%、ARC-AGI-2 15.9%(闭源 SOTA)。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 38.6% 模型 grok-4 · 版本 full set (April 3, 2025), with Python + Internet tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Humanity's Last Exam / State of the art · figure: Interactive bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 w/ Python + Internet 38.6

{
  "harness": null,
  "tools": [
    "Python",
    "Internet"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free Grok 4 bar on the same chart reads 25.4 (chart DOM) — separate condition, kept here in notes rather than a third row; compare only within identical tool conditions.

打开官方来源

arc-agi 15.9% 模型 grok-4 · 版本 2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier Intelligence · figure: Also rendered as bar chart; the DOM contains the numeric values · quote_snippet: state-of-the-art for closed models on ARC-AGI V2 with 15.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark arc-agi, variant 2. Prose and chart DOM agree (15.9). Claim scope is 'state-of-the-art for closed models' — Gemini 3 Deep Think's 45.1% with code execution (ARC Prize Verified) is a different tool+verification condition, not directly comparable.

打开官方来源

vending-bench $4694.15 net worth 模型 grok-4 · 版本 未说明 · 指标 net_worth · 单位 usd 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier Intelligence · quote_snippet: it dominates with $4694.15 net worth and 4569 units sold (averages across 5 runs)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "average across 5 runs",
  "judge": null
}

new-benchmark: vending-bench not yet in data/benchmarks/ — the ORIGINAL agentic Vending-Bench, distinct from vending-bench-2 (registered by google/gemini-3); scores of the two generations are not comparable. Secondary metric in prose: 4569 units sold. Human baseline also printed in prose: humans $844.05, 344 units (baseline, kept in notes — not a competitor model row).

打开官方来源

usamo-2025 37.5% 模型 grok-4 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: USAMO 2025 · figure: Olympiad Math Proofs bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 37.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free Grok 4 bar of the same chart (Heavy w/ Python covered by its own row above).

打开官方来源

gpqa 87.5% 模型 grok-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: GPQA · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 87.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free Grok 4 bar. Compare with Grok 3's 75.4 (non-reasoning) / 84.6 (Think) only with condition alignment.

打开官方来源

lcb 79% 模型 grok-4 · 版本 Jan-May 2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: LiveCodeBench (Jan - May) · figure: Competitive Coding bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 79

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free Grok 4 bar; Grok 4 w/ Python 79.3 sits between this and the Heavy bar (chart DOM), kept in notes.

打开官方来源

hmmt25 90% 模型 grok-4 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: HMMT 2025 · figure: Competitive Math bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 90

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free Grok 4 bar; Grok 4 w/ Python 93.9 kept in notes.

打开官方来源

aime-25 91.7% 模型 grok-4 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: AIME'25 · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 91.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free Grok 4 bar; Grok 4 w/ Python 98.8 kept in notes. Aggregation not stated on this chart — do not assume single-attempt; Grok 3's AIME'25 row was cons@64.

打开官方来源

Grok 4 Heavy

Grok 4 Heavy 是并行测试时计算版本(SuperGrok Heavy 档),在全部数学与科学项目上领先同代 Grok 4。评测亮点:HLE 全集 44.4%(+Python+Internet)、USAMO 2025 61.9%、AIME'25 100(+Python)、HMMT 2025 96.7%。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 50.7% 模型 grok-4-heavy · 版本 text-only subset · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier Intelligence · quote_snippet: is the first to score 50.7% on Humanity's Last Exam (text-only subset)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark hlehle, text-only subset variant. The Grok 4 Heavy section separately claims the model 'is the first model to score 50% on Humanity's Last Exam' — a rounded restatement of this same text-only result, not a separate full-set row. NOT comparable to tool-augmented HLE rows (see the full-set row below) nor to tool-free rows from other vendors without variant alignment.

打开官方来源

hlehle 44.4% 模型 grok-4-heavy · 版本 full set (April 3, 2025), with Python + Internet tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Humanity's Last Exam / State of the art · figure: Interactive bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python + Internet 44.4

{
  "harness": null,
  "tools": [
    "Python",
    "Internet"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Chart section explicitly labeled 'Full set (April 3, 2025) with Python and Internet tools'. Chart DOM values: Grok 4 Heavy w/ P+I 44.4 | Grok 4 w/ P+I 38.6 | Gemini Deep Research 26.9 | Grok 4 25.4 | o3 w/ P+I 24.9 | Gemini 2.5 Pro 21.6 | o3 21. Same extraction path as grok-4-6 batch-1 (values present in DOM text; not image-OCR). Competitor values kept in notes.

打开官方来源

usamo-2025 61.9% 模型 grok-4-heavy · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier Intelligence · figure: Also rendered as bar chart; the DOM contains the numeric values · quote_snippet: Grok 4 Heavy leads USAMO'25 with 61.9%

{
  "harness": null,
  "tools": [
    "Python"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: usamo-2025 (id pattern aligned with batch-3 usamo-2026 and aime-25). Chart DOM corroboration: USAMO 2025 chart bars Grok 4 Heavy w/ Python 61.9 | Gemini Deep Think 49.4 | Grok 4 37.5 | Gemini 2.5 Pro 34.5 | o3 21.7.

打开官方来源

gpqa 88.4% 模型 grok-4-heavy · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: GPQA · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 88.4

{
  "harness": null,
  "tools": [
    "Python"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

GPQA chart DOM: Grok 4 Heavy w/ Python 88.4 | Grok 4 87.5 | Gemini 2.5 Pro 86.4 | o3 83.3 | Claude Opus 4 79.6. The GPQA section exists on the live page but was missing from the web-reader markdown — values captured from rendered DOM. Shots/CoT not stated.

打开官方来源

lcb 79.4% 模型 grok-4-heavy · 版本 Jan-May 2025; with Python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: LiveCodeBench (Jan - May) · figure: Competitive Coding bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 79.4

{
  "harness": null,
  "tools": [
    "Python"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

LCB (Jan-May) chart DOM: Heavy w/ Python 79.4 | Grok 4 w/ Python 79.3 | Grok 4 79 | Gemini 2.5 Pro 74.2 | o3 72. Window differs from every other LCB row in the library — variant-tagged.

打开官方来源

hmmt25 96.7% 模型 grok-4-heavy · 版本 with Python · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: HMMT 2025 · figure: Competitive Math bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 96.7

{
  "harness": null,
  "tools": [
    "Python"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id hmmt-25 registered in batch 3 (CN batch) — no new id minted, id stays hmmt-25 for cross-batch consistency. Chart DOM: Heavy w/ Python 96.7 | Grok 4 w/ Python 93.9 | Grok 4 90 | Gemini 2.5 Pro 82.5 | o3 77.5 | Claude Opus 4 58.3.

打开官方来源

aime-25 100 模型 grok-4-heavy · 版本 with Python · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: AIME'25 · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 100

{
  "harness": null,
  "tools": [
    "Python"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aime-25 (registered batch 2). Chart DOM: Heavy w/ Python 100 | Grok 4 w/ Python 98.8 | o3 w/ Python 98.4 | Grok 4 91.7 | o3 88.9 | Gemini 2.5 Pro 88 | Claude Opus 4 75.5. Saturation territory — 1-point differences are noise without variance reporting. The AIME'25 section was missing from the web-reader markdown; captured from rendered DOM.

打开官方来源

未关联到本页变体的记录

这些记录不会分配给任意模型参与选型。

arc-agi ~8.6% 模型 claude-opus-4 · 版本 2 · 指标 未说明 · 单位 percent 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier Intelligence · quote_snippet: nearly double Opus's ~8.6%, +8pp over previous high

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor score independently printed in xAI prose (batch-1 口径: separate comparison_cited row). Value is approximate (~) as printed; chart DOM shows 8.6 for Claude Opus 4. model_id references a competitor model not in this release's models list (first-batch convention, flagged in batch-2 README).

打开官方来源

vending-bench $2077.41 net worth 模型 claude-opus-4 · 版本 未说明 · 指标 net_worth · 单位 usd 来源等级 A · comparison_cited · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Frontier Intelligence · quote_snippet: vastly outpacing Claude Opus 4 ($2077.41, 1412 units)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor net worth independently printed in prose; 1412 units secondary metric. Whether Anthropic's own runs used the same 5-run averaging is not stated by xAI — comparability caution.

打开官方来源