Grok 4 / Grok 4 Heavy
xAI / Grok · 2025-07-09 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Grok 4
Grok 4 是 xAI 具备原生工具调用与实时搜索的模型,官方称其为 HLE text-only 子集首个突破 50%(50.7%)的模型。评测集中于数学、科学与编码:AIME'25 91.7%(无工具)、HMMT 2025 90%、GPQA 87.5%、ARC-AGI-2 15.9%(闭源 SOTA)。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Humanity's Last Exam / State of the art · figure: Interactive bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 w/ Python + Internet 38.6
{
"harness": null,
"tools": [
"Python",
"Internet"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-free Grok 4 bar on the same chart reads 25.4 (chart DOM) — separate condition, kept here in notes rather than a third row; compare only within identical tool conditions.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Intelligence · figure: Also rendered as bar chart; the DOM contains the numeric values · quote_snippet: state-of-the-art for closed models on ARC-AGI V2 with 15.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark arc-agi, variant 2. Prose and chart DOM agree (15.9). Claim scope is 'state-of-the-art for closed models' — Gemini 3 Deep Think's 45.1% with code execution (ARC Prize Verified) is a different tool+verification condition, not directly comparable.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Intelligence · quote_snippet: it dominates with $4694.15 net worth and 4569 units sold (averages across 5 runs)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "average across 5 runs",
"judge": null
}new-benchmark: vending-bench not yet in data/benchmarks/ — the ORIGINAL agentic Vending-Bench, distinct from vending-bench-2 (registered by google/gemini-3); scores of the two generations are not comparable. Secondary metric in prose: 4569 units sold. Human baseline also printed in prose: humans $844.05, 344 units (baseline, kept in notes — not a competitor model row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: USAMO 2025 · figure: Olympiad Math Proofs bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 37.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-free Grok 4 bar of the same chart (Heavy w/ Python covered by its own row above).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: GPQA · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 87.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-free Grok 4 bar. Compare with Grok 3's 75.4 (non-reasoning) / 84.6 (Think) only with condition alignment.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: LiveCodeBench (Jan - May) · figure: Competitive Coding bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 79
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-free Grok 4 bar; Grok 4 w/ Python 79.3 sits between this and the Heavy bar (chart DOM), kept in notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: HMMT 2025 · figure: Competitive Math bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 90
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-free Grok 4 bar; Grok 4 w/ Python 93.9 kept in notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: AIME'25 · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 91.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-free Grok 4 bar; Grok 4 w/ Python 98.8 kept in notes. Aggregation not stated on this chart — do not assume single-attempt; Grok 3's AIME'25 row was cons@64.
Grok 4 Heavy
Grok 4 Heavy 是并行测试时计算版本(SuperGrok Heavy 档),在全部数学与科学项目上领先同代 Grok 4。评测亮点:HLE 全集 44.4%(+Python+Internet)、USAMO 2025 61.9%、AIME'25 100(+Python)、HMMT 2025 96.7%。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Intelligence · quote_snippet: is the first to score 50.7% on Humanity's Last Exam (text-only subset)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark hlehle, text-only subset variant. The Grok 4 Heavy section separately claims the model 'is the first model to score 50% on Humanity's Last Exam' — a rounded restatement of this same text-only result, not a separate full-set row. NOT comparable to tool-augmented HLE rows (see the full-set row below) nor to tool-free rows from other vendors without variant alignment.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Humanity's Last Exam / State of the art · figure: Interactive bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python + Internet 44.4
{
"harness": null,
"tools": [
"Python",
"Internet"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Chart section explicitly labeled 'Full set (April 3, 2025) with Python and Internet tools'. Chart DOM values: Grok 4 Heavy w/ P+I 44.4 | Grok 4 w/ P+I 38.6 | Gemini Deep Research 26.9 | Grok 4 25.4 | o3 w/ P+I 24.9 | Gemini 2.5 Pro 21.6 | o3 21. Same extraction path as grok-4-6 batch-1 (values present in DOM text; not image-OCR). Competitor values kept in notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Intelligence · figure: Also rendered as bar chart; the DOM contains the numeric values · quote_snippet: Grok 4 Heavy leads USAMO'25 with 61.9%
{
"harness": null,
"tools": [
"Python"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: usamo-2025 (id pattern aligned with batch-3 usamo-2026 and aime-25). Chart DOM corroboration: USAMO 2025 chart bars Grok 4 Heavy w/ Python 61.9 | Gemini Deep Think 49.4 | Grok 4 37.5 | Gemini 2.5 Pro 34.5 | o3 21.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: GPQA · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 88.4
{
"harness": null,
"tools": [
"Python"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}GPQA chart DOM: Grok 4 Heavy w/ Python 88.4 | Grok 4 87.5 | Gemini 2.5 Pro 86.4 | o3 83.3 | Claude Opus 4 79.6. The GPQA section exists on the live page but was missing from the web-reader markdown — values captured from rendered DOM. Shots/CoT not stated.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: LiveCodeBench (Jan - May) · figure: Competitive Coding bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 79.4
{
"harness": null,
"tools": [
"Python"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}LCB (Jan-May) chart DOM: Heavy w/ Python 79.4 | Grok 4 w/ Python 79.3 | Grok 4 79 | Gemini 2.5 Pro 74.2 | o3 72. Window differs from every other LCB row in the library — variant-tagged.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: HMMT 2025 · figure: Competitive Math bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 96.7
{
"harness": null,
"tools": [
"Python"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id hmmt-25 registered in batch 3 (CN batch) — no new id minted, id stays hmmt-25 for cross-batch consistency. Chart DOM: Heavy w/ Python 96.7 | Grok 4 w/ Python 93.9 | Grok 4 90 | Gemini 2.5 Pro 82.5 | o3 77.5 | Claude Opus 4 58.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: AIME'25 · figure: Bar chart; the rendered SVG DOM contains the numeric values · quote_snippet: Grok 4 Heavy w/ Python 100
{
"harness": null,
"tools": [
"Python"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-25 (registered batch 2). Chart DOM: Heavy w/ Python 100 | Grok 4 w/ Python 98.8 | o3 w/ Python 98.4 | Grok 4 91.7 | o3 88.9 | Gemini 2.5 Pro 88 | Claude Opus 4 75.5. Saturation territory — 1-point differences are noise without variance reporting. The AIME'25 section was missing from the web-reader markdown; captured from rendered DOM.
未关联到本页变体的记录
这些记录不会分配给任意模型参与选型。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Intelligence · quote_snippet: nearly double Opus's ~8.6%, +8pp over previous high
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor score independently printed in xAI prose (batch-1 口径: separate comparison_cited row). Value is approximate (~) as printed; chart DOM shows 8.6 for Claude Opus 4. model_id references a competitor model not in this release's models list (first-batch convention, flagged in batch-2 README).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Frontier Intelligence · quote_snippet: vastly outpacing Claude Opus 4 ($2077.41, 1412 units)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor net worth independently printed in prose; 1412 units secondary metric. Whether Anthropic's own runs used the same 5-run averaging is not stated by xAI — comparability caution.