← 模型目录

Grok 3 / Grok 3 mini

xAI / Grok · 2025-02-19 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Grok 3

Grok 3 被 xAI 定位为开启『推理代理时代』的 Beta 模型,主打 (Think) 深度推理配置并曾以代号 chocolate 登顶 LMArena。评测覆盖推理、编码、多模态理解与长上下文 RAG:AIME'25 93.3%、GPQA 84.6%、LCB 79.4%、Chatbot Arena 1402 Elo、LOFT(128k)83.3%。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime-25 93.3% 模型 grok-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: With our highest level of test-time compute (cons@64), Grok 3 (Think) achieved 93.3% on this competition

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think mode at highest test-time compute",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 64,
  "aggregation": "cons@64 (consensus over 64 samples)",
  "judge": null
}

new-benchmark: aime-25 (registered batch 2, not yet in data/benchmarks/). AIME 2025 released Feb 12 2025, 7 days before the post — contamination-safe freshness noted by xAI itself. cons@64 aggregation makes this NOT comparable to single-attempt AIME rows (e.g. Gemini 2.5 Pro prose claim explicitly excludes majority voting).

打开官方来源

gpqa 84.6% 模型 grok-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: Grok 3 (Think) also attained 84.6% on graduate-level expert reasoning (GPQA)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think mode",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark gpqa. Page's chart caption reads 'Graduate-Level Google-Proof Q&A (Diamond)', i.e. the Diamond subset; shots/CoT condition not stated for this prose row.

打开官方来源

lcb 79.4% 模型 grok-3 · 版本 Code Generation 10/1/2024-2/1/2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: 79.4% on LiveCodeBench for code generation and problem-solving

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think mode",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark lcb. Page chart caption pins the window: 'Code Generation: 10/1/2024 - 2/1/2025' — different snapshot from Grok 4's Jan-May 2025 window and Google's v5 windows; compare only within identical windows.

打开官方来源

arena 1402 Elo 模型 grok-3 · 版本 未说明 · 指标 arena_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Next-Generation Intelligence from xAI · quote_snippet: achieving an Elo score of 1402 in the Chatbot Arena

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard",
  "judge": null
}

Maps to existing benchmark arena (Chatbot Arena / LMArena). Elo snapshot at launch; the page separately notes an early Grok 3 version tested under the codename 'chocolate' topped the leaderboard. Leaderboard Elo moves over time — treat as dated snapshot.

打开官方来源

aime24 52.2% 模型 grok-3 · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: AIME'24 · quote_snippet: AIME'24 | 52.2% | 39.7% | — | 39.2% | 9.3% | 16.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

DOM table row, machine-read. Columns: Grok 3 Beta 52.2 | Grok 3 mini Beta 39.7 | Gemini 2.0 — | DeepSeek-V3 39.2 | GPT 4o 9.3 | Claude 3.5 Sonnet 16.0. Competitor values stay in notes per batch-1 口径 (not split into comparison_cited rows; they are table columns, not independent prose citations). Section states results are 'among non reasoning models' with reasoning turned off.

打开官方来源

gpqa 75.4% 模型 grok-3 · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: GPQA · quote_snippet: GPQA | 75.4% | 66.2% | 64.7% | 59.1% | 53.6% | 65.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Grok 3 Beta 75.4 | mini 66.2 | Gemini 2.0 64.7 | DeepSeek-V3 59.1 | GPT 4o 53.6 | Claude 3.5 Sonnet 65.0. Chart caption on the page identifies GPQA as the Diamond subset. Separate evidence row from the 84.6 Think row — different model_variant.

打开官方来源

lcb 57.0% 模型 grok-3 · 版本 Code Generation 10/1/2024-2/1/2025; non-reasoning · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LCB · quote_snippet: LCB | 57.0% | 41.5% | 36.0% | 33.1% | 32.3% | 40.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Grok 3 Beta 57.0 | mini 41.5 | Gemini 2.0 36.0 | DeepSeek-V3 33.1 | GPT 4o 32.3 | Claude 3.5 Sonnet 40.2. Same 10/1/2024-2/1/2025 window as the Think row per page caption.

打开官方来源

mmlu-pro 79.9% 模型 grok-3 · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMLU-pro · quote_snippet: MMLU-pro | 79.9% | 78.9% | 79.1% | 75.9% | 72.6% | 78.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Grok 3 Beta 79.9 | mini 78.9 | Gemini 2.0 79.1 | DeepSeek-V3 75.9 | GPT 4o 72.6 | Claude 3.5 Sonnet 78.0. Maps to existing benchmark mmlu-pro. Shots not stated.

打开官方来源

loft 83.3% 模型 grok-3 · 版本 128k · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LOFT (128k) · quote_snippet: LOFT (128k) | 83.3% | 83.1% | 75.6% | — | 78.0% | 69.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "average across 12 diverse tasks (per page prose)",
  "judge": null
}

new-benchmark: loft not yet in data/benchmarks/ (LOFT long-context RAG benchmark). Grok 3 Beta 83.3 | mini 83.1 | Gemini 2.0 75.6 | DeepSeek-V3 — | GPT 4o 78.0 | Claude 3.5 Sonnet 69.9. Prose: 'state-of-the-art accuracy (averaged across 12 diverse tasks)' targeting long-context RAG use cases, in the 1M-token context window paragraph.

打开官方来源

simpleqa 43.6% 模型 grok-3 · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: SimpleQA · quote_snippet: SimpleQA | 43.6% | 21.7% | 44.3% | 24.9% | 38.2% | 28.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Grok 3 Beta 43.6 | mini 21.7 | Gemini 2.0 44.3 | DeepSeek-V3 24.9 | GPT 4o 38.2 | Claude 3.5 Sonnet 28.4 — page frames this as factual accuracy; Gemini 2.0 slightly ahead here.

打开官方来源

mmmu 73.2% 模型 grok-3 · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMMU · quote_snippet: MMMU | 73.2% | 69.4% | 72.7% | — | 69.1% | 70.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Grok 3 Beta 73.2 | mini 69.4 | Gemini 2.0 72.7 | DeepSeek-V3 — | GPT 4o 69.1 | Claude 3.5 Sonnet 70.4. Image understanding, MMMU (test) per chart caption on page.

打开官方来源

egoschema 74.5% 模型 grok-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: EgoSchema · quote_snippet: EgoSchema | 74.5% | 74.3% | 71.9% | — | 72.2% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: egoschema (introduced in this batch by google/gemini-2-0). Grok 3 Beta 74.5 | mini 74.3 | Gemini 2.0 71.9 | GPT 4o 72.2. Video understanding task.

打开官方来源

aime24 93.3% 模型 grok-3 · 版本 Think chart · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · table: AIME'24 (Think) bar chart (rendered DOM text) · row: Grok 3 Beta (Think) · quote_snippet: Grok 3 Beta (Think) 93.3 | Grok 3 mini Beta (Think) 95.8 | DeepSeek-R1 79.8 | Gemini 2.0 Flash Thinking 73.3 | o1 83.3 | o3 mini (high) 87.3 | o3 mini (medium) 79.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think (reasoning on); chart prints no per-model test-time compute",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Chart prints 93.3 for Grok 3 (Think) on BOTH the AIME'24 and AIME'25 charts; the prose cons@64 sentence covers only AIME'25. mini 95.8 has its own row. Per-bar test-time compute not printed - aggregation unknown for this cell.

打开官方来源

mmmu 78% 模型 grok-3 · 版本 Think chart · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · table: MMMU (Multimodal Understanding) bar chart (rendered DOM text) · row: Grok 3 Beta (Think) · quote_snippet: Grok 3 Beta (Think) 78 | Gemini 2.0 Flash Thinking 75.4 | o1 78.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think (reasoning on)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Only MMMU appearance for the Think config; the 73.2 non-reasoning MMMU table cell is a separate row. Values machine-read from the rendered DOM 2026-09-01.

打开官方来源

Grok 3 mini

Grok 3 mini 是同批发布的成本高效推理小模型(Beta,支持 Think 配置),官方称其在 AIME 2024 达 95.8%。评测亮点:AIME'25 90.8%、LCB 80.4%、MMLU-pro 78.9%。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime24 95.8% 模型 grok-3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: Grok 3 mini reaches a new frontier in cost-efficient reasoning... reaching 95.8% on AIME 2024

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think mode",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark aime24. Sentence sits in the Think-section paragraph; sampling/aggregation not stated (mini's 95.8 exceeds flagship Think's AIME'24 number only because flagship AIME'24 was reported non-reasoning — see table row).

打开官方来源

lcb 80.4% 模型 grok-3-mini · 版本 Code Generation 10/1/2024-2/1/2025 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: reaching 95.8% on AIME 2024 and 80.4% on LiveCodeBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think mode",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark lcb, same 10/1/2024-2/1/2025 window as the flagship Think row.

打开官方来源

aime-25 90.8% 模型 grok-3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · table: AIME'25 (Think) bar chart (rendered DOM text) · row: Grok 3 mini Beta (Think) · quote_snippet: Grok 3 Beta (Think) 93.3 | Grok 3 mini Beta (Think) 90.8 | DeepSeek-R1 70 | Gemini 2.0 Flash Thinking 53.5 | o1 (medium) 79 | o3 mini (high) 86.5 | o3 mini (medium) 76.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think (reasoning on); chart prints no per-model test-time compute",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aime-25 (registered batch 2, not yet in data/benchmarks/). Chart values machine-read from the rendered DOM 2026-09-01; competitor pairing follows the DOM label sequence. The prose cons@64 statement is tied to the Grok 3 (Think) bar only - per-bar test-time compute is not printed on the chart, so this row's aggregation is unknown; do not compare 1:1 with single-attempt AIME rows.

打开官方来源

gpqa 84% 模型 grok-3-mini · 版本 Think chart (Diamond subset per caption) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Thinking Harder: Test-time Compute and Reasoning · table: GPQA (Think) bar chart (rendered DOM text) · row: Grok 3 mini Beta (Think) · quote_snippet: Grok 3 Beta (Think) 84.6 | Grok 3 mini Beta (Think) 84 | DeepSeek-R1 71.5 | Gemini 2.0 Flash Thinking 74.2 | o1 78 | o3 mini (high) 79.7 | o3 mini (medium) 76.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Think (reasoning on)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Chart caption reads Graduate-Level Google-Proof Q&A (Diamond). Values machine-read from the rendered DOM 2026-09-01; separate evidence row from the 84.6 Grok 3 (Think) prose value - different model_variant.

打开官方来源

aime24 39.7% 模型 grok-3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: AIME’24 · quote_snippet: AIME’24 | 52.2% | 39.7% | — | 39.2% | 9.3% | 16.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--aime24-base - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源

gpqa 66.2% 模型 grok-3-mini · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: GPQA · quote_snippet: GPQA | 75.4% | 66.2% | 64.7% | 59.1% | 53.6% | 65.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--gpqa-base - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源

lcb 41.5% 模型 grok-3-mini · 版本 Code Generation 10/1/2024-2/1/2025; non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LCB · quote_snippet: LCB | 57.0% | 41.5% | 36.0% | 33.1% | 32.3% | 40.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--lcb-base - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源

mmlu-pro 78.9% 模型 grok-3-mini · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMLU-pro · quote_snippet: MMLU-pro | 79.9% | 78.9% | 79.1% | 75.9% | 72.6% | 78.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--mmlu-pro - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源

loft 83.1% 模型 grok-3-mini · 版本 128k · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LOFT (128k) · quote_snippet: LOFT (128k) | 83.3% | 83.1% | 75.6% | — | 78.0% | 69.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--loft - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源

simpleqa 21.7% 模型 grok-3-mini · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: SimpleQA · quote_snippet: SimpleQA | 43.6% | 21.7% | 44.3% | 24.9% | 38.2% | 28.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--simpleqa - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源

mmmu 69.4% 模型 grok-3-mini · 版本 non-reasoning · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMMU · quote_snippet: MMMU | 73.2% | 69.4% | 72.7% | — | 69.1% | 70.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--mmmu - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源

egoschema 74.3% 模型 grok-3-mini · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: EgoSchema · quote_snippet: EgoSchema | 74.5% | 74.3% | 71.9% | — | 72.2% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-reasoning (reasoning turned off)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Companion row to xai-grok-3--egoschema - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.

打开官方来源