Grok 3 / Grok 3 mini
xAI / Grok · 2025-02-19 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Grok 3
Grok 3 被 xAI 定位为开启『推理代理时代』的 Beta 模型,主打 (Think) 深度推理配置并曾以代号 chocolate 登顶 LMArena。评测覆盖推理、编码、多模态理解与长上下文 RAG:AIME'25 93.3%、GPQA 84.6%、LCB 79.4%、Chatbot Arena 1402 Elo、LOFT(128k)83.3%。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: With our highest level of test-time compute (cons@64), Grok 3 (Think) achieved 93.3% on this competition
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think mode at highest test-time compute",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 64,
"aggregation": "cons@64 (consensus over 64 samples)",
"judge": null
}new-benchmark: aime-25 (registered batch 2, not yet in data/benchmarks/). AIME 2025 released Feb 12 2025, 7 days before the post — contamination-safe freshness noted by xAI itself. cons@64 aggregation makes this NOT comparable to single-attempt AIME rows (e.g. Gemini 2.5 Pro prose claim explicitly excludes majority voting).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: Grok 3 (Think) also attained 84.6% on graduate-level expert reasoning (GPQA)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think mode",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark gpqa. Page's chart caption reads 'Graduate-Level Google-Proof Q&A (Diamond)', i.e. the Diamond subset; shots/CoT condition not stated for this prose row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: 79.4% on LiveCodeBench for code generation and problem-solving
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think mode",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark lcb. Page chart caption pins the window: 'Code Generation: 10/1/2024 - 2/1/2025' — different snapshot from Grok 4's Jan-May 2025 window and Google's v5 windows; compare only within identical windows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Next-Generation Intelligence from xAI · quote_snippet: achieving an Elo score of 1402 in the Chatbot Arena
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Maps to existing benchmark arena (Chatbot Arena / LMArena). Elo snapshot at launch; the page separately notes an early Grok 3 version tested under the codename 'chocolate' topped the leaderboard. Leaderboard Elo moves over time — treat as dated snapshot.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: AIME'24 · quote_snippet: AIME'24 | 52.2% | 39.7% | — | 39.2% | 9.3% | 16.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DOM table row, machine-read. Columns: Grok 3 Beta 52.2 | Grok 3 mini Beta 39.7 | Gemini 2.0 — | DeepSeek-V3 39.2 | GPT 4o 9.3 | Claude 3.5 Sonnet 16.0. Competitor values stay in notes per batch-1 口径 (not split into comparison_cited rows; they are table columns, not independent prose citations). Section states results are 'among non reasoning models' with reasoning turned off.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: GPQA · quote_snippet: GPQA | 75.4% | 66.2% | 64.7% | 59.1% | 53.6% | 65.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Grok 3 Beta 75.4 | mini 66.2 | Gemini 2.0 64.7 | DeepSeek-V3 59.1 | GPT 4o 53.6 | Claude 3.5 Sonnet 65.0. Chart caption on the page identifies GPQA as the Diamond subset. Separate evidence row from the 84.6 Think row — different model_variant.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LCB · quote_snippet: LCB | 57.0% | 41.5% | 36.0% | 33.1% | 32.3% | 40.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Grok 3 Beta 57.0 | mini 41.5 | Gemini 2.0 36.0 | DeepSeek-V3 33.1 | GPT 4o 32.3 | Claude 3.5 Sonnet 40.2. Same 10/1/2024-2/1/2025 window as the Think row per page caption.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMLU-pro · quote_snippet: MMLU-pro | 79.9% | 78.9% | 79.1% | 75.9% | 72.6% | 78.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Grok 3 Beta 79.9 | mini 78.9 | Gemini 2.0 79.1 | DeepSeek-V3 75.9 | GPT 4o 72.6 | Claude 3.5 Sonnet 78.0. Maps to existing benchmark mmlu-pro. Shots not stated.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LOFT (128k) · quote_snippet: LOFT (128k) | 83.3% | 83.1% | 75.6% | — | 78.0% | 69.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "average across 12 diverse tasks (per page prose)",
"judge": null
}new-benchmark: loft not yet in data/benchmarks/ (LOFT long-context RAG benchmark). Grok 3 Beta 83.3 | mini 83.1 | Gemini 2.0 75.6 | DeepSeek-V3 — | GPT 4o 78.0 | Claude 3.5 Sonnet 69.9. Prose: 'state-of-the-art accuracy (averaged across 12 diverse tasks)' targeting long-context RAG use cases, in the 1M-token context window paragraph.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: SimpleQA · quote_snippet: SimpleQA | 43.6% | 21.7% | 44.3% | 24.9% | 38.2% | 28.4%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Grok 3 Beta 43.6 | mini 21.7 | Gemini 2.0 44.3 | DeepSeek-V3 24.9 | GPT 4o 38.2 | Claude 3.5 Sonnet 28.4 — page frames this as factual accuracy; Gemini 2.0 slightly ahead here.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMMU · quote_snippet: MMMU | 73.2% | 69.4% | 72.7% | — | 69.1% | 70.4%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Grok 3 Beta 73.2 | mini 69.4 | Gemini 2.0 72.7 | DeepSeek-V3 — | GPT 4o 69.1 | Claude 3.5 Sonnet 70.4. Image understanding, MMMU (test) per chart caption on page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: EgoSchema · quote_snippet: EgoSchema | 74.5% | 74.3% | 71.9% | — | 72.2% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: egoschema (introduced in this batch by google/gemini-2-0). Grok 3 Beta 74.5 | mini 74.3 | Gemini 2.0 71.9 | GPT 4o 72.2. Video understanding task.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · table: AIME'24 (Think) bar chart (rendered DOM text) · row: Grok 3 Beta (Think) · quote_snippet: Grok 3 Beta (Think) 93.3 | Grok 3 mini Beta (Think) 95.8 | DeepSeek-R1 79.8 | Gemini 2.0 Flash Thinking 73.3 | o1 83.3 | o3 mini (high) 87.3 | o3 mini (medium) 79.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think (reasoning on); chart prints no per-model test-time compute",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Chart prints 93.3 for Grok 3 (Think) on BOTH the AIME'24 and AIME'25 charts; the prose cons@64 sentence covers only AIME'25. mini 95.8 has its own row. Per-bar test-time compute not printed - aggregation unknown for this cell.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · table: MMMU (Multimodal Understanding) bar chart (rendered DOM text) · row: Grok 3 Beta (Think) · quote_snippet: Grok 3 Beta (Think) 78 | Gemini 2.0 Flash Thinking 75.4 | o1 78.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think (reasoning on)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Only MMMU appearance for the Think config; the 73.2 non-reasoning MMMU table cell is a separate row. Values machine-read from the rendered DOM 2026-09-01.
Grok 3 mini
Grok 3 mini 是同批发布的成本高效推理小模型(Beta,支持 Think 配置),官方称其在 AIME 2024 达 95.8%。评测亮点:AIME'25 90.8%、LCB 80.4%、MMLU-pro 78.9%。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: Grok 3 mini reaches a new frontier in cost-efficient reasoning... reaching 95.8% on AIME 2024
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think mode",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark aime24. Sentence sits in the Think-section paragraph; sampling/aggregation not stated (mini's 95.8 exceeds flagship Think's AIME'24 number only because flagship AIME'24 was reported non-reasoning — see table row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · quote_snippet: reaching 95.8% on AIME 2024 and 80.4% on LiveCodeBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think mode",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark lcb, same 10/1/2024-2/1/2025 window as the flagship Think row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · table: AIME'25 (Think) bar chart (rendered DOM text) · row: Grok 3 mini Beta (Think) · quote_snippet: Grok 3 Beta (Think) 93.3 | Grok 3 mini Beta (Think) 90.8 | DeepSeek-R1 70 | Gemini 2.0 Flash Thinking 53.5 | o1 (medium) 79 | o3 mini (high) 86.5 | o3 mini (medium) 76.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think (reasoning on); chart prints no per-model test-time compute",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-25 (registered batch 2, not yet in data/benchmarks/). Chart values machine-read from the rendered DOM 2026-09-01; competitor pairing follows the DOM label sequence. The prose cons@64 statement is tied to the Grok 3 (Think) bar only - per-bar test-time compute is not printed on the chart, so this row's aggregation is unknown; do not compare 1:1 with single-attempt AIME rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Thinking Harder: Test-time Compute and Reasoning · table: GPQA (Think) bar chart (rendered DOM text) · row: Grok 3 mini Beta (Think) · quote_snippet: Grok 3 Beta (Think) 84.6 | Grok 3 mini Beta (Think) 84 | DeepSeek-R1 71.5 | Gemini 2.0 Flash Thinking 74.2 | o1 78 | o3 mini (high) 79.7 | o3 mini (medium) 76.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Think (reasoning on)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Chart caption reads Graduate-Level Google-Proof Q&A (Diamond). Values machine-read from the rendered DOM 2026-09-01; separate evidence row from the 84.6 Grok 3 (Think) prose value - different model_variant.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: AIME’24 · quote_snippet: AIME’24 | 52.2% | 39.7% | — | 39.2% | 9.3% | 16.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--aime24-base - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: GPQA · quote_snippet: GPQA | 75.4% | 66.2% | 64.7% | 59.1% | 53.6% | 65.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--gpqa-base - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LCB · quote_snippet: LCB | 57.0% | 41.5% | 36.0% | 33.1% | 32.3% | 40.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--lcb-base - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMLU-pro · quote_snippet: MMLU-pro | 79.9% | 78.9% | 79.1% | 75.9% | 72.6% | 78.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--mmlu-pro - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: LOFT (128k) · quote_snippet: LOFT (128k) | 83.3% | 83.1% | 75.6% | — | 78.0% | 69.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--loft - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: SimpleQA · quote_snippet: SimpleQA | 43.6% | 21.7% | 44.3% | 24.9% | 38.2% | 28.4%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--simpleqa - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: MMMU · quote_snippet: MMMU | 73.2% | 69.4% | 72.7% | — | 69.1% | 70.4%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--mmmu - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Pretraining on a Massive Scale · table: Academic benchmarks among non reasoning models (DOM table) · row: EgoSchema · quote_snippet: EgoSchema | 74.5% | 74.3% | 71.9% | — | 72.2% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "non-reasoning (reasoning turned off)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Companion row to xai-grok-3--egoschema - same DOM table, Grok 3 mini Beta column promoted to its own evidence row in the 2026-09-01 audit. Competitor columns remain in the sibling row notes.