Gemini 3 Pro / Gemini 3 Deep Think
Google DeepMind / Gemini · 2025-11-18 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 3 Pro
Gemini 3 发布以「智能新纪元」为题呈现,Gemini 3 Pro 为旗舰型号(发布时 preview)。已收录评测横跨推理、代码、多模态与 Agent:亮点 LMArena 1501 Elo、AIME 2025(带代码执行)100%。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: It tops the LMArena Leaderboard with a breakthrough score of 1501 Elo.
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Maps to existing benchmark arena (Chatbot Arena / LMArena). Elo snapshot at release time; leaderboard moves, so treat as a dated snapshot.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: top scores on Humanity's Last Exam (37.5% without the usage of any tools)
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark hlehle. Tool-free condition stated on page (protocol.tools = []), which is the comparability-critical field; tool-augmented HLE numbers appear only in the chart image and were not machine-read.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: and GPQA Diamond (91.9%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark gpqa (canonical name is already GPQA Diamond), so benchmark_variant is null.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: achieving a new state-of-the-art of 23.4% on MathArena Apex
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: matharena-apex not yet in data/benchmarks/. Distinct from frontiermath and math500 already in the catalog; MathArena Apex is the hardest MathArena tier.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: redefines multimodal reasoning with 81% on MMMU-Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark mmmu with variant Pro (MMMU-Pro is a distinct reformatted edition of MMMU, not directly comparable to vanilla MMMU scores).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: 87.6% on Video-MMMU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: video-mmmu not yet in data/benchmarks/. Video multimodal understanding benchmark; no relation to the mmmu entry.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: scores a state-of-the-art 72.1% on SimpleQA Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark simpleqa with variant Verified. SimpleQA Verified is a decontaminated re-scored edition; not comparable to original SimpleQA numbers without variant handling.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · quote_snippet: tops the WebDev Arena leaderboard by scoring an impressive 1487 Elo
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}new-benchmark: webdev-arena not yet in data/benchmarks/. Human-preference arena for web dev; distinct from arena (Chatbot Arena) and arenahard.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · quote_snippet: It also scores 54.2% on Terminal-Bench 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark terminalbench, variant 2.0. Page describes it as testing tool use to operate a computer via terminal. Not comparable to Terminal-Bench 2.1 rows recorded for other vendors.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · quote_snippet: it greatly outperforms 2.5 Pro on SWE-bench Verified (76.2%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark swebench (canonical name is SWE-bench Verified), variant null. Gemini 2.5 Pro's own score is not printed on this page; only the relative claim.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Plan anything · figure: gemini_3_table_final_HLE_Tools_on.gif — Vending-Bench 2 row (Net worth (mean), higher is better), Gemini 3 Pro column; trend chart vending_bench_2_final.webp shows the same 5-run average curves · quote_snippet: topping the leaderboard on Vending-Bench 2, which tests longer horizon planning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": "one full simulated year of operation",
"run_count": null,
"aggregation": "mean net worth, average over 5 runs",
"judge": null
}new-benchmark: vending-bench-2 not yet in data/benchmarks/. Prose makes only the leaderboard-top claim; the number lives in the archived evaluation-table GIF and is visually transcribed (2026-09-01): Gemini 3 Pro $5,478.16 vs Gemini 2.5 Pro $573.64 / Claude Sonnet 4.5 $3,838.74 / GPT-5.1 $1,473.43. Chart header states "Average over 5 runs per model".
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: Humanity's Last Exam — With search and code execution · figure: gemini_3_table_final_HLE_Tools_on.gif — Humanity's Last Exam — With search and code execution row, Gemini 3 Pro column · quote_snippet: Humanity's Last Exam: 37.5% no tools / 45.8% with search and code execution
{
"harness": null,
"tools": [
"search",
"code execution"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Second protocol row of the HLE pair on the launch table. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini-only row — the 2.5 Pro / Claude Sonnet 4.5 / GPT-5.1 cells are em-dash. Not comparable to the no-tools hlehle row of the same release.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: ARC-AGI-2 — ARC Prize Verified · figure: gemini_3_table_final_HLE_Tools_on.gif — ARC-AGI-2 — ARC Prize Verified row, Gemini 3 Pro column · quote_snippet: ARC-AGI-2: visual reasoning puzzles, ARC Prize Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": "ARC Prize Verified"
}Maps to existing benchmark arc-agi, variant 2. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 4.9% / Claude Sonnet 4.5 13.6% / GPT-5.1 17.6%; the Deep Think chart (final_dt_blog_evals_2.gif) additionally shows GPT-5 Pro 15.8%. Same protocol family as the Deep Think 45.1% row (code execution + ARC Prize Verified).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: AIME 2025 — No tools · figure: gemini_3_table_final_HLE_Tools_on.gif — AIME 2025 — No tools row, Gemini 3 Pro column · quote_snippet: AIME 2025: mathematics
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 88.0% / Claude Sonnet 4.5 87.0% / GPT-5.1 94.0%. Distinct protocol from the with-code-execution row of the same release.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: AIME 2025 — With code execution · figure: gemini_3_table_final_HLE_Tools_on.gif — AIME 2025 — With code execution row, Gemini 3 Pro column · quote_snippet: AIME 2025: mathematics, with code execution
{
"harness": null,
"tools": [
"code execution"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Claude Sonnet 4.5 also 100%; Gemini 2.5 Pro and GPT-5.1 cells are em-dash. Not comparable to the no-tools row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: ScreenSpot-Pro — Screen understanding · figure: gemini_3_table_final_HLE_Tools_on.gif — ScreenSpot-Pro — Screen understanding row, Gemini 3 Pro column · quote_snippet: ScreenSpot-Pro: screen understanding
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: screenspot-pro already in data/benchmarks/. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 11.4% / Claude Sonnet 4.5 36.2% / GPT-5.1 3.5%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: CharXiv Reasoning — Information synthesis from complex charts · figure: gemini_3_table_final_HLE_Tools_on.gif — CharXiv Reasoning — Information synthesis from complex charts row, Gemini 3 Pro column · quote_snippet: CharXiv Reasoning: information synthesis from complex charts
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark charxiv-reasoning. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 69.6% / Claude Sonnet 4.5 68.5% / GPT-5.1 69.5%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: OmniDocBench 1.5 — OCR (Overall Edit Distance, lower is better) · figure: gemini_3_table_final_HLE_Tools_on.gif — OmniDocBench 1.5 — OCR (Overall Edit Distance, lower is better) row, Gemini 3 Pro column · quote_snippet: OmniDocBench 1.5: OCR
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark omnidocbench. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 0.145 / Claude Sonnet 4.5 0.145 / GPT-5.1 0.147. Lower is better, so higher-is-better cross-model comparisons must not be drawn against percent rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI (Elo Rating, higher is better) · figure: gemini_3_table_final_HLE_Tools_on.gif — LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI (Elo Rating, higher is better) row, Gemini 3 Pro column · quote_snippet: LiveCodeBench Pro: Elo rating, higher is better
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}LiveCodeBench Pro subset (Elo) — distinct from the v5 pass@1 windows used elsewhere; not comparable across variants. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 1,775 / Claude Sonnet 4.5 1,418 / GPT-5.1 2,243.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: τ2-bench — Agentic tool use · figure: gemini_3_table_final_HLE_Tools_on.gif — τ2-bench — Agentic tool use row, Gemini 3 Pro column · quote_snippet: τ2-bench: agentic tool use
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark tau2-bench. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 54.9% / Claude Sonnet 4.5 84.7% / GPT-5.1 80.2%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: FACTS Benchmark Suite — Held out internal grounding, parametric, MM, and search retrieval benchmarks · figure: gemini_3_table_final_HLE_Tools_on.gif — FACTS Benchmark Suite — Held out internal grounding, parametric, MM, and search retrieval benchmarks row, Gemini 3 Pro column · quote_snippet: FACTS Benchmark Suite: held out internal grounding, parametric, MM, and search retrieval benchmarks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: facts-suite not yet in data/benchmarks/. Distinct from factsg (FACTS Grounding, a single grounding eval): this row is Google's internal multi-capability FACTS suite aggregate. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 63.4% / Claude Sonnet 4.5 50.4% / GPT-5.1 50.8%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: MMMLU — Multilingual Q&A · figure: gemini_3_table_final_HLE_Tools_on.gif — MMMLU — Multilingual Q&A row, Gemini 3 Pro column · quote_snippet: MMMLU: multilingual Q&A
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark mmmlu. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 89.5% / Claude Sonnet 4.5 89.1% / GPT-5.1 91.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: Global PIQA — Commonsense reasoning across 100 Languages and Cultures · figure: gemini_3_table_final_HLE_Tools_on.gif — Global PIQA — Commonsense reasoning across 100 Languages and Cultures row, Gemini 3 Pro column · quote_snippet: Global PIQA: commonsense reasoning across 100 languages and cultures
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark global-piqa. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 91.5% / Claude Sonnet 4.5 90.1% / GPT-5.1 90.9%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: MRCR v2 (8-needle) — 128k (average) · figure: gemini_3_table_final_HLE_Tools_on.gif — MRCR v2 (8-needle) — 128k (average) row, Gemini 3 Pro column · quote_snippet: MRCR v2 (8-needle): long context performance, 128k average
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": "cumulative score at 128k (average)",
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 58.0% / Claude Sonnet 4.5 47.1% / GPT-5.1 61.6%. Same MRCR v2 8-needle methodology as the June 2.5 family table.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Build anything · row: MRCR v2 (8-needle) — 1M (pointwise) · figure: gemini_3_table_final_HLE_Tools_on.gif — MRCR v2 (8-needle) — 1M (pointwise) row, Gemini 3 Pro column · quote_snippet: MRCR v2 (8-needle): long context performance, 1M pointwise
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": "pointwise value at 1M context",
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 16.4%; Claude Sonnet 4.5 and GPT-5.1 cells read "not supported" in the GIF. Separate row from the 128k average because the aggregation differs.
Gemini 3 Deep Think
Gemini 3 Deep Think 为 Gemini 3 的增强推理模式,发布时访问仅限安全测试者。已收录 3 项页面文本评测:亮点 ARC-AGI-2(ARC Prize Verified)45.1%、GPQA Diamond 93.8%。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Gemini 3 Deep Think · quote_snippet: Humanity's Last Exam (41.0% without the use of tools)
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": "Deep Think (parallel reasoning; effort level not quantified)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Deep Think row reported on the Gemini 3 launch page; the dedicated Deep Think rollout post (gemini-3-deep-think release) repeats the same value, so both releases carry this evidence from their own source pages.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Gemini 3 Deep Think · quote_snippet: and GPQA Diamond (93.8%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Deep Think (parallel reasoning; effort level not quantified)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}GPQA Diamond 93.8% for Deep Think appears only on the Gemini 3 launch page; the dedicated rollout post omits GPQA, so this page is the sole A-tier source for this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Gemini 3 Deep Think · quote_snippet: an unprecedented 45.1% on ARC-AGI-2 (with code execution, ARC Prize Verified)
{
"harness": null,
"tools": [
"code execution"
],
"shots": null,
"reasoning_effort": "Deep Think (parallel reasoning; effort level not quantified)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "ARC Prize Verified"
}Maps to existing benchmark arc-agi, variant 2. Comparability-critical: code execution enabled and ARC Prize Verified validation; not comparable to tool-free ARC-AGI-2 rows.