← 模型目录

Gemini 3 Pro / Gemini 3 Deep Think

Google DeepMind / Gemini · 2025-11-18 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 3 Pro

Gemini 3 发布以「智能新纪元」为题呈现,Gemini 3 Pro 为旗舰型号(发布时 preview)。已收录评测横跨推理、代码、多模态与 Agent:亮点 LMArena 1501 Elo、AIME 2025(带代码执行)100%。

输入模态
文本 / 图像 / 视频
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

arena 1501 Elo 模型 gemini-3-pro · 版本 未说明 · 指标 arena_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: It tops the LMArena Leaderboard with a breakthrough score of 1501 Elo.

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard",
  "judge": null
}

Maps to existing benchmark arena (Chatbot Arena / LMArena). Elo snapshot at release time; leaderboard moves, so treat as a dated snapshot.

打开官方来源

hlehle 37.5% 模型 gemini-3-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: top scores on Humanity's Last Exam (37.5% without the usage of any tools)

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark hlehle. Tool-free condition stated on page (protocol.tools = []), which is the comparability-critical field; tool-augmented HLE numbers appear only in the chart image and were not machine-read.

打开官方来源

gpqa 91.9% 模型 gemini-3-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: and GPQA Diamond (91.9%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark gpqa (canonical name is already GPQA Diamond), so benchmark_variant is null.

打开官方来源

apex 23.4% 模型 gemini-3-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: achieving a new state-of-the-art of 23.4% on MathArena Apex

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: matharena-apex not yet in data/benchmarks/. Distinct from frontiermath and math500 already in the catalog; MathArena Apex is the hardest MathArena tier.

打开官方来源

mmmu 81% 模型 gemini-3-pro · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: redefines multimodal reasoning with 81% on MMMU-Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mmmu with variant Pro (MMMU-Pro is a distinct reformatted edition of MMMU, not directly comparable to vanilla MMMU scores).

打开官方来源

video-mmmu 87.6% 模型 gemini-3-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: 87.6% on Video-MMMU

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: video-mmmu not yet in data/benchmarks/. Video multimodal understanding benchmark; no relation to the mmmu entry.

打开官方来源

simpleqa 72.1% 模型 gemini-3-pro · 版本 Verified · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: State-of-the-art reasoning with unprecedented depth and nuance · quote_snippet: scores a state-of-the-art 72.1% on SimpleQA Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark simpleqa with variant Verified. SimpleQA Verified is a decontaminated re-scored edition; not comparable to original SimpleQA numbers without variant handling.

打开官方来源

webdev-arena 1487 Elo 模型 gemini-3-pro · 版本 未说明 · 指标 arena_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · quote_snippet: tops the WebDev Arena leaderboard by scoring an impressive 1487 Elo

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard",
  "judge": null
}

new-benchmark: webdev-arena not yet in data/benchmarks/. Human-preference arena for web dev; distinct from arena (Chatbot Arena) and arenahard.

打开官方来源

terminalbench 54.2% 模型 gemini-3-pro · 版本 2.0 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · quote_snippet: It also scores 54.2% on Terminal-Bench 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark terminalbench, variant 2.0. Page describes it as testing tool use to operate a computer via terminal. Not comparable to Terminal-Bench 2.1 rows recorded for other vendors.

打开官方来源

swebench 76.2% 模型 gemini-3-pro · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · quote_snippet: it greatly outperforms 2.5 Pro on SWE-bench Verified (76.2%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark swebench (canonical name is SWE-bench Verified), variant null. Gemini 2.5 Pro's own score is not printed on this page; only the relative claim.

打开官方来源

vending-bench-2 $5,478.16 (mean net worth, average over 5 runs) 模型 gemini-3-pro · 版本 未说明 · 指标 net_worth_mean · 单位 usd 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Plan anything · figure: gemini_3_table_final_HLE_Tools_on.gif — Vending-Bench 2 row (Net worth (mean), higher is better), Gemini 3 Pro column; trend chart vending_bench_2_final.webp shows the same 5-run average curves · quote_snippet: topping the leaderboard on Vending-Bench 2, which tests longer horizon planning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": "one full simulated year of operation",
  "run_count": null,
  "aggregation": "mean net worth, average over 5 runs",
  "judge": null
}

new-benchmark: vending-bench-2 not yet in data/benchmarks/. Prose makes only the leaderboard-top claim; the number lives in the archived evaluation-table GIF and is visually transcribed (2026-09-01): Gemini 3 Pro $5,478.16 vs Gemini 2.5 Pro $573.64 / Claude Sonnet 4.5 $3,838.74 / GPT-5.1 $1,473.43. Chart header states "Average over 5 runs per model".

打开官方来源

hlehle 45.8% (with search and code execution) 模型 gemini-3-pro · 版本 with search and code execution · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: Humanity's Last Exam — With search and code execution · figure: gemini_3_table_final_HLE_Tools_on.gif — Humanity's Last Exam — With search and code execution row, Gemini 3 Pro column · quote_snippet: Humanity's Last Exam: 37.5% no tools / 45.8% with search and code execution

{
  "harness": null,
  "tools": [
    "search",
    "code execution"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Second protocol row of the HLE pair on the launch table. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini-only row — the 2.5 Pro / Claude Sonnet 4.5 / GPT-5.1 cells are em-dash. Not comparable to the no-tools hlehle row of the same release.

打开官方来源

arc-agi 31.1% (ARC Prize Verified) 模型 gemini-3-pro · 版本 2 (ARC Prize Verified) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: ARC-AGI-2 — ARC Prize Verified · figure: gemini_3_table_final_HLE_Tools_on.gif — ARC-AGI-2 — ARC Prize Verified row, Gemini 3 Pro column · quote_snippet: ARC-AGI-2: visual reasoning puzzles, ARC Prize Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "ARC Prize Verified"
}

Maps to existing benchmark arc-agi, variant 2. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 4.9% / Claude Sonnet 4.5 13.6% / GPT-5.1 17.6%; the Deep Think chart (final_dt_blog_evals_2.gif) additionally shows GPT-5 Pro 15.8%. Same protocol family as the Deep Think 45.1% row (code execution + ARC Prize Verified).

打开官方来源

aime-25 95.0% (no tools) 模型 gemini-3-pro · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: AIME 2025 — No tools · figure: gemini_3_table_final_HLE_Tools_on.gif — AIME 2025 — No tools row, Gemini 3 Pro column · quote_snippet: AIME 2025: mathematics

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 88.0% / Claude Sonnet 4.5 87.0% / GPT-5.1 94.0%. Distinct protocol from the with-code-execution row of the same release.

打开官方来源

aime-25 100% (with code execution) 模型 gemini-3-pro · 版本 with code execution · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: AIME 2025 — With code execution · figure: gemini_3_table_final_HLE_Tools_on.gif — AIME 2025 — With code execution row, Gemini 3 Pro column · quote_snippet: AIME 2025: mathematics, with code execution

{
  "harness": null,
  "tools": [
    "code execution"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Claude Sonnet 4.5 also 100%; Gemini 2.5 Pro and GPT-5.1 cells are em-dash. Not comparable to the no-tools row.

打开官方来源

screenspot-pro 72.7% 模型 gemini-3-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: ScreenSpot-Pro — Screen understanding · figure: gemini_3_table_final_HLE_Tools_on.gif — ScreenSpot-Pro — Screen understanding row, Gemini 3 Pro column · quote_snippet: ScreenSpot-Pro: screen understanding

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: screenspot-pro already in data/benchmarks/. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 11.4% / Claude Sonnet 4.5 36.2% / GPT-5.1 3.5%.

打开官方来源

charxiv-reasoning 81.4% 模型 gemini-3-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: CharXiv Reasoning — Information synthesis from complex charts · figure: gemini_3_table_final_HLE_Tools_on.gif — CharXiv Reasoning — Information synthesis from complex charts row, Gemini 3 Pro column · quote_snippet: CharXiv Reasoning: information synthesis from complex charts

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark charxiv-reasoning. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 69.6% / Claude Sonnet 4.5 68.5% / GPT-5.1 69.5%.

打开官方来源

omnidocbench 0.115 (Overall Edit Distance, lower is better) 模型 gemini-3-pro · 版本 1.5 (Overall Edit Distance, lower is better) · 指标 overall_edit_distance · 单位 edit_distance 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: OmniDocBench 1.5 — OCR (Overall Edit Distance, lower is better) · figure: gemini_3_table_final_HLE_Tools_on.gif — OmniDocBench 1.5 — OCR (Overall Edit Distance, lower is better) row, Gemini 3 Pro column · quote_snippet: OmniDocBench 1.5: OCR

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark omnidocbench. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 0.145 / Claude Sonnet 4.5 0.145 / GPT-5.1 0.147. Lower is better, so higher-is-better cross-model comparisons must not be drawn against percent rows.

打开官方来源

lcb 2,439 Elo (LiveCodeBench Pro) 模型 gemini-3-pro · 版本 Pro (Codeforces/ICPC/IOI), Elo · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI (Elo Rating, higher is better) · figure: gemini_3_table_final_HLE_Tools_on.gif — LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI (Elo Rating, higher is better) row, Gemini 3 Pro column · quote_snippet: LiveCodeBench Pro: Elo rating, higher is better

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

LiveCodeBench Pro subset (Elo) — distinct from the v5 pass@1 windows used elsewhere; not comparable across variants. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 1,775 / Claude Sonnet 4.5 1,418 / GPT-5.1 2,243.

打开官方来源

tau2-bench 85.4% 模型 gemini-3-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: τ2-bench — Agentic tool use · figure: gemini_3_table_final_HLE_Tools_on.gif — τ2-bench — Agentic tool use row, Gemini 3 Pro column · quote_snippet: τ2-bench: agentic tool use

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark tau2-bench. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 54.9% / Claude Sonnet 4.5 84.7% / GPT-5.1 80.2%.

打开官方来源

facts-suite 70.5% 模型 gemini-3-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: FACTS Benchmark Suite — Held out internal grounding, parametric, MM, and search retrieval benchmarks · figure: gemini_3_table_final_HLE_Tools_on.gif — FACTS Benchmark Suite — Held out internal grounding, parametric, MM, and search retrieval benchmarks row, Gemini 3 Pro column · quote_snippet: FACTS Benchmark Suite: held out internal grounding, parametric, MM, and search retrieval benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: facts-suite not yet in data/benchmarks/. Distinct from factsg (FACTS Grounding, a single grounding eval): this row is Google's internal multi-capability FACTS suite aggregate. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 63.4% / Claude Sonnet 4.5 50.4% / GPT-5.1 50.8%.

打开官方来源

mmmlu 91.8% 模型 gemini-3-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: MMMLU — Multilingual Q&A · figure: gemini_3_table_final_HLE_Tools_on.gif — MMMLU — Multilingual Q&A row, Gemini 3 Pro column · quote_snippet: MMMLU: multilingual Q&A

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mmmlu. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 89.5% / Claude Sonnet 4.5 89.1% / GPT-5.1 91.0%.

打开官方来源

global-piqa 93.4% 模型 gemini-3-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: Global PIQA — Commonsense reasoning across 100 Languages and Cultures · figure: gemini_3_table_final_HLE_Tools_on.gif — Global PIQA — Commonsense reasoning across 100 Languages and Cultures row, Gemini 3 Pro column · quote_snippet: Global PIQA: commonsense reasoning across 100 languages and cultures

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark global-piqa. Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 91.5% / Claude Sonnet 4.5 90.1% / GPT-5.1 90.9%.

打开官方来源

mrcr 77.0% (v2 8-needle, 128k average) 模型 gemini-3-pro · 版本 v2 (8-needle), 128k (average) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: MRCR v2 (8-needle) — 128k (average) · figure: gemini_3_table_final_HLE_Tools_on.gif — MRCR v2 (8-needle) — 128k (average) row, Gemini 3 Pro column · quote_snippet: MRCR v2 (8-needle): long context performance, 128k average

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": "cumulative score at 128k (average)",
  "judge": null
}

Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 58.0% / Claude Sonnet 4.5 47.1% / GPT-5.1 61.6%. Same MRCR v2 8-needle methodology as the June 2.5 family table.

打开官方来源

mrcr 26.3% (v2 8-needle, 1M pointwise) 模型 gemini-3-pro · 版本 v2 (8-needle), 1M (pointwise) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Build anything · row: MRCR v2 (8-needle) — 1M (pointwise) · figure: gemini_3_table_final_HLE_Tools_on.gif — MRCR v2 (8-needle) — 1M (pointwise) row, Gemini 3 Pro column · quote_snippet: MRCR v2 (8-needle): long context performance, 1M pointwise

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": "pointwise value at 1M context",
  "judge": null
}

Verified visually 2026-09-01 by reading the archived GIF (models/2025-11-18-gemini-3/images/03.gif): Gemini 2.5 Pro 16.4%; Claude Sonnet 4.5 and GPT-5.1 cells read "not supported" in the GIF. Separate row from the 128k average because the aggregation differs.

打开官方来源

Gemini 3 Deep Think

Gemini 3 Deep Think 为 Gemini 3 的增强推理模式,发布时访问仅限安全测试者。已收录 3 项页面文本评测:亮点 ARC-AGI-2(ARC Prize Verified)45.1%、GPQA Diamond 93.8%。

输入模态
文本 / 图像 / 视频
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 41.0% 模型 gemini-3-deep-think · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 3 Deep Think · quote_snippet: Humanity's Last Exam (41.0% without the use of tools)

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": "Deep Think (parallel reasoning; effort level not quantified)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Deep Think row reported on the Gemini 3 launch page; the dedicated Deep Think rollout post (gemini-3-deep-think release) repeats the same value, so both releases carry this evidence from their own source pages.

打开官方来源

gpqa 93.8% 模型 gemini-3-deep-think · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 3 Deep Think · quote_snippet: and GPQA Diamond (93.8%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Deep Think (parallel reasoning; effort level not quantified)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

GPQA Diamond 93.8% for Deep Think appears only on the Gemini 3 launch page; the dedicated rollout post omits GPQA, so this page is the sole A-tier source for this row.

打开官方来源

arc-agi 45.1% 模型 gemini-3-deep-think · 版本 2 (ARC Prize Verified) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 3 Deep Think · quote_snippet: an unprecedented 45.1% on ARC-AGI-2 (with code execution, ARC Prize Verified)

{
  "harness": null,
  "tools": [
    "code execution"
  ],
  "shots": null,
  "reasoning_effort": "Deep Think (parallel reasoning; effort level not quantified)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "ARC Prize Verified"
}

Maps to existing benchmark arc-agi, variant 2. Comparability-critical: code execution enabled and ARC Prize Verified validation; not comparable to tool-free ARC-AGI-2 rows.

打开官方来源