Gemini 2.5 Pro
Google DeepMind / Gemini · 2025-03-25 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 2.5 Pro
谷歌将 Gemini 2.5 Pro 实验版定位为其最智能的 AI 模型与 2.5 系列首个思维(thinking)模型,原生多模态并配备 1M token 上下文窗口(200 万即将推出)。已收录评测覆盖推理、代码、长上下文与多模态理解等领域,HLE 无工具 18.8%、SWE-bench Verified 63.8%。
- 输入模态
- 文本 / 图像 / 音频 / 视频
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enhanced reasoning · quote_snippet: scores a state-of-the-art 18.8% across models without tool use on Humanity's Last Exam
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark hlehle. Tool-free condition is the comparability-critical field (protocol.tools = []); tool-augmented HLE rows from other vendors are not comparable. GIF table confirms the same 18.8% figure.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Advanced coding · quote_snippet: On SWE-Bench Verified... Gemini 2.5 Pro scores 63.8% with a custom agent setup
{
"harness": "custom agent setup",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark swebench (canonical SWE-bench Verified). Harness explicitly stated as a custom agent setup — not the scaffolding used by Claude/GPT entries in the same GIF table (footnote: 'All SWE-bench verified numbers follow official provider reports, using different scaffolding and infrastructure'), so cross-vendor SWE-bench comparisons are harness-confounded.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enhanced reasoning · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Science / GPQA diamond row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: 2.5 Pro leads in math and science benchmarks like GPQA and AIME 2025
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Named in prose with a leading claim but no printed number. Sentence also explicitly excludes test-time techniques (majority voting), i.e. single-attempt condition. GIF table (gemini_benchmarks_cropped_light2x_1PPmDuP.gif) visually reads GPQA diamond single attempt (pass@1) 84.0%5. Score transcribed 2026-09-01 from the archived GIF (prose itself prints no number): competitor single-attempt cells o3-mini High 79.7% / GPT-4.5 71.4% / Claude 3.7 Sonnet (64k thinking) 78.2% & 84.8% multiple / Grok 3 Beta 80.2% & 84.6% multiple / DeepSeek R1 71.5%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enhanced reasoning · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Mathematics / AIME 2025 row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: Without test-time techniques that increase cost, like majority voting, 2.5 Pro leads in... AIME 2025
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-25 (registered in batch 2, still not in data/benchmarks/). Named in prose, no printed number; GIF table visually reads AIME 2025 single attempt (pass@1) 86.7%. The no-majority-voting condition is protocol-critical vs cons@k rows (e.g. Grok 3 Think cons@64). Score transcribed 2026-09-01 from the archived GIF (prose itself prints no number): competitor single-attempt cells o3-mini High 86.5% / Claude 3.7 Sonnet 49.5% / Grok 3 Beta 77.3% single & 93.3% multiple / DeepSeek R1 70.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro · quote_snippet: debuts at #1 on LMArena by a significant margin
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Maps to existing benchmark arena (Chatbot Arena / LMArena). Rank-#1 claim without a printed Elo on this page (the June 2025 follow-up post later reports 1470 for the upgraded preview — different snapshot, different release).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Long context / MRCR row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: Updated March 26 with new MRCR (Multi Round Coreference Resolution) evaluations
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mrcr not yet in data/benchmarks/ (also used on the Gemini 2.0 chart at 1M, and as MRCR v2 8-needle in the June family table — version differences must be tracked as variants). Named via the page's own update line; numbers live in the GIF table. Scores transcribed 2026-09-01 from the archived GIF: 128k (average) 94.5% (1M pointwise recorded as the sibling row google-gemini-2-5-pro--mrcr-1m) — the GIF reports only these two MRCR rows; competitor 128k cells o3-mini High 61.4% / GPT-4.5 64.0%, all other cells em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro · row: Long context / MRCR — 1M (pointwise) · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Long context / MRCR — 1M (pointwise) row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: MRCR — 1M (pointwise)
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": "pointwise value at 1M context (per chart footnote)",
"judge": null
}GIF-table row added during the 2026-09-01 audit: separate row from the 128k (average) entry because the chart footnote defines different aggregations (128k cumulative vs 1M pointwise). Confirmed 2026-09-01 by reading the archived GIF (models/2025-03-25-gemini-2-5-pro/images/03.gif): value transcribed visually, all six competitor cells in the 1M (pointwise) row are em-dash (Gemini-only row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro (benchmark table) · row: Mathematics / AIME 2024 — single attempt (pass@1) · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Mathematics / AIME 2024 row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: AIME 2024 single attempt (pass@1)
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "single attempt (pass@1)",
"judge": null
}Row only exists inside the GIF benchmark table; visual reading gives 92.0%5, pending manual read. Footnote (visual read, partially garbled) states Gemini scores were run via AI Studio API on gemini-2.5-pro-exp-03-25 with default sampling, averaging over multiple trials for smaller benchmarks. Confirmed 2026-09-01 by reading the archived GIF (models/2025-03-25-gemini-2-5-pro/images/03.gif): competitor cells o3-mini High 87.3% / GPT-4.5 36.7% / Claude 3.7 Sonnet (64k thinking) 61.3% single & 80.0% multiple / Grok 3 Beta 83.9% single & 93.3% multiple / DeepSeek R1 79.8%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro (benchmark table) · row: Code generation / LiveCodeBench v5 — single attempt (pass@1) · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Code generation / LiveCodeBench v5 row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: LiveCodeBench v5 single attempt (pass@1)
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "single attempt (pass@1)",
"judge": null
}GIF-table row; visual reading 70.4%. LCB v5 window on this table differs from the v5 1/1/2025-5/1/2025 window used in the June family table (69.0% for Pro there) — treat windows as distinct variants. Confirmed 2026-09-01 from the archived GIF: competitor cells o3-mini High 74.1% (best in row) / Grok 3 Beta 70.6% single & 79.4% multiple / DeepSeek R1 64.3%; GPT-4.5 and Claude 3.7 cells are em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro (benchmark table) · row: Code editing / Aider Polyglot (whole / diff) · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Code editing / Aider Polyglot row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: Aider Polyglot: 74.0% / 68.6% (whole / diff)
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}GIF-table row; visual reading 74.0% whole / 68.6% diff edit formats. Maps to existing benchmark aider (Aider Polyglot leaderboard benchmark). Confirmed 2026-09-01 from the archived GIF: Gemini is the only whole-format entry; competitor diff-format cells o3-mini High 60.4% / GPT-4.5 44.9% / Claude 3.7 Sonnet (32k thinking) 64.9% / DeepSeek R1 56.9%; Grok 3 cell is em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro (benchmark table) · row: Factuality / SimpleQA · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Factuality / SimpleQA row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: Factuality / SimpleQA
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}GIF-table row. Maps to existing benchmark simpleqa. Confirmed 2026-09-01 from the archived GIF: competitor cells o3-mini High 13.8% / GPT-4.5 62.5% (best in row) / Grok 3 Beta 43.6% / DeepSeek R1 30.1%; Claude 3.7 cell is em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro (benchmark table) · row: Visual reasoning / MMMU — single attempt (pass@1) · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Visual reasoning / MMMU row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: Visual reasoning / MMMU single attempt (pass@1)
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "single attempt (pass@1)",
"judge": null
}GIF-table row; visual reading 81.7%. A second MMMU row (multiple attempts) is marked 'no MM support' for a competitor column in the visual read — Gemini column only carries the single-attempt figure. Confirmed 2026-09-01 from the archived GIF: competitor single-attempt cells GPT-4.5 74.4% / Claude 3.7 Sonnet (64k thinking) 75.0% / Grok 3 Beta 76.0% single & 78.0% multiple; o3-mini and DeepSeek R1 carry "no MM support".
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro (benchmark table) · row: Image understanding / Vibe-Eval (Reka) · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Image understanding / Vibe-Eval (Reka) row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: Image understanding / Vibe-Eval (Reka)
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "Gemini as judge (per footnote)"
}new-benchmark: vibe-eval (introduced in this batch by gemini-2-0). Visual reading 69.4%. Judge = Gemini (self-judging bias risk, explicitly stated in footnote). Confirmed 2026-09-01 from the archived GIF: Gemini is the only scored column (judge = Gemini per footnote); o3-mini and DeepSeek R1 carry "no MM support", remaining cells em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Pro (benchmark table) · row: Multilingual performance / Global MMLU (Lite) · figure: gemini_benchmarks_cropped_light2x_1PPmDuP.gif — Multilingual performance / Global MMLU (Lite) row, Gemini 2.5 Pro Experimental (03-25) column · quote_snippet: Multilingual performance / Global MMLU (Lite)
{
"harness": "AI Studio API, model id gemini-2.5-pro-exp-03-25, default sampling",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: global-mmlu-lite not yet in data/benchmarks/ (Global MMLU Lite, distinct from mmlu/mmlu-pro). Visual reading 89.8%. Confirmed 2026-09-01 from the archived GIF: all five competitor cells are em-dash (Gemini-only row).