← 模型目录

Gemini 3 Flash

Google DeepMind / Gemini · 2025-12-17 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 3 Flash

Gemini 3 Flash 以「为速度打造的前沿智能」为定位发布,上线即为 Gemini 应用与搜索 AI Mode 默认模型(发布时为 preview)。已收录 25 项评测横跨推理、代码、多模态与 Agent:亮点 AIME 2025(带代码执行)99.7%、SWE-bench Verified 78%。

输入模态
文本 / 图像 / 视频
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 0.5 / 输出 3

本变体的评测证据

gpqa 90.4% 模型 gemini-3-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 3 Flash: frontier intelligence at scale · quote_snippet: PhD-level reasoning and knowledge benchmarks like GPQA Diamond (90.4%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark gpqa (canonical GPQA Diamond). Thinking level / sampling not stated in prose.

打开官方来源

hlehle 33.7% 模型 gemini-3-flash · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 3 Flash: frontier intelligence at scale · quote_snippet: Humanity's Last Exam (33.7% without tools)

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free condition explicit (protocol.tools = []); not comparable to tool-augmented HLE rows.

打开官方来源

mmmu 81.2% 模型 gemini-3-flash · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 3 Flash: frontier intelligence at scale · quote_snippet: state-of-the-art performance with an impressive score of 81.2% on MMMU Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mmmu with variant Pro (reformatted edition; not comparable to vanilla MMMU).

打开官方来源

swebench 78% 模型 gemini-3-flash · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: For developers: intelligence that keeps up · quote_snippet: On SWE-bench Verified... Gemini 3 Flash achieves a score of 78%, outperforming not only the 2.5 series, but also Gemini 3 Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark swebench (SWE-bench Verified). Harness/scaffold not stated — later Gemini model cards (3.6/3.7) publish harness detail this post omits.

打开官方来源

arena 官方未公布数值 模型 gemini-3-flash · 版本 未说明 · 指标 arena_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 3 Flash: frontier intelligence at scale · figure: Pareto scatter (LMArena Elo vs price), image only · quote_snippet: Performance, here, is measured by LMArena Elo Score.

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": "leaderboard",
  "judge": null
}

LMArena Elo named in prose as the performance axis of the Pareto claim; the Elo value itself lives only in the scatter image.

打开官方来源

hlehle 43.5% (with search and code execution) 模型 gemini-3-flash · 版本 with search and code execution · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: Humanity's Last Exam — With search and code execution · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — Humanity's Last Exam — With search and code execution row, Gemini 3 Flash Thinking column · quote_snippet: HLE full set (text + MM): no tools vs with search and code execution

{
  "harness": null,
  "tools": [
    "search",
    "code execution"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 45.8% / GPT-5.2 Extra high 45.5%; 2.5 Flash, 2.5 Pro and Claude Sonnet 4.5 cells em-dash. Distinct protocol from the no-tools hlehle row of the same release.

打开官方来源

arc-agi 33.6% (ARC Prize Verified) 模型 gemini-3-flash · 版本 2 (ARC Prize Verified) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: ARC-AGI-2 — ARC Prize Verified · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — ARC-AGI-2 — ARC Prize Verified row, Gemini 3 Flash Thinking column · quote_snippet: ARC-AGI-2: visual reasoning puzzles

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "ARC Prize Verified"
}

Maps to existing benchmark arc-agi, variant 2. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 31.1% / 2.5 Flash 2.5% / 2.5 Pro 4.9% / Claude Sonnet 4.5 13.6% / GPT-5.2 Extra high 52.9% (best in row); Grok 4.1 Fast em-dash.

打开官方来源

aime-25 95.2% (no tools) 模型 gemini-3-flash · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: AIME 2025 — No tools · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — AIME 2025 — No tools row, Gemini 3 Flash Thinking column · quote_snippet: AIME 2025: mathematics

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 95.0% / 2.5 Flash 72.0% / 2.5 Pro 88.0% / Claude Sonnet 4.5 87.0% / GPT-5.2 Extra high 100% (best in row) / Grok 4.1 Fast 91.9%. Distinct protocol from the with-code-execution row.

打开官方来源

aime-25 99.7% (with code execution) 模型 gemini-3-flash · 版本 with code execution · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: AIME 2025 — With code execution · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — AIME 2025 — With code execution row, Gemini 3 Flash Thinking column · quote_snippet: AIME 2025: mathematics, with code execution

{
  "harness": null,
  "tools": [
    "code execution"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 100% / 2.5 Flash 75.7% / Claude Sonnet 4.5 100%; 2.5 Pro, GPT-5.2 and Grok cells em-dash. Not comparable to the no-tools row.

打开官方来源

screenspot-pro 69.1% 模型 gemini-3-flash · 版本 no tools unless specified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: ScreenSpot-Pro — Screen understanding (No tools unless specified) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — ScreenSpot-Pro — Screen understanding (No tools unless specified) row, Gemini 3 Flash Thinking column · quote_snippet: ScreenSpot-Pro: screen understanding

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark screenspot-pro. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 72.7% / 2.5 Flash 3.9% / 2.5 Pro 11.4% / Claude Sonnet 4.5 36.2% / GPT-5.2 Extra high 86.3% (with python — different tool condition); Grok em-dash.

打开官方来源

charxiv-reasoning 80.3% 模型 gemini-3-flash · 版本 no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: CharXiv Reasoning — Information synthesis from complex charts (No tools) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — CharXiv Reasoning — Information synthesis from complex charts (No tools) row, Gemini 3 Flash Thinking column · quote_snippet: CharXiv Reasoning: information synthesis from complex charts

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark charxiv-reasoning. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 81.4% / 2.5 Flash 63.7% / 2.5 Pro 69.6% / Claude Sonnet 4.5 68.5% / GPT-5.2 Extra high 82.1% (best in row); Grok em-dash.

打开官方来源

omnidocbench 0.121 (Overall Edit Distance, lower is better) 模型 gemini-3-flash · 版本 1.5 (Overall Edit Distance, lower is better) · 指标 overall_edit_distance · 单位 edit_distance 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: OmniDocBench 1.5 — OCR (Overall Edit Distance, lower is better) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — OmniDocBench 1.5 — OCR (Overall Edit Distance, lower is better) row, Gemini 3 Flash Thinking column · quote_snippet: OmniDocBench 1.5: OCR

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark omnidocbench. Lower is better. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 0.115 (best in row) / 2.5 Flash 0.154 / 2.5 Pro 0.145 / Claude Sonnet 4.5 0.145 / GPT-5.2 Extra high 0.143; Grok em-dash.

打开官方来源

video-mmmu 86.9% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: Video-MMMU — Knowledge acquisition from videos · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — Video-MMMU — Knowledge acquisition from videos row, Gemini 3 Flash Thinking column · quote_snippet: Video-MMMU: knowledge acquisition from videos

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark video-mmmu. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 87.6% / 2.5 Flash 79.2% / 2.5 Pro 83.6% / Claude Sonnet 4.5 77.8% / GPT-5.2 Extra high 85.9%; Grok em-dash.

打开官方来源

lcb 2316 Elo (LiveCodeBench Pro) 模型 gemini-3-flash · 版本 Pro (Codeforces/ICPC/IOI), Elo · 指标 elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI (Elo rating, higher is better) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI (Elo rating, higher is better) row, Gemini 3 Flash Thinking column · quote_snippet: LiveCodeBench Pro: Elo rating, higher is better

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

LiveCodeBench Pro subset (Elo) — distinct from the v5 pass@1 windows used elsewhere. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 2439 / 2.5 Flash 1143 / 2.5 Pro 1775 / Claude Sonnet 4.5 1418 / GPT-5.2 Extra high 2393; Grok em-dash.

打开官方来源

terminalbench 47.6% (Terminus-2 harness) 模型 gemini-3-flash · 版本 2.0 (Terminus-2 harness) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: Terminal-bench 2.0 — Agentic terminal coding (Terminus-2 harness) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — Terminal-bench 2.0 — Agentic terminal coding (Terminus-2 harness) row, Gemini 3 Flash Thinking column · quote_snippet: Terminal-bench 2.0: agentic terminal coding

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark terminalbench, variant 2.0 with Terminus-2 harness. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 54.2% / 2.5 Flash 16.9% / 2.5 Pro 32.6% / Claude Sonnet 4.5 42.8%; GPT-5.2 and Grok em-dash.

打开官方来源

tau2-bench 90.2% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: τ2-bench — Agentic tool use · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — τ2-bench — Agentic tool use row, Gemini 3 Flash Thinking column · quote_snippet: τ2-bench: agentic tool use

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark tau2-bench. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 90.7% / 2.5 Flash 79.5% / 2.5 Pro 77.8% / Claude Sonnet 4.5 87.2%; GPT-5.2 and Grok em-dash.

打开官方来源

toolathlon 49.4% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: Toolathlon — Long horizon real-world software tasks · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — Toolathlon — Long horizon real-world software tasks row, Gemini 3 Flash Thinking column · quote_snippet: Toolathlon: long horizon real-world software tasks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark toolathlon. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 36.4% / 2.5 Flash 3.7% / 2.5 Pro 10.5% / Claude Sonnet 4.5 38.9% / GPT-5.2 Extra high 46.3%; Grok em-dash. Gemini 3 Flash best in row.

打开官方来源

mcp-atlas 57.4% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: MCP Atlas — Multi-step workflows using MCP · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — MCP Atlas — Multi-step workflows using MCP row, Gemini 3 Flash Thinking column · quote_snippet: MCP Atlas: multi-step workflows using MCP

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mcp-atlas. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 54.1% / 2.5 Flash 3.4% / 2.5 Pro 8.8% / Claude Sonnet 4.5 43.8% / GPT-5.2 Extra high 60.6% (best in row); Grok em-dash.

打开官方来源

vending-bench-2 $3,635 (net worth mean) 模型 gemini-3-flash · 版本 未说明 · 指标 net_worth_mean · 单位 usd 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: Vending-Bench 2 — Agentic long term coherence (Net worth (mean), higher is better) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — Vending-Bench 2 — Agentic long term coherence (Net worth (mean), higher is better) row, Gemini 3 Flash Thinking column · quote_snippet: Vending-Bench 2: net worth mean

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark vending-bench-2. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro $5,478 (best in row) / 2.5 Flash $549 / 2.5 Pro $574 / Claude Sonnet 4.5 $3,839 / GPT-5.2 Extra high $3,952 / Grok 4.1 Fast $1,107.

打开官方来源

facts-suite 61.9% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: FACTS Benchmark Suite — Factuality benchmark across grounding, parametric, search, and MM · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — FACTS Benchmark Suite — Factuality benchmark across grounding, parametric, search, and MM row, Gemini 3 Flash Thinking column · quote_snippet: FACTS Benchmark Suite: factuality across grounding, parametric, search, MM

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: facts-suite (same suite aggregate introduced on the Gemini 3 launch page; distinct from factsg / FACTS Grounding). Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 70.5% / 2.5 Flash 50.4% / 2.5 Pro 63.4% / Claude Sonnet 4.5 48.9% / GPT-5.2 Extra high 61.4% / Grok 4.1 Fast 42.1%.

打开官方来源

simpleqa-verified 68.7% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: SimpleQA Verified — Parametric knowledge · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — SimpleQA Verified — Parametric knowledge row, Gemini 3 Flash Thinking column · quote_snippet: SimpleQA Verified: parametric knowledge

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark simpleqa-verified (distinct from simpleqa). Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 72.1% / 2.5 Flash 28.1% / 2.5 Pro 54.5% / Claude Sonnet 4.5 29.3% / GPT-5.2 Extra high 38.0% / Grok 4.1 Fast 19.5%. Gemini 3 Flash best in row.

打开官方来源

mmmlu 91.8% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: MMMLU — Multilingual Q&A · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — MMMLU — Multilingual Q&A row, Gemini 3 Flash Thinking column · quote_snippet: MMMLU: multilingual Q&A

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mmmlu. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 91.8% (tied) / 2.5 Flash 86.6% / 2.5 Pro 89.5% / Claude Sonnet 4.5 89.1% / GPT-5.2 Extra high 89.6% / Grok 4.1 Fast 86.8%. Gemini 3 Flash tied best in row.

打开官方来源

global-piqa 92.8% 模型 gemini-3-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: Global PIQA — Commonsense reasoning across 100 Languages and Cultures · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — Global PIQA — Commonsense reasoning across 100 Languages and Cultures row, Gemini 3 Flash Thinking column · quote_snippet: Global PIQA: commonsense reasoning across 100 languages and cultures

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark global-piqa. Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 93.4% (best in row) / 2.5 Flash 90.2% / 2.5 Pro 91.5% / Claude Sonnet 4.5 90.1% / GPT-5.2 Extra high 91.2% / Grok 4.1 Fast 85.6%.

打开官方来源

mrcr 67.2% (v2 8-needle, 128k average) 模型 gemini-3-flash · 版本 v2 (8-needle), 128k (average) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: MRCR v2 (8-needle) — 128k (average) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — MRCR v2 (8-needle) — 128k (average) row, Gemini 3 Flash Thinking column · quote_snippet: MRCR v2 (8-needle): long context performance, 128k average

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": "cumulative score at 128k (average)",
  "judge": null
}

Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 77.0% / 2.5 Flash 54.3% / 2.5 Pro 58.0% / Claude Sonnet 4.5 47.1% / GPT-5.2 Extra high 81.9% (best in row) / Grok 4.1 Fast 54.6%.

打开官方来源

mrcr 22.1% (v2 8-needle, 1M pointwise) 模型 gemini-3-flash · 版本 v2 (8-needle), 1M (pointwise) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark table · row: MRCR v2 (8-needle) — 1M (pointwise) · figure: gemini-3-flash_final_benchmark-t.width-1200.format-webp.webp — MRCR v2 (8-needle) — 1M (pointwise) row, Gemini 3 Flash Thinking column · quote_snippet: MRCR v2 (8-needle): long context performance, 1M pointwise

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "run_count": null,
  "aggregation": "pointwise value at 1M context",
  "judge": null
}

Verified visually 2026-09-01 by reading the archived table image (models/2025-12-17-gemini-3-flash/images/03.webp): Gemini 3 Pro 26.3% / 2.5 Flash 21.0% / 2.5 Pro 16.4% / Grok 4.1 Fast 6.1%; Claude Sonnet 4.5 and GPT-5.2 cells read "not supported". Separate row from the 128k average because the aggregation differs.

打开官方来源