← 模型目录

Gemini 2.5 Flash / Gemini 2.5 Flash-Lite

Google DeepMind / Gemini · 2025-06-17 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 2.5 Flash

Gemini 2.5 Flash 是 Gemini 2.5 家族 GA 扩容公告中的主力 Flash 档模型,默认开启思考并同时报告非思考配置。评测覆盖科学/数学/编码/事实性/视觉与多语言:思考配置 GPQA Diamond 82.8%、AIME 2025 72.0%、MRCR v2 8-needle(1M)21.0%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 0.3 / 输出 2.5 · 非思考与思考配置同价;表中标注为 no caching 定价

本变体的评测证据

hlehle 11.0% (no tools, thinking) 模型 gemini-2-5-flash · 版本 no tools; thinking config · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Reasoning & knowledge / Humanity's Last Exam (no tools) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Reasoning & knowledge / Humanity's Last Exam (no tools) row, Gemini 2.5 Flash Thinking column · quote_snippet: Humanity's Last Exam (no tools)

{
  "harness": "AI Studio API with default sampling settings",
  "tools": [],
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt, no majority voting or parallel test-time compute",
  "judge": null
}

Score only in the GIF family table; visual reading: Flash Thinking 11.0% (Flash non-thinking 8.4%; Flash-Lite thinking 6.9%, non-thinking 5.1%; Pro thinking 21.6%). Page prose names no benchmark explicitly. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 5.1% / thinking 6.9%; Flash non-thinking 8.4%; Pro thinking 21.6%.

打开官方来源

gpqa 82.8% (thinking) 模型 gemini-2-5-flash · 版本 thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Science / GPQA diamond · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Science / GPQA diamond row, Gemini 2.5 Flash Thinking column · quote_snippet: Science / GPQA diamond

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

Visual reading: Flash Thinking 82.8% (Flash non-thinking 78.3%; Flash-Lite thinking 66.7%, non-thinking 64.6%; Pro thinking 86.4%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 64.6% / thinking 66.7%; Flash non-thinking 78.3%; Pro thinking 86.4%.

打开官方来源

aime-25 72.0% (thinking) 模型 gemini-2-5-flash · 版本 thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Mathematics / AIME 2025 · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Mathematics / AIME 2025 row, Gemini 2.5 Flash Thinking column · quote_snippet: Mathematics / AIME 2025

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

new-benchmark: aime-25 (registered batch 2, not yet in data/benchmarks/). Visual reading: Flash Thinking 72.0% (Flash non-thinking 61.6%; Flash-Lite thinking 63.1%, non-thinking 49.8%; Pro thinking 88.0%). Footnote (visual read): AIME 2025 numbers sourced from matharena.ai where provider numbers unavailable. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 49.8% / thinking 63.1%; Flash non-thinking 61.6%; Pro thinking 88.0%.

打开官方来源

lcb 55.4% (v5 1/1/2025-5/1/2025, thinking) 模型 gemini-2-5-flash · 版本 v5 1/1/2025-5/1/2025; thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Code generation / LiveCodeBench (v5: 1/1/2025-5/1/2025) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Code generation / LiveCodeBench (v5: 1/1/2025-5/1/2025) row, Gemini 2.5 Flash Thinking column · quote_snippet: LiveCodeBench (v5: 1/1/2025-5/1/2025)

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

Visual reading: Flash Thinking 55.4% (Flash non-thinking 41.1%; Flash-Lite thinking 34.3%, non-thinking 33.7%; Pro thinking 69.0%). Window differs from the 2.5 Pro March table's v5 and from Grok 3/Llama 4 windows — always compare by window variant. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 33.7% / thinking 34.3%; Flash non-thinking 41.1%; Pro thinking 69.0%.

打开官方来源

aider 56.7% (Polyglot, thinking) 模型 gemini-2-5-flash · 版本 Polyglot; thinking config · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Code editing / Aider Polyglot · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Code editing / Aider Polyglot row, Gemini 2.5 Flash Thinking column · quote_snippet: Code editing / Aider Polyglot

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 3,
  "aggregation": "pass rate average of 3 trials",
  "judge": null
}

Visual reading: Flash Thinking 56.7% (Flash non-thinking 44.0%; Flash-Lite thinking 27.1%, non-thinking 26.7%; Pro thinking 82.2%). Footnote explicitly states 3-trial averaging and that Aider results differ from the official leaderboard due to non-default evaluation settings. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 26.7% / thinking 27.1%; Flash non-thinking 44.0%; Pro thinking 82.2%. Table footnote: Aider Polyglot score is the pass-rate average of 3 trials.

打开官方来源

swebench 48.9% (Verified, single attempt, thinking) 模型 gemini-2-5-flash · 版本 Verified, single attempt; thinking config · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Agentic coding / SWE-bench Verified — single attempt · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Agentic coding / SWE-bench Verified — single attempt row, Gemini 2.5 Flash Thinking column · quote_snippet: SWE-bench Verified — single attempt

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

Visual reading: Flash Thinking 48.9% (Flash non-thinking 50.0% — non-thinking scores higher here; Flash-Lite thinking 27.6%, non-thinking 31.6%; Pro thinking 59.6%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 31.6% / thinking 27.6%; Flash non-thinking 50.0% (non-thinking scores higher); Pro thinking 59.6%.

打开官方来源

swebench 60.3% (Verified, multiple attempts, thinking) 模型 gemini-2-5-flash · 版本 Verified, multiple attempts; thinking config · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Agentic coding / SWE-bench Verified — multiple attempts · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Agentic coding / SWE-bench Verified — multiple attempts row, Gemini 2.5 Flash Thinking column · quote_snippet: SWE-bench Verified — multiple attempts

{
  "harness": "AI Studio API with default sampling settings; Google scaffolding draws multiple trajectories and re-scores with the model's own judgement",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "test-time selection of the candidate answer (multiple attempts)",
  "judge": "model's own judgement (re-scoring)"
}

Visual reading: Flash Thinking 60.3% (Flash non-thinking 60.0%; Flash-Lite thinking 44.9%, non-thinking 42.6%; Pro thinking 67.2%). Multiple-attempts numbers are NOT comparable to single-attempt rows from other vendors — aggregation differs by design. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 42.6% / thinking 44.9%; Flash non-thinking 60.0%; Pro thinking 67.2%.

打开官方来源

simpleqa 26.9% (thinking) 模型 gemini-2-5-flash · 版本 thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Factuality / SimpleQA · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Factuality / SimpleQA row, Gemini 2.5 Flash Thinking column · quote_snippet: Factuality / SimpleQA

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

Visual reading: Flash Thinking 26.9% (Flash non-thinking 25.8%; Flash-Lite thinking 13.0%, non-thinking 10.7%; Pro thinking 54.0%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 10.7% / thinking 13.0%; Flash non-thinking 25.8%; Pro thinking 54.0%.

打开官方来源

factsg 85.3% (thinking) 模型 gemini-2-5-flash · 版本 thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Factuality / FACTS Grounding · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Factuality / FACTS Grounding row, Gemini 2.5 Flash Thinking column · quote_snippet: Factuality / FACTS Grounding

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

Maps to existing benchmark factsg. Visual reading: Flash Thinking 85.3% (Flash non-thinking 83.4%; Flash-Lite thinking 86.8%, non-thinking 84.1%; Pro thinking 87.8%) — Flash-Lite thinking outscores Flash here, useful cost/quality decision point. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 84.1% / thinking 86.8% (Flash-Lite thinking is best in row); Flash non-thinking 83.4%; Pro thinking 87.8%.

打开官方来源

mmmu 79.7% (thinking) 模型 gemini-2-5-flash · 版本 thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Visual reasoning / MMMU · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Visual reasoning / MMMU row, Gemini 2.5 Flash Thinking column · quote_snippet: Visual reasoning / MMMU

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

Visual reading: Flash Thinking 79.7% (Flash non-thinking 76.9%; Flash-Lite both configs 72.9%; Pro thinking 82.0%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 72.9% / thinking 72.9%; Flash non-thinking 76.9%; Pro thinking 82.0%.

打开官方来源

vibe-eval 65.4% (thinking) 模型 gemini-2-5-flash · 版本 thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Image understanding / Vibe-Eval (Reka) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Image understanding / Vibe-Eval (Reka) row, Gemini 2.5 Flash Thinking column · quote_snippet: Image understanding / Vibe-Eval (Reka)

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": "Gemini as judge (per footnote)"
}

new-benchmark: vibe-eval (introduced in this batch by gemini-2-0). Visual reading: Flash Thinking 65.4% (Flash non-thinking 66.2%; Flash-Lite thinking 57.5%, non-thinking 51.3%; Pro thinking 67.2%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 51.3% / thinking 57.5%; Flash non-thinking 66.2%; Pro thinking 67.2%.

打开官方来源

mrcr 54.3% (v2 8-needle, 128k average, thinking) 模型 gemini-2-5-flash · 版本 v2 8-needle, 128k average; thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Long context / MRCR v2 (8-needle) — 128k (average) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Long context / MRCR v2 (8-needle) — 128k (average) row, Gemini 2.5 Flash Thinking column · quote_snippet: Long context — MRCR v2 (8-needle)

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "cumulative score at 128k (per footnote, chosen for comparability)",
  "judge": null
}

new-benchmark: mrcr (introduced in this batch). Visual reading: Flash Thinking 54.3% (Flash non-thinking 34.1%; Flash-Lite thinking 30.6%, non-thinking 16.6%; Pro thinking 58.0%). Footnote states MRCR v2 methodology changed vs previously published results and is not publicly available yet. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 16.6% / thinking 30.6%; Flash non-thinking 34.1%; Pro thinking 58.0%.

打开官方来源

mrcr 21.0% (v2 8-needle, 1M pointwise, thinking) 模型 gemini-2-5-flash · 版本 v2 8-needle, 1M pointwise; thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Long context / MRCR v2 (8-needle) — 1M (pointwise) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Long context / MRCR v2 (8-needle) — 1M (pointwise) row, Gemini 2.5 Flash Thinking column · quote_snippet: MRCR v2 — 1M (pointwise)

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pointwise value at 1M context (per footnote)",
  "judge": null
}

Visual reading: Flash Thinking 21.0% (Flash non-thinking 16.8%; Flash-Lite thinking 5.4%, non-thinking 4.1%; Pro thinking 16.4% — Flash outperforms Pro at 1M pointwise). Separate row from 128k average because the footnote defines different aggregations. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 4.1% / thinking 5.4%; Flash non-thinking 16.8%; Pro thinking 16.4% (Flash outperforms Pro at 1M pointwise).

打开官方来源

global-mmlu-lite 88.4% (thinking) 模型 gemini-2-5-flash · 版本 thinking config · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Multilingual performance / Global MMLU (Lite) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Multilingual performance / Global MMLU (Lite) row, Gemini 2.5 Flash Thinking column · quote_snippet: Multilingual performance / Global MMLU (Lite)

{
  "harness": "AI Studio API with default sampling settings",
  "tools": null,
  "shots": null,
  "reasoning_effort": "thinking",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1 single attempt",
  "judge": null
}

new-benchmark: global-mmlu-lite (introduced in this batch by gemini-2-5-pro). Visual reading: Flash Thinking 88.4% (Flash non-thinking 85.8%; Flash-Lite thinking 84.5%, non-thinking 81.1%; Pro thinking 89.2%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 81.1% / thinking 84.5%; Flash non-thinking 85.8%; Pro thinking 89.2%.

打开官方来源

Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite(Preview 06-17)定位为 2.5 家族中成本最优的一档,官方称其在编码、数学、科学、推理与多模态基准上全面优于 2.0 Flash-Lite。评测覆盖同套基准:思考配置 GPQA Diamond 66.7%、AIME 2025 63.1%、SWE-bench Verified 27.6%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 0.1 / 输出 0.4 · Preview 06-17 定价;非思考与思考配置同价,表中标注为 no caching

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。