Gemini 2.5 Flash / Gemini 2.5 Flash-Lite
Google DeepMind / Gemini · 2025-06-17 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 2.5 Flash
Gemini 2.5 Flash 是 Gemini 2.5 家族 GA 扩容公告中的主力 Flash 档模型,默认开启思考并同时报告非思考配置。评测覆盖科学/数学/编码/事实性/视觉与多语言:思考配置 GPQA Diamond 82.8%、AIME 2025 72.0%、MRCR v2 8-needle(1M)21.0%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 0.3 / 输出 2.5 · 非思考与思考配置同价;表中标注为 no caching 定价
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Reasoning & knowledge / Humanity's Last Exam (no tools) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Reasoning & knowledge / Humanity's Last Exam (no tools) row, Gemini 2.5 Flash Thinking column · quote_snippet: Humanity's Last Exam (no tools)
{
"harness": "AI Studio API with default sampling settings",
"tools": [],
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt, no majority voting or parallel test-time compute",
"judge": null
}Score only in the GIF family table; visual reading: Flash Thinking 11.0% (Flash non-thinking 8.4%; Flash-Lite thinking 6.9%, non-thinking 5.1%; Pro thinking 21.6%). Page prose names no benchmark explicitly. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 5.1% / thinking 6.9%; Flash non-thinking 8.4%; Pro thinking 21.6%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Science / GPQA diamond · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Science / GPQA diamond row, Gemini 2.5 Flash Thinking column · quote_snippet: Science / GPQA diamond
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}Visual reading: Flash Thinking 82.8% (Flash non-thinking 78.3%; Flash-Lite thinking 66.7%, non-thinking 64.6%; Pro thinking 86.4%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 64.6% / thinking 66.7%; Flash non-thinking 78.3%; Pro thinking 86.4%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Mathematics / AIME 2025 · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Mathematics / AIME 2025 row, Gemini 2.5 Flash Thinking column · quote_snippet: Mathematics / AIME 2025
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}new-benchmark: aime-25 (registered batch 2, not yet in data/benchmarks/). Visual reading: Flash Thinking 72.0% (Flash non-thinking 61.6%; Flash-Lite thinking 63.1%, non-thinking 49.8%; Pro thinking 88.0%). Footnote (visual read): AIME 2025 numbers sourced from matharena.ai where provider numbers unavailable. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 49.8% / thinking 63.1%; Flash non-thinking 61.6%; Pro thinking 88.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Code generation / LiveCodeBench (v5: 1/1/2025-5/1/2025) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Code generation / LiveCodeBench (v5: 1/1/2025-5/1/2025) row, Gemini 2.5 Flash Thinking column · quote_snippet: LiveCodeBench (v5: 1/1/2025-5/1/2025)
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}Visual reading: Flash Thinking 55.4% (Flash non-thinking 41.1%; Flash-Lite thinking 34.3%, non-thinking 33.7%; Pro thinking 69.0%). Window differs from the 2.5 Pro March table's v5 and from Grok 3/Llama 4 windows — always compare by window variant. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 33.7% / thinking 34.3%; Flash non-thinking 41.1%; Pro thinking 69.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Code editing / Aider Polyglot · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Code editing / Aider Polyglot row, Gemini 2.5 Flash Thinking column · quote_snippet: Code editing / Aider Polyglot
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 3,
"aggregation": "pass rate average of 3 trials",
"judge": null
}Visual reading: Flash Thinking 56.7% (Flash non-thinking 44.0%; Flash-Lite thinking 27.1%, non-thinking 26.7%; Pro thinking 82.2%). Footnote explicitly states 3-trial averaging and that Aider results differ from the official leaderboard due to non-default evaluation settings. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 26.7% / thinking 27.1%; Flash non-thinking 44.0%; Pro thinking 82.2%. Table footnote: Aider Polyglot score is the pass-rate average of 3 trials.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Agentic coding / SWE-bench Verified — single attempt · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Agentic coding / SWE-bench Verified — single attempt row, Gemini 2.5 Flash Thinking column · quote_snippet: SWE-bench Verified — single attempt
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}Visual reading: Flash Thinking 48.9% (Flash non-thinking 50.0% — non-thinking scores higher here; Flash-Lite thinking 27.6%, non-thinking 31.6%; Pro thinking 59.6%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 31.6% / thinking 27.6%; Flash non-thinking 50.0% (non-thinking scores higher); Pro thinking 59.6%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Agentic coding / SWE-bench Verified — multiple attempts · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Agentic coding / SWE-bench Verified — multiple attempts row, Gemini 2.5 Flash Thinking column · quote_snippet: SWE-bench Verified — multiple attempts
{
"harness": "AI Studio API with default sampling settings; Google scaffolding draws multiple trajectories and re-scores with the model's own judgement",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "test-time selection of the candidate answer (multiple attempts)",
"judge": "model's own judgement (re-scoring)"
}Visual reading: Flash Thinking 60.3% (Flash non-thinking 60.0%; Flash-Lite thinking 44.9%, non-thinking 42.6%; Pro thinking 67.2%). Multiple-attempts numbers are NOT comparable to single-attempt rows from other vendors — aggregation differs by design. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 42.6% / thinking 44.9%; Flash non-thinking 60.0%; Pro thinking 67.2%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Factuality / SimpleQA · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Factuality / SimpleQA row, Gemini 2.5 Flash Thinking column · quote_snippet: Factuality / SimpleQA
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}Visual reading: Flash Thinking 26.9% (Flash non-thinking 25.8%; Flash-Lite thinking 13.0%, non-thinking 10.7%; Pro thinking 54.0%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 10.7% / thinking 13.0%; Flash non-thinking 25.8%; Pro thinking 54.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Factuality / FACTS Grounding · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Factuality / FACTS Grounding row, Gemini 2.5 Flash Thinking column · quote_snippet: Factuality / FACTS Grounding
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}Maps to existing benchmark factsg. Visual reading: Flash Thinking 85.3% (Flash non-thinking 83.4%; Flash-Lite thinking 86.8%, non-thinking 84.1%; Pro thinking 87.8%) — Flash-Lite thinking outscores Flash here, useful cost/quality decision point. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 84.1% / thinking 86.8% (Flash-Lite thinking is best in row); Flash non-thinking 83.4%; Pro thinking 87.8%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Visual reasoning / MMMU · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Visual reasoning / MMMU row, Gemini 2.5 Flash Thinking column · quote_snippet: Visual reasoning / MMMU
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}Visual reading: Flash Thinking 79.7% (Flash non-thinking 76.9%; Flash-Lite both configs 72.9%; Pro thinking 82.0%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 72.9% / thinking 72.9%; Flash non-thinking 76.9%; Pro thinking 82.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Image understanding / Vibe-Eval (Reka) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Image understanding / Vibe-Eval (Reka) row, Gemini 2.5 Flash Thinking column · quote_snippet: Image understanding / Vibe-Eval (Reka)
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": "Gemini as judge (per footnote)"
}new-benchmark: vibe-eval (introduced in this batch by gemini-2-0). Visual reading: Flash Thinking 65.4% (Flash non-thinking 66.2%; Flash-Lite thinking 57.5%, non-thinking 51.3%; Pro thinking 67.2%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 51.3% / thinking 57.5%; Flash non-thinking 66.2%; Pro thinking 67.2%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Long context / MRCR v2 (8-needle) — 128k (average) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Long context / MRCR v2 (8-needle) — 128k (average) row, Gemini 2.5 Flash Thinking column · quote_snippet: Long context — MRCR v2 (8-needle)
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "cumulative score at 128k (per footnote, chosen for comparability)",
"judge": null
}new-benchmark: mrcr (introduced in this batch). Visual reading: Flash Thinking 54.3% (Flash non-thinking 34.1%; Flash-Lite thinking 30.6%, non-thinking 16.6%; Pro thinking 58.0%). Footnote states MRCR v2 methodology changed vs previously published results and is not publicly available yet. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 16.6% / thinking 30.6%; Flash non-thinking 34.1%; Pro thinking 58.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Long context / MRCR v2 (8-needle) — 1M (pointwise) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Long context / MRCR v2 (8-needle) — 1M (pointwise) row, Gemini 2.5 Flash Thinking column · quote_snippet: MRCR v2 — 1M (pointwise)
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pointwise value at 1M context (per footnote)",
"judge": null
}Visual reading: Flash Thinking 21.0% (Flash non-thinking 16.8%; Flash-Lite thinking 5.4%, non-thinking 4.1%; Pro thinking 16.4% — Flash outperforms Pro at 1M pointwise). Separate row from 128k average because the footnote defines different aggregations. Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 4.1% / thinking 5.4%; Flash non-thinking 16.8%; Pro thinking 16.4% (Flash outperforms Pro at 1M pointwise).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Introducing Gemini 2.5 Flash-Lite (family benchmark table) · row: Multilingual performance / Global MMLU (Lite) · figure: gemini_2-5_benchmarks_margin_light2x_1.gif — Multilingual performance / Global MMLU (Lite) row, Gemini 2.5 Flash Thinking column · quote_snippet: Multilingual performance / Global MMLU (Lite)
{
"harness": "AI Studio API with default sampling settings",
"tools": null,
"shots": null,
"reasoning_effort": "thinking",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass@1 single attempt",
"judge": null
}new-benchmark: global-mmlu-lite (introduced in this batch by gemini-2-5-pro). Visual reading: Flash Thinking 88.4% (Flash non-thinking 85.8%; Flash-Lite thinking 84.5%, non-thinking 81.1%; Pro thinking 89.2%). Confirmed 2026-09-01 by reading the archived GIF (models/2025-06-17-gemini-2-5-flash/images/03.gif): sibling cells — Flash-Lite Preview 06-17 non-thinking 81.1% / thinking 84.5%; Flash non-thinking 85.8%; Pro thinking 89.2%.
Gemini 2.5 Flash-Lite
Gemini 2.5 Flash-Lite(Preview 06-17)定位为 2.5 家族中成本最优的一档,官方称其在编码、数学、科学、推理与多模态基准上全面优于 2.0 Flash-Lite。评测覆盖同套基准:思考配置 GPQA Diamond 66.7%、AIME 2025 63.1%、SWE-bench Verified 27.6%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 0.1 / 输出 0.4 · Preview 06-17 定价;非思考与思考配置同价,表中标注为 no caching
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。