Gemini 3.5 Flash
Google DeepMind / Gemini · 2026-05-19 · 通用模型
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 3.5 Flash
Google 在 I/O 2026 发布 Gemini 3.5 首个模型 Gemini 3.5 Flash,定位「frontier intelligence with action」,并成为 Gemini 应用与 Search AI Mode 的默认模型。评测横跨终端智能体、多模态理解与抽象推理,亮点如 Terminal-Bench 2.1 76.2%、MCP Atlas 83.6%,页面并称输出速度约为其他前沿模型 4 倍。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding · quote_snippet: outperforming Gemini 3.1 Pro on challenging coding and agentic benchmarks like Terminal-Bench 2.1 (76.2%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark terminalbench, variant 2.1. Comparison baseline named as Gemini 3.1 Pro (that 3.1 release has no library entry yet — inventory gap flagged in README).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding · quote_snippet: GDPval-AA (1656 Elo)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard-style Elo",
"judge": null
}Reuses candidate id gdpval-aa (registered batch 1). Page writes 'GDPval-AA' without the v2 suffix used on later Google pages (3.6/3.7 write GDPVal-AA v2) — variant left null here; do not assume same edition as v2 rows without checking.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding · quote_snippet: MCP Atlas (83.6%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id mcp-atlas (registered batch 1). Muse Glimmer model card later reports 'MCP Atlas (Public)' 75.5 — the Public-subset qualifier matters for comparability.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding · quote_snippet: leading in multimodal understanding (84.2% on CharXiv Reasoning)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id charxiv-reasoning (registered batch 2). The Gemini 3.6 model page later splits CharXiv into no-tools 84.2 / with-tools 84.9 for 3.5 Flash — this row's condition (likely no-tools) is not stated on this page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: SWE-Bench Pro (Public) — Diverse agentic coding tasks, single attempt · figure: gemini-3-5__benchmarks__light.gif — SWE-Bench Pro (Public) — Diverse agentic coding tasks, single attempt row, Gemini 3.5 Flash column · quote_snippet: SWE-Bench Pro (Public): single attempt
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 49.6% / Gemini 3.1 Pro 54.2% / Claude Opus 4.7 64.3% (best in row) / GPT-5.5 58.6%; Sonnet 4.6 em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: Toolathlon — Real-world general tool use · figure: gemini-3-5__benchmarks__light.gif — Toolathlon — Real-world general tool use row, Gemini 3.5 Flash column · quote_snippet: Toolathlon: real-world general tool use
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 49.4% / GPT-5.5 55.6%; 3.1 Pro, Sonnet 4.6 and Opus 4.7 em-dash. Gemini 3.5 Flash best in row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: OSWorld-Verified — Agentic computer use · figure: gemini-3-5__benchmarks__light.gif — OSWorld-Verified — Agentic computer use row, Gemini 3.5 Flash column · quote_snippet: OSWorld-Verified: agentic computer use
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 65.1% / Gemini 3.1 Pro 76.2% / Claude Sonnet 4.6 72.5% / Claude Opus 4.7 78.0% / GPT-5.5 78.7% (best in row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: Finance Agent v2 — Financial analysis and decision-making · figure: gemini-3-5__benchmarks__light.gif — Finance Agent v2 — Financial analysis and decision-making row, Gemini 3.5 Flash column · quote_snippet: Finance Agent v2: financial analysis and decision-making
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 42.6% / Gemini 3.1 Pro 43.0% / Claude Sonnet 4.6 51.0% / Claude Opus 4.7 51.5% / GPT-5.5 51.8%. Gemini 3.5 Flash best in row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: MMMU-Pro — Multimodal understanding and reasoning, no tools · figure: gemini-3-5__benchmarks__light.gif — MMMU-Pro — Multimodal understanding and reasoning, no tools row, Gemini 3.5 Flash column · quote_snippet: MMMU-Pro: multimodal understanding and reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 81.2% / Gemini 3.1 Pro 80.5% / Claude Sonnet 4.6 74.5% / Claude Opus 4.7 75.2% / GPT-5.5 81.2%. Gemini 3.5 Flash best in row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: Blueprint-Bench 2 — Agentic spatial reasoning, normalized score · figure: gemini-3-5__benchmarks__light.gif — Blueprint-Bench 2 — Agentic spatial reasoning, normalized score row, Gemini 3.5 Flash column · quote_snippet: Blueprint-Bench 2: agentic spatial reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 0.0% / Gemini 3.1 Pro 26.5% / Claude Sonnet 4.6 6.7% / Claude Opus 4.7 24.5% / GPT-5.5 36.2% (best in row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: MRCR v2 (8-needle) — 128k (average) · figure: gemini-3-5__benchmarks__light.gif — MRCR v2 (8-needle) — 128k (average) row, Gemini 3.5 Flash column · quote_snippet: MRCR v2 (8-needle): long context performance, 128k average
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 67.2% / Gemini 3.1 Pro 84.9% / Claude Sonnet 4.6 84.9% / Claude Opus 4.7 59.3% / GPT-5.5 94.8% (best in row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: MRCR v2 (8-needle) — 1M (pointwise) · figure: gemini-3-5__benchmarks__light.gif — MRCR v2 (8-needle) — 1M (pointwise) row, Gemini 3.5 Flash column · quote_snippet: MRCR v2 (8-needle): long context performance, 1M pointwise
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 22.1% / Gemini 3.1 Pro 26.3%; Sonnet 4.6, Opus 4.7 and GPT-5.5 em-dash (no 1M context). Separate row from the 128k average because the aggregation differs.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: Humanity's Last Exam — Academic reasoning, full set (text + MM) · figure: gemini-3-5__benchmarks__light.gif — Humanity's Last Exam — Academic reasoning, full set (text + MM) row, Gemini 3.5 Flash column · quote_snippet: Humanity's Last Exam: academic reasoning, full set
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 33.7% / Gemini 3.1 Pro 44.4% / Claude Sonnet 4.6 33.2% / Claude Opus 4.7 46.9% (best in row) / GPT-5.5 41.4%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 3.5 Flash: frontier performance for agents and coding (benchmark table) · row: ARC-AGI-2 — Abstract reasoning puzzles · figure: gemini-3-5__benchmarks__light.gif — ARC-AGI-2 — Abstract reasoning puzzles row, Gemini 3.5 Flash column · quote_snippet: ARC-AGI-2: abstract reasoning puzzles
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": "ARC Prize Verified"
}Verified visually 2026-09-01 by reading the archived GIF (models/2026-05-19-gemini-3-5-flash/images/03.gif): Gemini 3 Flash 33.6% / Gemini 3.1 Pro 77.1% / Claude Sonnet 4.6 58.3% / Claude Opus 4.7 75.8% / GPT-5.5 84.6% (best in row).