← 模型目录

Gemini 3.6 Flash / Gemini 3.5 Flash-Lite / Gemini 3.5 Flash Cyber

Google DeepMind / Gemini · 2026-07-21 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 3.6 Flash

发布文将 Flash 系列定位为以效率、时延与可靠性支撑规模化 AI 智能体的主力款,3.6 Flash 在 3.5 Flash 基础上提升编码与知识工作且输出 token 减少 17%。4 项评测集中于代码、机器学习工程与计算机使用,亮点为 OSWorld-Verified 83.0% 与 MLE-Bench 63.9%。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 1.5 / 输出 7.5

本变体的评测证据

deepswe 49% 模型 gemini-3-6-flash · 版本 v1.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.6 Flash: More efficient and better quality than 3.5 Flash · quote_snippet: higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id deepswe (registered batch 1). 3.5 Flash baseline = 37%. The Gemini 3.6 model page (deepmind.google/models/gemini/flash/) labels this row DeepSWE v1.1 and the intro cites up-to-65% token reduction on Datacurve's DeepSWE.

打开官方来源

mlebench 63.9% 模型 gemini-3-6-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.6 Flash: More efficient and better quality than 3.5 Flash · quote_snippet: significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mlebench. 3.5 Flash baseline = 49.7%.

打开官方来源

osworld 83.0% 模型 gemini-3-6-flash · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.6 Flash: More efficient and better quality than 3.5 Flash · quote_snippet: improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%)

{
  "harness": null,
  "tools": [
    "computer use as built-in client-side tool"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark osworld, variant Verified. 3.5 Flash baseline = 78.4%. Page notes computer use is now a built-in client-side tool via Gemini API / Enterprise — tool availability is comparability-relevant.

打开官方来源

gdpval-aa 1421 Elo 模型 gemini-3-6-flash · 版本 v2 · 指标 knowledge_work_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.6 Flash: More efficient and better quality than 3.5 Flash · quote_snippet: It outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard-style Elo",
  "judge": null
}

Reuses candidate id gdpval-aa, variant v2 (matches grok-4-6 batch-1 row GDPVal-AA v2). 3.5 Flash baseline = 1349. Cross-vendor note: the Gemini 3.6 model page shows Grok 4.5 at 1535 and Claude Sonnet 5 at 1607 on the same metric — do not mix with the May GDPval-AA (non-v2-labeled) 1656 row.

打开官方来源

Gemini 3.5 Flash-Lite

发布文将其定位为 3.5 系列最快的模型,面向低时延与高吞吐场景(Artificial Analysis 实测 350 输出 token/s)。5 项评测显示对 3.1 Flash-Lite 的代际提升,亮点为 Terminal-Bench 2.1 54%(对比 31%)与 GDM-MRCR v2 72.2%。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 0.3 / 输出 2.5

本变体的评测证据

terminalbench 54% 模型 gemini-3-5-flash-lite · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.5 Flash-Lite: Built to scale agentic workflows · quote_snippet: a significant step up in coding and agentic tasks as seen in Terminal-Bench 2.1 (54% vs 31%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "across thinking levels (minimal/low/medium/high configurable)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Baseline = Gemini 3.1 Flash-Lite 31% (first official 3.1 Flash-Lite score in the library, published on this later post).

打开官方来源

mrcr 72.2% 模型 gemini-3-5-flash-lite · 版本 GDM-MRCR v2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.5 Flash-Lite: Built to scale agentic workflows · quote_snippet: long context as seen in GDM-MRCR v2 (72.2% vs. 60.1%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark-adjacent: reuses mrcr (introduced by google/gemini-2-0 in batch 4) with GDM-MRCR v2 variant label; baseline = 3.1 Flash-Lite 60.1%. Context length for the 72.2% figure not stated (model page uses 128k average).

打开官方来源

gdpval-aa 1140 Elo 模型 gemini-3-5-flash-lite · 版本 v2 · 指标 knowledge_work_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.5 Flash-Lite: Built to scale agentic workflows · quote_snippet: real-world task execution as seen in GDPval-AA v2 (1140 vs. 642)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard-style Elo",
  "judge": null
}

Baseline = 3.1 Flash-Lite 642. Same-release family comparison: Muse Glimmer-30B later reports GDPVal-AA v2 953 (small open model bracket).

打开官方来源

swebench-pro 54.2% 模型 gemini-3-5-flash-lite · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.5 Flash-Lite: Built to scale agentic workflows · quote_snippet: even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id swebench-pro (registered batch 2). Comparison target here is Gemini 3 Flash 49.6% (not 3.1 Flash-Lite) — the baseline named on the page for this row.

打开官方来源

osworld 74.0% 模型 gemini-3-5-flash-lite · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.5 Flash-Lite: Built to scale agentic workflows · quote_snippet: and OSWorld-Verified (74.0% vs. 65.1%)

{
  "harness": null,
  "tools": [
    "computer use as built-in tool"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Baseline = Gemini 3 Flash 65.1%.

打开官方来源

Gemini 3.5 Flash Cyber

发布文将其定位为面向网络安全特化的 3.5 Flash 变体,独家用于 CodeMender 限量试点。本次发布仅报告 1 项数值:CodeMender 多智能体内的 CyberGym 83.2%。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

cybergym 83.2% 模型 gemini-3-5-flash-cyber · 版本 within CodeMender, multi-agent · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 3.5 Flash Cyber in CodeMender: finding and fixing vulnerabilities efficiently · figure: gemini-3-5-flash-cyber__evals__c...webp (images/18) — CyberGym bar chart, Gemini 3.5 Flash Cyber in CodeMender (Max 5 model calls) bar · quote_snippet: 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym

{
  "harness": "CodeMender multi-agent setup, max 5 model calls",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cybergym not yet in data/benchmarks/ — distinct from batch-1 cybergym-adjacent ids (cybergym vs exploitgym/cybergym naming must be checked at migration; batch 1 already registered `cybergym`, verify before adding). Qualitative claim only, no printed score; harness (multi-agent CodeMender) is the comparability-critical field. Score transcribed 2026-09-01 from the archived chart: Gemini 3.5 Flash Cyber in CodeMender (max 5 model calls) 83.2%; competitor bars run in different agents — Mythos Preview in Anthropic agent 83.1% / GPT-5.6 Sol in OpenAI agent 83.6% / Mythos 5 in Anthropic agent 83.8% / GPT-5.5-Cyber in OpenAI agent 85.6% — so the "competitive at the frontier" claim spans non-identical harnesses.

打开官方来源