Gemini 3.7 Flash
Google DeepMind / Gemini · 2026-08-13 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 3.7 Flash
Gemini 3.7 Flash 被定位为面向编码与代理的旗舰 workhorse 模型(most intelligent workhorse model for coding and agents)。评测聚焦编码与代理工作流并延伸到文档/视频理解:DeepSWE v1.1 65.3%、Terminal-bench 2.1 85.8%、WebDev Arena 1588 Elo、GDP.pdf 34.0%。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 0.75 / 输出 3.75 · 促销价(原 3.6 Flash 半价),2026-12-31 到期;2027-01-01 起 $1.50/$7.50
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Better intelligence for complex workflows · quote_snippet: generating production-ready code as seen in FrontierCode 1.1 Main (43.6% vs 34.4%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id frontier-code (registered batch 3; page label 'FrontierCode 1.1 Main'). 3.6 Flash baseline = 34.4%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Better intelligence for complex workflows · quote_snippet: DeepSWE v1.1 (65.3% vs 49.0%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}3.6 Flash baseline = 49.0% (matches the 3.6 release row). Same benchmark family as Grok 4.5's DeepSWE 1.0/1.1 chart rows and Muse Spark 1.2's 59.3% — harness/runner differs per vendor (Datacurve eval, provider harnesses), variant field must align before comparing.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Better intelligence for complex workflows · quote_snippet: It outperforms 3.6 Flash on Arena.ai's WebDev Arena with an Elo score of 1588 vs 1538
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard",
"judge": null
}Reuses candidate id webdev-arena (registered batch 2). 3.6 Flash baseline = 1538.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Better intelligence for complex workflows · quote_snippet: significantly outperforms 3.6 Flash on the GDP.pdf benchmark (34.0% vs 22.0%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gdp-pdf not yet in data/benchmarks/ — the PDF-document comprehension edition of GDPVal; distinct from gdpval-aa (Elo edition). 3.6 Flash baseline = 22.0%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Better intelligence for complex workflows · quote_snippet: surpasses 3.6 Flash in AutomationBench... (30.4% vs 17.0%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id automationbench (registered batch 1). 3.6 Flash baseline = 17.0%. Model card marks this set private — cross-vendor comparison impossible without set access.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: Artificial Analysis Intelligence Index — Composite model intelligence · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — Artificial Analysis Intelligence Index — Composite model intelligence row, Gemini 3.7 Flash column · quote_snippet: Artificial Analysis Intelligence Index: composite model intelligence
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Third-party index (Artificial Analysis) republished on Google's release page: attribution third_party_reported. Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 52 / Claude Sonnet 5 55 / GPT-5.6 Terra 57 / Muse Spark 1.2 57 (tied best in row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: Terminal-bench 2.1 — Agentic terminal coding · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — Terminal-bench 2.1 — Agentic terminal coding row, Gemini 3.7 Flash column · quote_snippet: Terminal-bench 2.1: agentic terminal coding
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 78.0% / Claude Sonnet 5 80.4% / GPT-5.6 Terra 87.4% (best in row) / Muse Spark 1.2 82.9%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: Terminal-bench 3.0 — General agent capabilities · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — Terminal-bench 3.0 — General agent capabilities row, Gemini 3.7 Flash column · quote_snippet: Terminal-bench 3.0: general agent capabilities
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 5.4% / Claude Sonnet 5 14.6% / GPT-5.6 Terra 20.8% (best in row); Muse Spark 1.2 em-dash. Distinct benchmark generation from Terminal-bench 2.1 — not comparable.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: GDPval-AA v2 — Knowledge work (Elo) · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — GDPval-AA v2 — Knowledge work (Elo) row, Gemini 3.7 Flash column · quote_snippet: GDPval-AA v2: knowledge work, Elo
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 1422 / Claude Sonnet 5 1598 / GPT-5.6 Terra 1578 / Muse Spark 1.2 1628 (best in row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: Harvey LAB-AA — Complex legal workflows · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — Harvey LAB-AA — Complex legal workflows row, Gemini 3.7 Flash column · quote_snippet: Harvey LAB-AA: complex legal workflows
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 85.1% / Claude Sonnet 5 90.1% / GPT-5.6 Terra 85.2%; Muse Spark 1.2 em-dash. Gemini 3.7 Flash best in row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: CharXiv Reasoning — Information synthesis from complex charts, no tools · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — CharXiv Reasoning — Information synthesis from complex charts, no tools row, Gemini 3.7 Flash column · quote_snippet: CharXiv Reasoning: no tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 85.2% / Claude Sonnet 5 77.0% / GPT-5.6 Terra 85.9% (best in row); Muse Spark 1.2 em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: CharXiv Reasoning — Information synthesis from complex charts, with tools · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — CharXiv Reasoning — Information synthesis from complex charts, with tools row, Gemini 3.7 Flash column · quote_snippet: CharXiv Reasoning: with tools
{
"harness": null,
"tools": [
"chart tools"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 89.4% (best in row) / Claude Sonnet 5 88.3%; GPT-5.6 Terra and Muse Spark 1.2 em-dash. Distinct tool condition from the no-tools row — not comparable.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: LVBench — Long video understanding · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — LVBench — Long video understanding row, Gemini 3.7 Flash column · quote_snippet: LVBench: long video understanding
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 84.2% / Claude Sonnet 5 68.5% / GPT-5.6 Terra 78.9%; Muse Spark 1.2 em-dash. Gemini 3.7 Flash best in row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: GDM-MRCR v2 (8-needle) — 128k (average) · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — GDM-MRCR v2 (8-needle) — 128k (average) row, Gemini 3.7 Flash column · quote_snippet: GDM-MRCR v2 (8-needle): long context performance, 128k average
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 91.8% / Claude Sonnet 5 81.5% / GPT-5.6 Terra 93.5%; Muse Spark 1.2 em-dash. Gemini 3.7 Flash best in row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: OSWorld-2.0 — Agentic computer use · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — OSWorld-2.0 — Agentic computer use row, Gemini 3.7 Flash column · quote_snippet: OSWorld-2.0: agentic computer use
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 33.8% / GPT-5.6 Terra 50.2% (best in row); Claude Sonnet 5 and Muse Spark 1.2 em-dash. Distinct from OSWorld-Verified lane used by earlier releases.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: Agent's Last Exam — Multimodal desktop and OS agent tasks, pass rate · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — Agent's Last Exam — Multimodal desktop and OS agent tasks, pass rate row, Gemini 3.7 Flash column · quote_snippet: Agent's Last Exam: multimodal desktop/OS agent tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 24.2% / Claude Sonnet 5 33.3% (best in row) / GPT-5.6 Terra 28.0%; Muse Spark 1.2 em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: HLE-Verified — Multidisciplinary expert reasoning · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — HLE-Verified — Multidisciplinary expert reasoning row, Gemini 3.7 Flash column · quote_snippet: HLE-Verified: multidisciplinary expert reasoning
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 51.2% / Claude Sonnet 5 31.0% / GPT-5.6 Terra 51.1%; Muse Spark 1.2 em-dash. Gemini 3.7 Flash best in row. Distinct lane from the full-set HLE rows of earlier releases (Verified harness).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: BioMysteryBench — Bioinformatics research reasoning, human solvable · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — BioMysteryBench — Bioinformatics research reasoning, human solvable row, Gemini 3.7 Flash column · quote_snippet: BioMysteryBench: human solvable split
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 80.6% / Claude Sonnet 5 87.5% (best in row) / GPT-5.6 Terra 83.8%; Muse Spark 1.2 em-dash.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: BioMysteryBench — Bioinformatics research reasoning, human difficult · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — BioMysteryBench — Bioinformatics research reasoning, human difficult row, Gemini 3.7 Flash column · quote_snippet: BioMysteryBench: human difficult split
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 41.2% / Claude Sonnet 5 34.1% / GPT-5.6 Terra 49.4% (best in row); Muse Spark 1.2 em-dash. Distinct split from the human-solvable row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmark table · row: LABBench2 — Biology real-world research tasks · figure: gemini-3-7-flash__evals__benchma.width-1200.format-webp.webp — LABBench2 — Biology real-world research tasks row, Gemini 3.7 Flash column · quote_snippet: LABBench2: biology real-world research tasks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: labbench2 not yet in data/benchmarks/. Verified visually 2026-09-01 by reading the archived table image (models/2026-08-13-gemini-3-7-flash/images/17.webp): Gemini 3.6 Flash 76.1% / Claude Sonnet 5 80.1% / GPT-5.6 Terra 81.2%; Muse Spark 1.2 em-dash. Gemini 3.7 Flash best in row.