← 模型目录

Gemini 2.0 Flash / Project Mariner

Google DeepMind / Gemini · 2024-12-11 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 2.0 Flash

发布文称其为面向智能体时代(agentic era)的 Gemini 2.0 家族首个模型。13 项评测横跨知识、代码、数学、多模态与长上下文,亮点为 Natural2Code 92.9% 与 MATH 89.7%。

输入模态
文本 / 图像 / 音频 / 视频
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu-pro 76.4% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — General / MMLU-Pro row, Gemini 2.0 Flash Experimental column · quote_snippet: 2.0 Flash even outperforms 1.5 Pro on key benchmarks, at twice the speed

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Score lives only inside the GIF benchmark chart (2.0 Flash vs 1.5 Flash 002 vs 1.5 Pro 002). Prose only makes the relative claim quoted. Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 67.3% / 1.5 Pro 002 75.8%.

打开官方来源

natural2code 92.9% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Code / Natural2Code row, Gemini 2.0 Flash Experimental column · quote_snippet: Natural2Code: code generation across Python, Java, C++, JS, Go

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: natural2code not yet in data/benchmarks/. Google-internal held-out HumanEval-like set (not leaked on the web per chart description). Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 79.8% / 1.5 Pro 002 85.4%.

打开官方来源

bird-sql 56.9% 模型 gemini-2-0-flash · 版本 Dev · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Code / Bird-SQL (Dev) row, Gemini 2.0 Flash Experimental column · quote_snippet: Bird-SQL (Dev): converting natural language questions into executable SQL

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: bird-sql not yet in data/benchmarks/. Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 45.6% / 1.5 Pro 002 54.4%.

打开官方来源

lcb 35.1% 模型 gemini-2-0-flash · 版本 Code Generation 06/01/2024-10/05/2024 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Code / LiveCodeBench (Code Generation) row, Gemini 2.0 Flash Experimental column · quote_snippet: LiveCodeBench: code generation in Python, subset 06/01/2024 - 10/05/2024

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark lcb; window variant 06/01/2024-10/05/2024 — a different snapshot from the 10/01/2024-02/01/2025 window used by Grok 3 and Llama 4 releases; not directly comparable. Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 30.0% / 1.5 Pro 002 34.3%.

打开官方来源

factsg 83.6% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Factuality / FACTS Grounding row, Gemini 2.0 Flash Experimental column · quote_snippet: FACTS Grounding: factually correct responses given documents, held out internal dataset

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark factsg (FACTS Grounding). Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 82.9% / 1.5 Pro 002 80.0%.

打开官方来源

math 89.7% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Math / MATH row, Gemini 2.0 Flash Experimental column · quote_snippet: MATH: challenging math problems incl. algebra, geometry, pre-calculus

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: math (full Hendrycks MATH) not yet in data/benchmarks/ — distinct from existing math500 (MATH-500 subset); scores not interchangeable. Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 77.9% / 1.5 Pro 002 86.5%.

打开官方来源

hiddenmath 63.0% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Math / HiddenMath row, Gemini 2.0 Flash Experimental column · quote_snippet: HiddenMath: competition-level math, held out AIME/AMC-like dataset, not leaked

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hiddenmath not yet in data/benchmarks/ (Google-internal held-out set). Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 47.2% / 1.5 Pro 002 52.0%.

打开官方来源

gpqa 62.1% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Reasoning / GPQA (diamond) row, Gemini 2.0 Flash Experimental column · quote_snippet: GPQA (diamond): questions written by domain experts in biology, physics, chemistry

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark gpqa (canonical name already GPQA Diamond). Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 51.0% / 1.5 Pro 002 59.1%.

打开官方来源

mrcr 69.2% (1M context) 模型 gemini-2-0-flash · 版本 1M · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Long Context / MRCR (1M) row, Gemini 2.0 Flash Experimental column · quote_snippet: MRCR (1M): novel, diagnostic long-context understanding evaluation

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mrcr not yet in data/benchmarks/ (Google long-context eval; later releases report MRCR v2 8-needle — different version, not comparable). Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 71.9% / 1.5 Pro 002 82.6% (best in row).

打开官方来源

mmmu 70.7% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Image / MMMU row, Gemini 2.0 Flash Experimental column · quote_snippet: MMMU: multi-discipline college-level multimodal understanding and reasoning

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mmmu. Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 62.3% / 1.5 Pro 002 65.9%.

打开官方来源

vibe-eval 56.3% 模型 gemini-2-0-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Image / Vibe-Eval (Reka) row, Gemini 2.0 Flash Experimental column · quote_snippet: Vibe-Eval (Reka): visual understanding in chat models, Gemini Flash model as rater

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "Gemini Flash model as rater (LLM judge)"
}

new-benchmark: vibe-eval not yet in data/benchmarks/. Judge explicitly stated in the chart description (self-preference risk: Google model rated by Google model). Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 48.9% / 1.5 Pro 002 53.9%.

打开官方来源

covost2 39.2 BLEU 模型 gemini-2-0-flash · 版本 21 languages · 指标 未说明 · 单位 bleu 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Audio / CoVoST2 (21 lang) row, Gemini 2.0 Flash Experimental column · quote_snippet: CoVoST2 (21 lang): automatic speech translation, BLEU score

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: covost2 not yet in data/benchmarks/. Unit is BLEU (not percent) per chart footnote. Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 37.4 / 1.5 Pro 002 40.1 (best in row).

打开官方来源

egoschema 71.5% 模型 gemini-2-0-flash · 版本 test · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Gemini 2.0 Flash · figure: gemini_benchmarks_narrow_light2x.gif — Video / EgoSchema (test) row, Gemini 2.0 Flash Experimental column · quote_snippet: EgoSchema (test): video analysis across multiple domains

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: egoschema not yet in data/benchmarks/ (video understanding; also appears in xAI Grok 3 table the same quarter). Confirmed 2026-09-01 by reading the archived GIF frame (models/2024-12-11-gemini-2-0/images/03.gif): value transcribed visually, competitor cells 1.5 Flash 002 66.8% / 1.5 Pro 002 71.2%.

打开官方来源

Project Mariner

发布文将 Project Mariner 定位为基于 Gemini 2.0 构建的浏览器操作研究原型,能理解浏览器页面并代用户完成多步任务。本次发布为其报告的唯一数值是 WebVoyager 83.5%。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

webvoyager 83.5% 模型 project-mariner · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Project Mariner: agents that can help you accomplish complex tasks · quote_snippet: Project Mariner achieved a state-of-the-art result of 83.5% working as a single agent setup

{
  "harness": "single agent setup (experimental Chrome extension)",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: webvoyager not yet in data/benchmarks/. Only prose-scoreable row on the page: the score belongs to Project Mariner (agent prototype built with Gemini 2.0), not to the Gemini 2.0 Flash model itself — model_id reflects that. The page text explicitly defines the benchmark as testing agent performance on end-to-end real world web tasks.

打开官方来源