← 模型目录

Gemini 3.1 Pro

Google DeepMind / Gemini · 2026-02-19 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 3.1 Pro

Google 将 Gemini 3.1 Pro 定位为「为最复杂任务打造的更聪明模型」,基于 Gemini 3 Pro,发布时为 preview(thinking 档 low/medium/high)。已收录 19 项评测横跨推理、代码、多模态与长上下文:亮点 GPQA 94.3%、SWE-bench Verified 80.6%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 44.4% 模型 gemini-3-1-pro · 版本 full set (text + MM), no tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text), Humanity's Last Exam row · row: Humanity's Last Exam — Academic reasoning (full set, text + MM) / No tools · quote_snippet: Humanity's Last Exam... No tools | 44.4% | 37.5% | 33.2% | 40.0% | 34.5% | —

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark hlehle. Full-set text+MM edition with the no-tools condition isolated as its own row (competitors: Gemini 3 Pro 37.5 / Sonnet 4.6 33.2 / Opus 4.6 40.0 / GPT-5.2 34.5 — table columns, kept in notes per batch-1 口径). Companion row with Search+Code tools recorded separately.

打开官方来源

hlehle 51.4% 模型 gemini-3-1-pro · 版本 full set (text + MM), Search (blocklist) + Code · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text), Humanity's Last Exam row, tool-augmented condition · row: Humanity's Last Exam / Search (blocklist) + Code · quote_snippet: Search (blocklist) + Code | 51.4% | 45.8% | 49.0% | 53.1% | 45.5% | —

{
  "harness": null,
  "tools": [
    "Search (blocklist)",
    "Code execution"
  ],
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-augmented lane of the same HLE row (Opus 4.6 leads at 53.1). NOT comparable to the no-tools row or to other vendors' tool-free HLE numbers.

打开官方来源

arc-agi 77.1% 模型 gemini-3-1-pro · 版本 2, ARC Prize Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: ARC-AGI-2 — ARC Prize Verified · quote_snippet: ARC-AGI-2 | ARC Prize Verified | 77.1% | 31.1% | 58.3% | 68.8% | 52.9% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": "ARC Prize Verified"
}

Maps to existing benchmark arc-agi variant 2. The release blog's headline number ('verified score of 77.1%... more than double 3 Pro'). Third-party-verified lane — distinct from Gemini 3 Deep Think's 45.1% with-code-execution row (different tool condition).

打开官方来源

gpqa 94.3% 模型 gemini-3-1-pro · 版本 no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: GPQA Diamond — Scientific knowledge / No tools · quote_snippet: GPQA Diamond | No tools | 94.3% | 91.9% | 89.9% | 91.3% | 92.4% | —

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark gpqa (canonical Diamond).

打开官方来源

terminalbench 68.5% 模型 gemini-3-1-pro · 版本 2.0, Terminus-2 harness · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: Terminal-Bench 2.0 — Agentic terminal coding / Terminus-2 harness · quote_snippet: Terminal-Bench 2.0 | Terminus-2 harness | 68.5% | 56.9% | 59.1% | 65.4% | 54.0% | 64.7%

{
  "harness": "Terminus-2",
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark terminalbench variant 2.0. Same-harness lane (Terminus-2) as Muse Glimmer's TB 2.1 row — but different TB VERSION (2.0 vs 2.1); still not interchangeable. Table also carries an 'Other best self-reported harness' sub-row where GPT-5.3-Codex reaches 77.3 (Codex harness) — GPT-5.2 62.2 (Codex); Gemini columns are '—' there, so it is recorded here rather than as a comparison_cited row.

打开官方来源

swebench 80.6% 模型 gemini-3-1-pro · 版本 Verified, single attempt · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: SWE-Bench Verified — Agentic coding / Single attempt · quote_snippet: SWE-Bench Verified | Single attempt | 80.6% | 76.2% | 79.6% | 80.8% | 80.0% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "single attempt (pass@1 lane)",
  "judge": null
}

Maps to existing benchmark swebench. Opus 4.6 edges this row at 80.8 (table column, in notes). Single-attempt condition explicit — distinct from Google's 'multiple attempts' scaffolding lane used on other pages.

打开官方来源

swebench-pro 54.2% 模型 gemini-3-1-pro · 版本 Public, single attempt · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: SWE-Bench Pro (Public) — Diverse agentic coding tasks / Single attempt · quote_snippet: SWE-Bench Pro (Public) | Single attempt | 54.2% | 43.3% | — | — | 55.6% | 56.8%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "single attempt",
  "judge": null
}

Reuses candidate id swebench-pro, variant Public — same lane later used for Gemini 3.6's 58.7% and Grok 4.5's 64.7% (Grok's page does not state the Public qualifier; variant alignment needed before cross-vendor ranking).

打开官方来源

lcb 2887 Elo 模型 gemini-3-1-pro · 版本 LiveCodeBench Pro (Codeforces/ICPC/IOI problems), Elo · 指标 competition_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI / Elo · quote_snippet: LiveCodeBench Pro | Elo | 2887 | 2439 | — | — | 2393 | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Elo",
  "judge": null
}

Maps to existing benchmark lcb with variant Pro — LiveCodeBench Pro is a competitive-programming Elo edition, NOT comparable to pass-rate LCB windows (10/01-02/01 etc.) used elsewhere in the library; Elo vs percent metrics must never be merged.

打开官方来源

scicode 59% 模型 gemini-3-1-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: SciCode — Scientific research coding · quote_snippet: SciCode | | 59% | 56% | 47% | 52% | 52% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id scicode (registered batch 3). Same benchmark as Muse Glimmer's 43.6% (High reasoning).

打开官方来源

apex-agents 33.5% 模型 gemini-3-1-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: APEX-Agents — Long horizon professional tasks · quote_snippet: APEX-Agents | | 33.5% | 18.4% | — | 29.8% | 23.0% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id apex-agents (registered batch 1). Nearly doubles Gemini 3 Pro (18.4).

打开官方来源

gdpval-aa 1317 Elo 模型 gemini-3-1-pro · 版本 未说明 · 指标 knowledge_work_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: GDPval-AA — Elo / Expert tasks · quote_snippet: GDPval-AA | Elo | 1317 | 1195 | 1633 | 1606 | 1462 | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard-style Elo",
  "judge": null
}

Reuses candidate id gdpval-aa. Page label is GDPval-AA WITHOUT the v2 suffix used on Google's later pages and Meta's methodology (GDPVal-AA v2, 220 tasks, AA Stirrup harness, human baseline 1000) — variant left null here and flagged: same-family id, edition uncertainty recorded rather than assumed.

打开官方来源

tau2-bench 90.8% 模型 gemini-3-1-pro · 版本 Retail · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text), τ2-bench row · row: τ2-bench — Agentic and tool use / Retail · quote_snippet: τ2-bench | Retail | 90.8% | 85.3% | 91.7% | 91.9% | 82.0% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark tau2-bench variant Retail (canonical id for the τ²/τ2 generation; claude-haiku-4-5 and gpt-5 use the same tau2-bench Retail/Telecom lanes). Page row label is τ2-bench. 2026-09-01 audit: id corrected from tau-bench to tau2-bench — tau-bench is the τ1 generation in the catalog.

打开官方来源

tau2-bench 99.3% 模型 gemini-3-1-pro · 版本 Telecom · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text), τ2-bench row · row: τ2-bench — Agentic and tool use / Telecom · quote_snippet: Telecom | 99.3% | 98.0% | 97.9% | 99.3% | 98.7% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Saturation territory (tied with Opus 4.6 at 99.3); 1-point differences are noise. Secondary sources tie Grok 4.3's 'τ²-Bench Telecom 98%' claim to this same τ2 Telecom lane.

打开官方来源

mcp-atlas 69.2% 模型 gemini-3-1-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: MCP Atlas — Multi-step workflows using MCP · quote_snippet: MCP Atlas | | 69.2% | 54.1% | 61.3% | 59.5% | 60.6% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id mcp-atlas. No Public/Private split stated on this page (Muse Glimmer's row is explicitly the Public split; Meta's methodology describes 500+500 public/private splits) — split alignment needed before comparing 69.2 vs 75.5.

打开官方来源

browsecomp 85.9% 模型 gemini-3-1-pro · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: BrowseComp — Agentic search / Search + Python + Browse · quote_snippet: BrowseComp | Search + Python + Browse | 85.9% | 59.2% | 74.7% | 84.0% | 65.8% | —

{
  "harness": null,
  "tools": [
    "Search",
    "Python",
    "Browse"
  ],
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id browsecomp (registered batch 1). Tool stack explicit in the row label — comparability-critical vs tool-free BrowseComp numbers.

打开官方来源

mmmu 80.5% 模型 gemini-3-1-pro · 版本 Pro, no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: MMMU-Pro — Multimodal understanding and reasoning / No tools · quote_snippet: MMMU-Pro | No tools | 80.5% | 81.0% | 74.5% | 73.9% | 79.5% | —

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mmmu variant Pro. Gemini 3 Pro slightly ahead (81.0) on this row.

打开官方来源

mmmlu 92.6% 模型 gemini-3-1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text) · row: MMMLU — Multilingual Q&A · quote_snippet: MMMLU | | 92.6% | 91.8% | 89.3% | 91.1% | 89.6% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmlu (Multilingual MMLU Q&A) not in data/benchmarks/ — distinct from global-mmlu-lite (Lite edition used on other Google pages) and from multilingual-mmlu (Meta's label); three sibling ids now exist for multilingual MMLU variants, migration must keep them separate or unify with explicit edition variants.

打开官方来源

mrcr 84.9% 模型 gemini-3-1-pro · 版本 v2 (8-needle), 128k average · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text), MRCR v2 (8-needle) row · row: MRCR v2 (8-needle) / 128k (average) · quote_snippet: MRCR v2 (8-needle) | 128k (average) | 84.9% | 77.0% | 84.9% | 84.0% | 83.8% | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "cumulative score at 128k",
  "judge": null
}

Reuses mrcr (introduced batch 4) with v2 8-needle variant — same edition/lane as Gemini 2.0's 1M row base family and 3.5 Flash-Lite's 72.2%.

打开官方来源

mrcr 26.3% 模型 gemini-3-1-pro · 版本 v2 (8-needle), 1M pointwise · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance · table: Benchmark table (DOM text), MRCR v2 (8-needle) row · row: MRCR v2 (8-needle) / 1M (pointwise) · quote_snippet: 1M (pointwise) | 26.3% | 26.3% | Not supported | Not supported | Not supported | —

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "Thinking (High)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pointwise value at 1M context",
  "judge": null
}

1M pointwise lane — identical to Gemini 3 Pro (26.3); competitors marked 'Not supported' (no 1M context). Separate row from 128k average because the aggregations differ.

打开官方来源