Gemini 3.1 Pro
Google DeepMind / Gemini · 2026-02-19 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 3.1 Pro
Google 将 Gemini 3.1 Pro 定位为「为最复杂任务打造的更聪明模型」,基于 Gemini 3 Pro,发布时为 preview(thinking 档 low/medium/high)。已收录 19 项评测横跨推理、代码、多模态与长上下文:亮点 GPQA 94.3%、SWE-bench Verified 80.6%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text), Humanity's Last Exam row · row: Humanity's Last Exam — Academic reasoning (full set, text + MM) / No tools · quote_snippet: Humanity's Last Exam... No tools | 44.4% | 37.5% | 33.2% | 40.0% | 34.5% | —
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark hlehle. Full-set text+MM edition with the no-tools condition isolated as its own row (competitors: Gemini 3 Pro 37.5 / Sonnet 4.6 33.2 / Opus 4.6 40.0 / GPT-5.2 34.5 — table columns, kept in notes per batch-1 口径). Companion row with Search+Code tools recorded separately.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text), Humanity's Last Exam row, tool-augmented condition · row: Humanity's Last Exam / Search (blocklist) + Code · quote_snippet: Search (blocklist) + Code | 51.4% | 45.8% | 49.0% | 53.1% | 45.5% | —
{
"harness": null,
"tools": [
"Search (blocklist)",
"Code execution"
],
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Tool-augmented lane of the same HLE row (Opus 4.6 leads at 53.1). NOT comparable to the no-tools row or to other vendors' tool-free HLE numbers.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: ARC-AGI-2 — ARC Prize Verified · quote_snippet: ARC-AGI-2 | ARC Prize Verified | 77.1% | 31.1% | 58.3% | 68.8% | 52.9% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": "ARC Prize Verified"
}Maps to existing benchmark arc-agi variant 2. The release blog's headline number ('verified score of 77.1%... more than double 3 Pro'). Third-party-verified lane — distinct from Gemini 3 Deep Think's 45.1% with-code-execution row (different tool condition).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: GPQA Diamond — Scientific knowledge / No tools · quote_snippet: GPQA Diamond | No tools | 94.3% | 91.9% | 89.9% | 91.3% | 92.4% | —
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark gpqa (canonical Diamond).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: Terminal-Bench 2.0 — Agentic terminal coding / Terminus-2 harness · quote_snippet: Terminal-Bench 2.0 | Terminus-2 harness | 68.5% | 56.9% | 59.1% | 65.4% | 54.0% | 64.7%
{
"harness": "Terminus-2",
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark terminalbench variant 2.0. Same-harness lane (Terminus-2) as Muse Glimmer's TB 2.1 row — but different TB VERSION (2.0 vs 2.1); still not interchangeable. Table also carries an 'Other best self-reported harness' sub-row where GPT-5.3-Codex reaches 77.3 (Codex harness) — GPT-5.2 62.2 (Codex); Gemini columns are '—' there, so it is recorded here rather than as a comparison_cited row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: SWE-Bench Verified — Agentic coding / Single attempt · quote_snippet: SWE-Bench Verified | Single attempt | 80.6% | 76.2% | 79.6% | 80.8% | 80.0% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "single attempt (pass@1 lane)",
"judge": null
}Maps to existing benchmark swebench. Opus 4.6 edges this row at 80.8 (table column, in notes). Single-attempt condition explicit — distinct from Google's 'multiple attempts' scaffolding lane used on other pages.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: SWE-Bench Pro (Public) — Diverse agentic coding tasks / Single attempt · quote_snippet: SWE-Bench Pro (Public) | Single attempt | 54.2% | 43.3% | — | — | 55.6% | 56.8%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "single attempt",
"judge": null
}Reuses candidate id swebench-pro, variant Public — same lane later used for Gemini 3.6's 58.7% and Grok 4.5's 64.7% (Grok's page does not state the Public qualifier; variant alignment needed before cross-vendor ranking).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: LiveCodeBench Pro — Competitive coding problems from Codeforces, ICPC, and IOI / Elo · quote_snippet: LiveCodeBench Pro | Elo | 2887 | 2439 | — | — | 2393 | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Elo",
"judge": null
}Maps to existing benchmark lcb with variant Pro — LiveCodeBench Pro is a competitive-programming Elo edition, NOT comparable to pass-rate LCB windows (10/01-02/01 etc.) used elsewhere in the library; Elo vs percent metrics must never be merged.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: SciCode — Scientific research coding · quote_snippet: SciCode | | 59% | 56% | 47% | 52% | 52% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id scicode (registered batch 3). Same benchmark as Muse Glimmer's 43.6% (High reasoning).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: APEX-Agents — Long horizon professional tasks · quote_snippet: APEX-Agents | | 33.5% | 18.4% | — | 29.8% | 23.0% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id apex-agents (registered batch 1). Nearly doubles Gemini 3 Pro (18.4).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: GDPval-AA — Elo / Expert tasks · quote_snippet: GDPval-AA | Elo | 1317 | 1195 | 1633 | 1606 | 1462 | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard-style Elo",
"judge": null
}Reuses candidate id gdpval-aa. Page label is GDPval-AA WITHOUT the v2 suffix used on Google's later pages and Meta's methodology (GDPVal-AA v2, 220 tasks, AA Stirrup harness, human baseline 1000) — variant left null here and flagged: same-family id, edition uncertainty recorded rather than assumed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text), τ2-bench row · row: τ2-bench — Agentic and tool use / Retail · quote_snippet: τ2-bench | Retail | 90.8% | 85.3% | 91.7% | 91.9% | 82.0% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark tau2-bench variant Retail (canonical id for the τ²/τ2 generation; claude-haiku-4-5 and gpt-5 use the same tau2-bench Retail/Telecom lanes). Page row label is τ2-bench. 2026-09-01 audit: id corrected from tau-bench to tau2-bench — tau-bench is the τ1 generation in the catalog.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text), τ2-bench row · row: τ2-bench — Agentic and tool use / Telecom · quote_snippet: Telecom | 99.3% | 98.0% | 97.9% | 99.3% | 98.7% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Saturation territory (tied with Opus 4.6 at 99.3); 1-point differences are noise. Secondary sources tie Grok 4.3's 'τ²-Bench Telecom 98%' claim to this same τ2 Telecom lane.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: MCP Atlas — Multi-step workflows using MCP · quote_snippet: MCP Atlas | | 69.2% | 54.1% | 61.3% | 59.5% | 60.6% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id mcp-atlas. No Public/Private split stated on this page (Muse Glimmer's row is explicitly the Public split; Meta's methodology describes 500+500 public/private splits) — split alignment needed before comparing 69.2 vs 75.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: BrowseComp — Agentic search / Search + Python + Browse · quote_snippet: BrowseComp | Search + Python + Browse | 85.9% | 59.2% | 74.7% | 84.0% | 65.8% | —
{
"harness": null,
"tools": [
"Search",
"Python",
"Browse"
],
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id browsecomp (registered batch 1). Tool stack explicit in the row label — comparability-critical vs tool-free BrowseComp numbers.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: MMMU-Pro — Multimodal understanding and reasoning / No tools · quote_snippet: MMMU-Pro | No tools | 80.5% | 81.0% | 74.5% | 73.9% | 79.5% | —
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark mmmu variant Pro. Gemini 3 Pro slightly ahead (81.0) on this row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text) · row: MMMLU — Multilingual Q&A · quote_snippet: MMMLU | | 92.6% | 91.8% | 89.3% | 91.1% | 89.6% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmlu (Multilingual MMLU Q&A) not in data/benchmarks/ — distinct from global-mmlu-lite (Lite edition used on other Google pages) and from multilingual-mmlu (Meta's label); three sibling ids now exist for multilingual MMLU variants, migration must keep them separate or unify with explicit edition variants.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text), MRCR v2 (8-needle) row · row: MRCR v2 (8-needle) / 128k (average) · quote_snippet: MRCR v2 (8-needle) | 128k (average) | 84.9% | 77.0% | 84.9% | 84.0% | 83.8% | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "cumulative score at 128k",
"judge": null
}Reuses mrcr (introduced batch 4) with v2 8-needle variant — same edition/lane as Gemini 2.0's 1M row base family and 3.5 Flash-Lite's 72.2%.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Performance · table: Benchmark table (DOM text), MRCR v2 (8-needle) row · row: MRCR v2 (8-needle) / 1M (pointwise) · quote_snippet: 1M (pointwise) | 26.3% | 26.3% | Not supported | Not supported | Not supported | —
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "Thinking (High)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pointwise value at 1M context",
"judge": null
}1M pointwise lane — identical to Gemini 3 Pro (26.3); competitors marked 'Not supported' (no 1M context). Separate row from 128k average because the aggregations differ.