← 模型目录

Gemini 3 Deep Think

Google DeepMind / Gemini · 2025-12-04 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 3 Deep Think

谷歌将 Gemini 3 Deep Think 定位为 Gemini 应用内面向 Google AI Ultra 订阅用户推出的增强推理模式,专攻挑战最先进模型的复杂数学、科学与逻辑问题。已收录评测为通用推理与抽象推理两项,HLE 41.0%(无工具)、ARC-AGI-2 45.1%(带代码执行)。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 41.0% 模型 gemini-3-deep-think · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Body (single-paragraph rollout announcement) · quote_snippet: industry leading on rigorous benchmarks like Humanity's Last Exam (41.0% without the use of tools)

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": "Deep Think (parallel reasoning over multiple hypotheses)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Tool-free condition stated on page (protocol.tools = []). Same value as the row on the Gemini 3 launch page; recorded here from this page's own source URL for release-level provenance.

打开官方来源

arc-agi 45.1% 模型 gemini-3-deep-think · 版本 2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Body (single-paragraph rollout announcement) · quote_snippet: ARC-AGI-2 (an unprecedented 45.1% with code execution)

{
  "harness": null,
  "tools": [
    "code execution"
  ],
  "shots": null,
  "reasoning_effort": "Deep Think (parallel reasoning over multiple hypotheses)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark arc-agi, variant 2. Code execution is an explicit protocol condition; the launch page additionally qualifies the number as 'ARC Prize Verified', which this rollout post does not repeat.

打开官方来源