← 模型目录

Gemini 4 Argon

Google DeepMind / Gemini · 2026-09-30 · 通用模型

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Gemini 4 Argon

Google 的新前沿模型,面向真实软件工程、法律与金融等企业知识工作和网络安全防御,输出 token 上限提升到 1M。收录的 7 项评测以编码、智能体与安全为主,如 DeepSWE v1.1 77.9%、LVBench 91.7%;目前仅向可信网络防御者开放。

输入模态
官方资料未说明
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 2 / 输出 10 · 介绍期价格,缓存输入价比输入价低 95%;介绍期结束后为输入 $4、输出 $20

本变体的评测证据

deepswe 77.9% 模型 gemini-4-argon · 版本 DeepSWE v1.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling coding and enterprise workflows across domains · row: DeepSWE v1.1 · quote_snippet: sets a new state of the art on DeepSWE v1.1 (77.9%)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

页面称这是该基准的新最高水平。

打开官方来源

automationbench 51.3% 模型 gemini-4-argon · 版本 AutomationBench · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling coding and enterprise workflows across domains · row: AutomationBench · quote_snippet: Argon ranks #1 with a score of 51.3%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

页面称由 Zapier 提出,排名第一。

打开官方来源

lvbench 91.7% 模型 gemini-4-argon · 版本 LVBench · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling coding and enterprise workflows across domains · row: LVBench · quote_snippet: Argon is state of the art with a score of 91.7%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

长视频理解基准,页面称为最高水平。

打开官方来源

cwe-bench 68% 模型 gemini-4-argon · 版本 CWE-bench v1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Leading in defensive cybersecurity · row: CWE-bench v1 · quote_snippet: Argon ties for first place with a top score of 68%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: CWE-bench 衡量模型修复安全漏洞的能力;页面称与其他模型并列第一。

打开官方来源

vals-index 该来源尚无可读数值 模型 gemini-4-argon · 版本 Vals Index · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling coding and enterprise workflows across domains · row: Vals Index · quote_snippet: Argon is the leading model on the Vals Index

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: Vals Index 衡量金融、编码、法律与税务工作的经济影响,按各行业占美国 GDP 的比重加权;页面称 Argon 领先,未给数值。

打开官方来源

vals-finance-agent 该来源尚无可读数值 模型 gemini-4-argon · 版本 Vals Finance Agent v2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-10-03 · 距快照 1 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Enabling coding and enterprise workflows across domains · row: Vals Finance Agent v2 · quote_snippet: similarly leading performance across ... Vals Finance Agent v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: 多步骤金融研究基准;页面称表现领先,未给数值。

打开官方来源