Gemini 4 Argon
Google DeepMind / Gemini · 2026-09-30 · 通用模型
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Gemini 4 Argon
Google 的新前沿模型,面向真实软件工程、法律与金融等企业知识工作和网络安全防御,输出 token 上限提升到 1M。收录的 7 项评测以编码、智能体与安全为主,如 DeepSWE v1.1 77.9%、LVBench 91.7%;目前仅向可信网络防御者开放。
- 输入模态
- 官方资料未说明
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- USD 输入 2 / 输出 10 · 介绍期价格,缓存输入价比输入价低 95%;介绍期结束后为输入 $4、输出 $20
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enabling coding and enterprise workflows across domains · row: DeepSWE v1.1 · quote_snippet: sets a new state of the art on DeepSWE v1.1 (77.9%)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面称这是该基准的新最高水平。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enabling coding and enterprise workflows across domains · row: AutomationBench · quote_snippet: Argon ranks #1 with a score of 51.3%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面称由 Zapier 提出,排名第一。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enabling coding and enterprise workflows across domains · row: LVBench · quote_snippet: Argon is state of the art with a score of 91.7%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}长视频理解基准,页面称为最高水平。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Leading in defensive cybersecurity · row: CWE-bench v1 · quote_snippet: Argon ties for first place with a top score of 68%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: CWE-bench 衡量模型修复安全漏洞的能力;页面称与其他模型并列第一。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enabling coding and enterprise workflows across domains · row: Vals Index · quote_snippet: Argon is the leading model on the Vals Index
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: Vals Index 衡量金融、编码、法律与税务工作的经济影响,按各行业占美国 GDP 的比重加权;页面称 Argon 领先,未给数值。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enabling coding and enterprise workflows across domains · row: Vals Finance Agent v2 · quote_snippet: similarly leading performance across ... Vals Finance Agent v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: 多步骤金融研究基准;页面称表现领先,未给数值。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Enabling coding and enterprise workflows across domains · row: Harvey's Legal Agent Benchmark · quote_snippet: Harvey's Legal Agent Benchmark (legal research and drafting)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}页面称表现领先,未给数值。