Grok 4.7
xAI / Grok · 2026-09-21 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Grok 4.7
xAI 发布的最新旗舰编码与知识工作模型,基于 2.1T 参数底座与长程强化学习,在 CursorBench 4.0 (46.3%)、DeepSWE v1.1 (71.0%) 与 Terminal-Bench 4.0 (38.0%) 等多小时自主长程任务上展现出更强的自我检验与 500K 上下文管理能力。
- 输入模态
- 文本 / 代码 / 图像
- 上下文
- 500K
- 参数
- 2.1T
- 价格(每百万 tokens)
- USD 输入 2 / 输出 6 · 另有 fast 变体,价格为标准价的 2 倍
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: CursorBench 4.0 · quote_snippet: CursorBench 4.0 | 46.3%
{
"harness": "Cursor IDE Agent Harness",
"tools": "terminal + file edit + linter",
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}CursorBench 4.0 长程编码任务自报分 46.3%
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: DeepSWE v1.1 · quote_snippet: DeepSWE v1.1 | 71.0%
{
"harness": "DeepSWE Harness",
"tools": "bash + git",
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSWE v1.1 pass rate 71.0%
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: EEBench (Electrical Engineering) · quote_snippet: EEBench (Electrical Engineering) | 64.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}EEBench 电气工程学科领域推理基准。new-benchmark
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: AA Briefcase v1.1 · quote_snippet: AA Briefcase v1.1 | 1,657
{
"harness": "Artificial Analysis Briefcase Suite",
"tools": "office + spreadsheet + browser",
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Artificial Analysis Briefcase v1.1 办公综合任务 Elo 得分 1,657
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: Terminal-Bench 4.0 · quote_snippet: Terminal-Bench 4.0 | 38.0%
{
"harness": "Terminal-Bench CLI Harness",
"tools": "bash",
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Terminal-Bench 4.0 终端长程操作基准自报分 38.0%
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: HealthBench Professional · quote_snippet: HealthBench Professional | 56.7%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}HealthBench 临床医学专业评测 56.7%
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: Harvey Legal Agent Benchmark · quote_snippet: Harvey Legal Agent Benchmark | 19.6%
{
"harness": "Harvey Legal Agent Framework",
"tools": "legal retrieval + drafting",
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Harvey Legal Agent Benchmark 复杂法律智能体评测 19.6%
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: LatchBio Biosafety · quote_snippet: LatchBio Biosafety | 62.4%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}LatchBio 生物安全基准测试。new-benchmark
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations · table: Main Benchmark Table · row: HackerBench v0.3 Dual-Use Risk · quote_snippet: HackerBench v0.3 Dual-Use Risk | 3.3%
{
"harness": "HackerBench Evaluation Harness",
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}HackerBench v0.3 双用途网络安全风险提示准入率(越低越安全)。new-benchmark