← 模型目录

Grok 4.7

xAI / Grok · 2026-09-21 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Grok 4.7

xAI 发布的最新旗舰编码与知识工作模型,基于 2.1T 参数底座与长程强化学习,在 CursorBench 4.0 (46.3%)、DeepSWE v1.1 (71.0%) 与 Terminal-Bench 4.0 (38.0%) 等多小时自主长程任务上展现出更强的自我检验与 500K 上下文管理能力。

输入模态
文本 / 代码 / 图像
上下文
500K
参数
2.1T
价格(每百万 tokens)
USD 输入 2 / 输出 6 · 另有 fast 变体,价格为标准价的 2 倍

本变体的评测证据

cursor-bench 46.3% 模型 grok-4-7 · 版本 4.0 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: CursorBench 4.0 · quote_snippet: CursorBench 4.0 | 46.3%

{
  "harness": "Cursor IDE Agent Harness",
  "tools": "terminal + file edit + linter",
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

CursorBench 4.0 长程编码任务自报分 46.3%

打开官方来源

deepswe 71.0% 模型 grok-4-7 · 版本 1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: DeepSWE v1.1 · quote_snippet: DeepSWE v1.1 | 71.0%

{
  "harness": "DeepSWE Harness",
  "tools": "bash + git",
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

DeepSWE v1.1 pass rate 71.0%

打开官方来源

eebench 64.0% 模型 grok-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: EEBench (Electrical Engineering) · quote_snippet: EEBench (Electrical Engineering) | 64.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

EEBench 电气工程学科领域推理基准。new-benchmark

打开官方来源

aa-briefcase 1,657 模型 grok-4-7 · 版本 1.1 · 指标 elo · 单位 elo 来源等级 A · third_party_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: AA Briefcase v1.1 · quote_snippet: AA Briefcase v1.1 | 1,657

{
  "harness": "Artificial Analysis Briefcase Suite",
  "tools": "office + spreadsheet + browser",
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Artificial Analysis Briefcase v1.1 办公综合任务 Elo 得分 1,657

打开官方来源

terminalbench 38.0% 模型 grok-4-7 · 版本 4.0 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: Terminal-Bench 4.0 · quote_snippet: Terminal-Bench 4.0 | 38.0%

{
  "harness": "Terminal-Bench CLI Harness",
  "tools": "bash",
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Terminal-Bench 4.0 终端长程操作基准自报分 38.0%

打开官方来源

healthbench 56.7% 模型 grok-4-7 · 版本 Professional · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: HealthBench Professional · quote_snippet: HealthBench Professional | 56.7%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

HealthBench 临床医学专业评测 56.7%

打开官方来源

harvey-lab 19.6% 模型 grok-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: Harvey Legal Agent Benchmark · quote_snippet: Harvey Legal Agent Benchmark | 19.6%

{
  "harness": "Harvey Legal Agent Framework",
  "tools": "legal retrieval + drafting",
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Harvey Legal Agent Benchmark 复杂法律智能体评测 19.6%

打开官方来源

latchbio 62.4% 模型 grok-4-7 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: LatchBio Biosafety · quote_snippet: LatchBio Biosafety | 62.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

LatchBio 生物安全基准测试。new-benchmark

打开官方来源

hackerbench 3.3% 模型 grok-4-7 · 版本 0.3 · 指标 risky_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-22 · 距快照 12 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations · table: Main Benchmark Table · row: HackerBench v0.3 Dual-Use Risk · quote_snippet: HackerBench v0.3 Dual-Use Risk | 3.3%

{
  "harness": "HackerBench Evaluation Harness",
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

HackerBench v0.3 双用途网络安全风险提示准入率(越低越安全)。new-benchmark

打开官方来源