← 模型目录

GPT-6 Astra

OpenAI · 2026-09-03 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

GPT-6 Astra

OpenAI 首个将 Preparedness 框架推至 Critical 等级的旗舰前沿模型,在计算机直接操控、复杂多步长程自主 Agent、软件工程与零日(zero-day)漏洞挖掘上实现重大跃迁;在 ARC-AGI-3 和 ExploitBench 达到饱和得分。

输入模态
文本 / 图像 / 代码 / 计算机操作
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 10 / 输出 50 · API 定价每百万输入 10 美元,输出 50 美元

本变体的评测证据

arc-agi 99.9% 模型 gpt-6-astra · 版本 3 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-04 · 距快照 30 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance Benchmarks · quote_snippet: saturates ARC-AGI-3 at 99.9%

{
  "harness": "OpenAI agentic harness with retained reasoning and custom compaction",
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

OpenAI 官方博客公布数值 99.9%。官方注明使用了带有 retained reasoning 与自定义 compaction 的自研 agentic harness;而在 ARC Prize 官方标准无状态 harness 下评测值约为 62.7%,两者协议完全不同,不可直接比较。

打开官方来源

frontiermath 97.6% 模型 gpt-6-astra · 版本 Tier 4 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-04 · 距快照 30 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance Benchmarks · quote_snippet: FrontierMath Tier 4: 97.6%

{
  "harness": null,
  "tools": [
    "python"
  ],
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

官方公布研究级数学难题 FrontierMath Tier 4 得分 97.6%。

打开官方来源

exploitbench 100% 模型 gpt-6-astra · 版本 未说明 · 指标 ladder-progress · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-04 · 距快照 30 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance Benchmarks · quote_snippet: ExploitBench: 100%

{
  "harness": "Daybreak secure sandboxed environment",
  "tools": [
    "shell",
    "code"
  ],
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

首个在 ExploitBench 达成 16 级完整利用阶梯 100% 推进的模型;安全防护达到 Preparedness Critical 等级,相关攻击性利用能力仅受限于 Daybreak 计划安全防护机制。

打开官方来源