DeepSeek V4 Pro
DeepSeek · 2026-08-13 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
DeepSeek V4 Pro
DeepSeek 发布 V4 Pro 正式版,同步上线 APP/网页端/API(API 模型名不变),官方定位为正式版增强了 Agent 能力、在生产环境中的性能表现提升尤为显著。本发布收录的 10 项评测全部为 Agent 向(编码智能体、网络安全、工具调用、数据科学与自动化智能体),官方以单张对比图给出与 GLM-5.2 / Kimi-K3 / Opus-4.8 / Fable 5 的对比,数值以人工读图确认前不计入公开计数。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: HLE (wo/w tools) · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 1 行, 双值=无工具/带工具)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致:模型直读 + MiniMax OCR,2026-09-02),未经人工复核不得翻 verified。V4-Pro-0813 = 42.7/60.0(无工具/带工具);DeepSeek-V4-Flash-0731 37.8/51.5;V4-Pro-Preview 37.7/48.2;V4-Flash-Preview 34.8/45.1;GLM-5.2 40.5/54.7;Kimi-K3 43.5/56.0;Opus-4.8 49.8/57.9;Fable 5 (w/ fallback) 53.3/63.0。表内 V4-Pro-Preview 值与既有 deepseek-v4.json(预览版)verified 行一致(37.7 / w-tools 48.2),互证读图可信。表注 Code Agent 协议(DeepSeek Harness 极简模式 / max / top_p 0.95 / temp 1.0)明示只适用 Code Agent 任务,本行非 Code Agent → protocol 全 null。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: Terminal Bench 2.1 · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 2 行)
{
"harness": "DeepSeek Harness 极简模式",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 87.9;Flash-0731 82.7;V4-Pro-Preview 72.1;V4-Flash-Preview 61.8;GLM-5.2 81.0;Kimi-K3 88.3;Opus-4.8 85.0;Fable 5 (w/ fallback) 88.0。Code Agent 任务 → 表注协议按行记入 protocol(harness=DeepSeek Harness 极简模式 / max 档位 / topp=0.95 / temperature=1.0;表注并声明其他框架下结果可能略有不同)。同表 Flash-0731 82.7 与 Opus-4.8 85.0 与已人工核验的 260821 Vision-Exp 表(deepseek-v4-flash-vision-exp.json)逐值一致,互证读图可信。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: NL2Repo · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 3 行)
{
"harness": "DeepSeek Harness 极简模式",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 61.5;Flash-0731 54.2;V4-Pro-Preview 38.5;V4-Flash-Preview 39.4;GLM-5.2 48.9;Kimi-K3 未报告;Opus-4.8 69.7;Fable 5 (w/ fallback) 未报告。Code Agent 任务 → 表注协议按行记入 protocol。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: Cybergym · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 4 行)
{
"harness": "DeepSeek Harness 极简模式",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 83.3;Flash-0731 76.7;V4-Pro-Preview 52.7;V4-Flash-Preview 38.7;GLM-5.2 未报告;Kimi-K3 80.0;Opus-4.8 78.3;Fable 5 (w/ fallback) 83.1。CyberGym 为代码级安全任务(PoC 复现)→ 按 Code Agent 任务记入表注协议。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: DeepSWE · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 5 行)
{
"harness": "DeepSeek Harness 极简模式",
"tools": null,
"shots": null,
"reasoning_effort": "max",
"temperature": 1,
"top_p": 0.95,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 62.7;Flash-0731 54.4;V4-Pro-Preview 12.8;V4-Flash-Preview 7.3;GLM-5.2 46.2;Kimi-K3 67.5;Opus-4.8 58.0;Fable 5 (w/ fallback) 70.0。SWE 编码智能体任务 → 表注协议按行记入 protocol。注意:本表 Kimi-K3 列 67.5 与 GLM-5.2 页记录一致,但 Kimi 自家脚注曾自记 67.3(跨厂冲突已在 README 待裁定清单);跨厂商 DeepSWE harness 不一(260821 页 DeepSeek 曾用自有 harness),比较须按协议区分。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: Toolathlon-Verified · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 6 行)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 74.1;Flash-0731 70.3;V4-Pro-Preview 55.9;V4-Flash-Preview 49.7;GLM-5.2 59.9;Kimi-K3 76.5;Opus-4.8 76.2;Fable 5 (w/ fallback) 77.9。行为工具调用智能体评测而非表注定义的 Code Agent 任务,协议适用性未明示 → protocol 全 null。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: Agents' Last Exam · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 7 行)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 25.7;Flash-0731 25.2;V4-Pro-Preview 16.5;V4-Flash-Preview 15.8;GLM-5.2 23.8;Kimi-K3 27.6;Opus-4.8 25.7;Fable 5 (w/ fallback) 未报告。通用智能体评测,非表注定义的 Code Agent 任务 → protocol 全 null。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: AutomationBench (Public) · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 8 行)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 31.8;Flash-0731 25.1;V4-Pro-Preview 12.8;V4-Flash-Preview 10.8;GLM-5.2 12.9;Kimi-K3 30.8;Opus-4.8 27.2;Fable 5 (w/ fallback) 29.1。自动化智能体评测,非表注定义的 Code Agent 任务 → protocol 全 null。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: DSBench-FullStack · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 9 行)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: dsbench-fullstack(DSBench 的 DeepSeek FullStack 划分,页面未定义划分语义;与既有 dsbench-hard 平行,已同步建实体 data/benchmarks/dsbench-fullstack.json)。视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 71.1;Flash-0731 68.7;V4-Pro-Preview 41.8;V4-Flash-Preview 37.0;GLM-5.2 61.8;Kimi-K3 73.7;Opus-4.8 71.6;Fable 5 (w/ fallback) 77.2。数据科学智能体评测,非表注定义的 Code Agent 任务 → protocol 全 null。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agent 能力大幅提升 · row: DSBench-Hard · figure: models/2026-08-13-deepseek-v4-pro/images/v4_260813_benchmark_table_cn.png (第 10 行)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}视觉转写(双路径一致),未经人工复核不得翻 verified。V4-Pro-0813 = 67.2;Flash-0731 59.6;V4-Pro-Preview 31.1;V4-Flash-Preview 25.8;GLM-5.2 54.5;Kimi-K3 63.0;Opus-4.8 71.7;Fable 5 (w/ fallback) 68.3。同表 Flash-0731 59.6 与已人工核验的 260821 Vision-Exp 表逐值一致,互证读图可信。数据科学智能体评测,非表注定义的 Code Agent 任务 → protocol 全 null。