DeepSeek-V4-Pro / DeepSeek-V4-Flash
DeepSeek · 2026-04-24 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
DeepSeek-V4-Pro
DeepSeek 发布 V4 预览版「迈入百万上下文普惠时代」,V4-Pro 为旗舰:1M 上下文、思考/非思考双模式、reasoning_effort high/max。评测覆盖推理、编码、智能体与百万级长上下文,亮点如 GPQA Diamond 90.1%、Codeforces 3206。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: MMLU-Pro · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V4-Pro 87.5, V4-Flash 86.2. Columns: K2.6 Thinking 87.1, GLM-5.1 86.0, Opus-4.6 Max 89.1, GPT-5.4 xHigh 87.5, Gemini-3.1-Pro High 91.0. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 86.2;竞品列 K2.6 Thinking 87.1, GLM-5.1 Thinking 86, Opus-4.6 Max 89.1, GPT-5.4 xHigh 87.5, Gemini-3.1-Pro High 91。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SimpleQA-Verified · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: simpleqa-verified (Verified split, distinct from simpleqa) not yet in data/benchmarks/. Vision read: V4-Pro 57.9, V4-Flash 34.1; K2.6 36.9, GLM-5.1 38.1, Opus 46.2, GPT-5.4 45.3, Gemini 75.6. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 34.1;竞品列 K2.6 Thinking 36.9, GLM-5.1 Thinking 38.1, Opus-4.6 Max 46.2, GPT-5.4 xHigh 45.3, Gemini-3.1-Pro High 75.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Chinese-SimpleQA · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: chinese-simpleqa not yet in data/benchmarks/. Vision read: V4-Pro 84.4, V4-Flash 78.9; K2.6 75.9, GLM-5.1 75.0, Opus 76.2, GPT-5.4 76.8, Gemini 85.9. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 78.9;竞品列 K2.6 Thinking 75.9, GLM-5.1 Thinking 75, Opus-4.6 Max 76.2, GPT-5.4 xHigh 76.8, Gemini-3.1-Pro High 85.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: GPQA Diamond · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 90.1, V4-Flash 88.1; K2.6 90.5, GLM-5.1 86.2, Opus 91.3, GPT-5.4 93.0, Gemini 94.3. Matches GLM-5.2's V4-Pro column (90.1). 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 88.1;竞品列 K2.6 Thinking 90.5, GLM-5.1 Thinking 86.2, Opus-4.6 Max 91.3, GPT-5.4 xHigh 93, Gemini-3.1-Pro High 94.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: HLE · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 37.7, V4-Flash 34.8; K2.6 36.4, GLM-5.1 34.7, Opus 40.0, GPT-5.4 39.8, Gemini 44.4. Matches GLM-5.2's column (37.7). 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 34.8;竞品列 K2.6 Thinking 36.4, GLM-5.1 Thinking 34.7, Opus-4.6 Max 40, GPT-5.4 xHigh 39.8, Gemini-3.1-Pro High 44.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: LiveCodeBench · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 93.5, V4-Flash 91.6; K2.6 89.6, Opus 88.8, Gemini 91.7. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 91.6;竞品列 K2.6 Thinking 89.6, GLM-5.1 Thinking -, Opus-4.6 Max 88.8, GPT-5.4 xHigh -, Gemini-3.1-Pro High 91.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Codeforces · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 3206, V4-Flash 3052; Opus 3168, GPT-5.4 3052. Elo rating. The separate summary chart (v4-benchmark.png) also shows 3206 - two images agree. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 3052;竞品列 K2.6 Thinking -, GLM-5.1 Thinking -, Opus-4.6 Max -, GPT-5.4 xHigh 3168, Gemini-3.1-Pro High 3052。 注意:主表该行 K2.6/GLM-5.1/Opus-4.6 均为 '-',GPT-5.4 xHigh=3168、Gemini-3.1-Pro=3052;而汇总图 v4-benchmark.png (images/03.png) 把 3168 标在 Opus-4.6-Max、3052 标在 GPT-5.4-xHigh,两张官方图竞品归属互相矛盾;仅本行 V4 值 3206 两图一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: HMMT 2026 Feb · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hmmt-26. Vision read: V4-Pro 95.2, V4-Flash 94.8; K2.6 92.7, GLM-5.1 89.4, Opus 96.2, GPT-5.4 97.7, Gemini 94.7. Matches GLM-5.2's V4-Pro column (95.2). 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 94.8;竞品列 K2.6 Thinking 92.7, GLM-5.1 Thinking 89.4, Opus-4.6 Max 96.2, GPT-5.4 xHigh 97.7, Gemini-3.1-Pro High 94.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: IMOAnswerBench · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 89.8, V4-Flash 88.4; K2.6 86.0, GLM-5.1 83.8, Opus 75.3, GPT-5.4 91.4, Gemini 81.0. Matches GLM-5.2's column (89.8). 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 88.4;竞品列 K2.6 Thinking 86, GLM-5.1 Thinking 83.8, Opus-4.6 Max 75.3, GPT-5.4 xHigh 91.4, Gemini-3.1-Pro High 81。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Apex · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: apex (full Apex benchmark; distinct from apex-agents and from the Apex Shortlist row) not yet in data/benchmarks/. Vision read: V4-Pro 38.3, V4-Flash 33.0; K2.6 24.0, GLM-5.1 11.5, Opus 34.5, GPT-5.4 54.1, Gemini 60.9. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 33;竞品列 K2.6 Thinking 24, GLM-5.1 Thinking 11.5, Opus-4.6 Max 34.5, GPT-5.4 xHigh 54.1, Gemini-3.1-Pro High 60.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Apex Shortlist · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: apex-shortlist not yet in data/benchmarks/. Vision read: V4-Pro 90.2, V4-Flash 85.7; K2.6 75.5, GLM-5.1 72.4, Opus 85.9, GPT-5.4 78.1, Gemini 89.1. Also on the summary chart (90.2) - two images agree. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 85.7;竞品列 K2.6 Thinking 75.5, GLM-5.1 Thinking 72.4, Opus-4.6 Max 85.9, GPT-5.4 xHigh 78.1, Gemini-3.1-Pro High 89.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: MRCR 1M · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mrcr-1m not yet in data/benchmarks/. Vision read: V4-Pro 83.5, V4-Flash 78.7; Opus 92.9, Gemini 76.3. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 78.7;竞品列 K2.6 Thinking -, GLM-5.1 Thinking -, Opus-4.6 Max 92.9, GPT-5.4 xHigh -, Gemini-3.1-Pro High 76.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: CorpusQA 1M · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: corpusqa-1m not yet in data/benchmarks/. Vision read: V4-Pro 62.0, V4-Flash 60.5; Opus 71.7, Gemini 53.8. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 60.5;竞品列 K2.6 Thinking -, GLM-5.1 Thinking -, Opus-4.6 Max 71.7, GPT-5.4 xHigh -, Gemini-3.1-Pro High 53.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Terminal Bench 2.0 · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 67.9, V4-Flash 56.9; K2.6 66.7, GLM-5.1 63.5, Opus 65.4, GPT-5.4 75.1, Gemini 68.5. Harness undisclosed. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 56.9;竞品列 K2.6 Thinking 66.7, GLM-5.1 Thinking 63.5, Opus-4.6 Max 65.4, GPT-5.4 xHigh 75.1, Gemini-3.1-Pro High 68.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SWE Verified · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 80.6, V4-Flash 79.0; K2.6 80.2, Opus 80.8, Gemini 80.6. Summary chart also 80.6 - agree. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 79;竞品列 K2.6 Thinking 80.2, GLM-5.1 Thinking -, Opus-4.6 Max 80.8, GPT-5.4 xHigh -, Gemini-3.1-Pro High 80.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SWE Pro · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 55.4, V4-Flash 52.6; K2.6 58.6, GLM-5.1 58.4, Opus 57.3, GPT-5.4 57.7, Gemini 54.2. Matches GLM-5.2's column (55.4). 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 52.6;竞品列 K2.6 Thinking 58.6, GLM-5.1 Thinking 58.4, Opus-4.6 Max 57.3, GPT-5.4 xHigh 57.7, Gemini-3.1-Pro High 54.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SWE Multilingual · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual. Vision read: V4-Pro 76.2, V4-Flash 73.3; K2.6 76.7, GLM-5.1 73.3, Opus 77.5. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 73.3;竞品列 K2.6 Thinking 76.7, GLM-5.1 Thinking 73.3, Opus-4.6 Max 77.5, GPT-5.4 xHigh -, Gemini-3.1-Pro High -。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: BrowseComp · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp. Vision read: V4-Pro 83.4, V4-Flash 73.2; K2.6 83.2, GLM-5.1 79.3, Opus 83.7, GPT-5.4 82.7, Gemini 85.9. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 73.2;竞品列 K2.6 Thinking 83.2, GLM-5.1 Thinking 79.3, Opus-4.6 Max 83.7, GPT-5.4 xHigh 82.7, Gemini-3.1-Pro High 85.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: HLE w/tools · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 48.2, V4-Flash 45.1; K2.6 54.0, GLM-5.1 50.4, Opus 53.1, GPT-5.4 52.0, Gemini 51.6. Matches GLM-5.2's column (48.2). 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 45.1;竞品列 K2.6 Thinking 54, GLM-5.1 Thinking 50.4, Opus-4.6 Max 53.1, GPT-5.4 xHigh 52, Gemini-3.1-Pro High 51.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: GDPval-AA · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}gdpval-aa id already introduced by prior batches. Elo unit. Vision read: V4-Pro 1554, V4-Flash 1395; K2.6 1482, GLM-5.1 1535, Opus 1619, GPT-5.4 1674, Gemini 1314. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 1395;竞品列 K2.6 Thinking 1482, GLM-5.1 Thinking 1535, Opus-4.6 Max 1619, GPT-5.4 xHigh 1674, Gemini-3.1-Pro High 1314。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: MCPAtlas Public · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read: V4-Pro 73.6, V4-Flash 69.0; K2.6 66.6, GLM-5.1 71.8, Opus 73.8, GPT-5.4 67.2, Gemini 69.2. Matches GLM-5.2's column (73.6). 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 69;竞品列 K2.6 Thinking 66.6, GLM-5.1 Thinking 71.8, Opus-4.6 Max 73.8, GPT-5.4 xHigh 67.2, Gemini-3.1-Pro High 69.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Toolathlon · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}toolathlon id already introduced by prior batches. Vision read: V4-Pro 51.8, V4-Flash 47.8; K2.6 50.0, GLM-5.1 40.7, Opus 47.2, GPT-5.4 54.6, Gemini 48.8. Summary chart also 51.8 - agree. 视觉转写自归档图 images/04.png(api-docs v4-benchmark-2.png 全表,2026-09-01 复核)。同表 V4-Flash Max 47.8;竞品列 K2.6 Thinking 50, GLM-5.1 Thinking 40.7, Opus-4.6 Max 47.2, GPT-5.4 xHigh 54.6, Gemini-3.1-Pro High 48.8。
DeepSeek-V4-Flash
同公告中的 DeepSeek-V4-Flash 定位为经济档:更小的参数与激活量,推理接近 Pro、世界知识更轻。评测同样覆盖推理/编码/长上下文,亮点如 LiveCodeBench 91.6%、SWE Verified 79.0%。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: MMLU-Pro · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SimpleQA-Verified · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Chinese-SimpleQA · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: GPQA Diamond · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: HLE · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: LiveCodeBench · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Codeforces · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: HMMT 2026 Feb · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: IMOAnswerBench · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Apex · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Apex Shortlist · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: MRCR 1M · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: CorpusQA 1M · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Terminal Bench 2.0 · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SWE Verified · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SWE Pro · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: SWE Multilingual · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: BrowseComp · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: HLE w/tools · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: GDPval-AA · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: MCPAtlas Public · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek-V4-Pro:性能比肩顶级闭源模型 · row: Toolathlon · figure: images/04.png (archive of api-docs.deepseek.com/zh-cn/img/v4-benchmark-2.png, full table)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}DeepSeek-V4-Flash (max effort) 列值——该发布第二模型的同表列(04.webp 全量基准表 DS-V4-Flash Max 列),此前仅记录在 V4-Pro 行 notes 中,2026-09-01 升级为独立 evidence 行。视觉转写自归档图(2026-09-01 复核)。同表竞品列(K2.6 Thinking / GLM-5.1 Thinking / Opus-4.6 Max / GPT-5.4 xHigh / Gemini-3.1-Pro High)见对应 V4-Pro 行 notes。