DeepSeek-V3.1
DeepSeek · 2025-08-21 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
DeepSeek-V3.1
DeepSeek-V3.1 发布文定位为首个混合推理架构模型:同一模型同时服务 deepseek-chat(非思考)与 deepseek-reasoner(思考),主打思考效率提升与 Agent 能力增强。评测分编码代理、搜索代理与思考效率三组:SWE-bench Verified 66、BrowseComp 30、AIME 2025(V3.1-Think)88.4。
- 输入模态
- 文本
- 上下文
- 128K
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 编程智能体 · row: SWE · figure: images/02.webp (archive of cdn.deepseek.com v3.1_benchmark_1.webp coding-agent 表) · quote_snippet: 在代码修复测评 SWE 与命令行终端环境下的复杂任务(Terminal-Bench)测试中,DeepSeek-V3.1 相比之前的 DeepSeek 系列模型有明显提高
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose claim verified (SWE named with explicit improvement claim vs prior DeepSeek series); absolute score lives only in image v3.1_benchmark_1.webp. Vision-assisted read (unconfirmed, per goal.md 12.5 not promoted to value): SWE-bench Verified DeepSeek-V3.1 66.0 vs DeepSeek-V3-0324 45.4 vs DeepSeek-R1-0528 44.6. 视觉转写自归档图 images/02.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:45.4 / 44.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 编程智能体 · row: Terminal-Bench · figure: images/02.webp (archive of cdn.deepseek.com v3.1_benchmark_1.webp coding-agent 表) · quote_snippet: 在代码修复测评 SWE 与命令行终端环境下的复杂任务(Terminal-Bench)测试中,DeepSeek-V3.1 相比之前的 DeepSeek 系列模型有明显提高
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose claim verified; score in image only. Vision-assisted read of the archived chart image (confirmed 2026-09-01): Terminal-Bench V3.1 31.3 vs V3-0324 13.3 vs R1-0528 5.7. Version of Terminal-Bench not printed on page. 视觉转写自归档图 images/02.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:13.3 / 5.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 搜索智能体 · row: browsecomp · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表) · quote_snippet: 在需要多步推理的复杂搜索测试(browsecomp)与多学科专家级难题测试(HLE)上,DeepSeek-V3.1 性能已大幅领先 R1-0528
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Prose claim verified; score in image only. Vision-assisted read of the archived chart image (confirmed 2026-09-01): Browsecomp V3.1 30.0 vs R1-0528 8.9. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:8.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 搜索智能体 · row: HLE · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表) · quote_snippet: 在需要多步推理的复杂搜索测试(browsecomp)与多学科专家级难题测试(HLE)上,DeepSeek-V3.1 性能已大幅领先 R1-0528
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose claim verified; score in image only. Vision-assisted read of the archived chart image (confirmed 2026-09-01): HLE V3.1 29.8 vs R1-0528 24.8. Tool condition for this row not stated in prose; do not compare against HLE no-tools rows in other releases. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:24.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 编程智能体 · table: benchmark chart image v3.1_benchmark_1.webp · row: SWE-bench Multilingual · figure: images/02.webp (archive of cdn.deepseek.com v3.1_benchmark_1.webp coding-agent 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 54.5 vs V3-0324 29.3 vs R1-0528 30.5. 视觉转写自归档图 images/02.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:29.3 / 30.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: Browsecomp_zh · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp-zh not yet in data/benchmarks/ (also appears in kimi-k2-thinking.json this batch). Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 49.2 vs R1-0528 35.7. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:35.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: xbench-DeepSearch · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: xbench-deepsearch not yet in data/benchmarks/. Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 71.2 vs R1-0528 55.0. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:55.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: Frames · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: frames not yet in data/benchmarks/ (also in kimi-k2-thinking.json this batch). Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 83.7 vs R1-0528 82.0. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:82.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: SimpleQA · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 93.4 vs R1-0528 92.3. Note SimpleQA 'correct' here is much higher than Kimi K2's 31.0 'Correct' - tool/agentic conditions differ across vendors; not comparable. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:92.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: Seal0 · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: seal-0 not yet in data/benchmarks/ (also in kimi-k2-thinking.json this batch). Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 42.6 vs R1-0528 29.7. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:29.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考效率提升 · table: thinking efficiency chart image v3.1_benchmark_3.webp · row: AIME 2025 · figure: images/04.webp (archive of cdn.deepseek.com v3.1_benchmark_3.webp thinking-efficiency 图) · quote_snippet: V3.1-Think 在输出 token 数减少 20%-50% 的情况下,各项任务的平均表现与 R1-0528 持平
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "output-token comparison: V3.1-Think 15,889 vs R1-0528 22,615 tokens (AIME 2025)",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived chart image (confirmed 2026-09-01): AIME 2025 V3.1-Think 88.4 at 15,889 output tokens vs R1-0528 87.5 at 22,615 tokens. The chart pairs accuracy with token cost - both dimensions belong to the claim. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:87.5。 V3.1-Think 输出 15,889 tokens,R1-0528 22,615。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考效率提升 · table: thinking efficiency chart image v3.1_benchmark_3.webp · row: GPQA Diamond · figure: images/04.webp (archive of cdn.deepseek.com v3.1_benchmark_3.webp thinking-efficiency 图)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "output-token comparison: V3.1-Think 4,122 vs R1-0528 7,678 tokens (GPQA Diamond)",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived chart image (confirmed 2026-09-01): GPQA Diamond V3.1-Think 80.1 at 4,122 tokens vs R1-0528 81.0 at 7,678 tokens - slightly lower accuracy at ~46% fewer tokens. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:81.0。 V3.1-Think 输出 4,122 tokens,R1-0528 7,678。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考效率提升 · table: thinking efficiency chart image v3.1_benchmark_3.webp · row: LiveCodeBench · figure: images/04.webp (archive of cdn.deepseek.com v3.1_benchmark_3.webp thinking-efficiency 图)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "output-token comparison: V3.1-Think 13,977 vs R1-0528 19,352 tokens (LiveCodeBench)",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived chart image (confirmed 2026-09-01): LiveCodeBench V3.1-Think 74.8 at 13,977 tokens vs R1-0528 73.3 at 19,352 tokens. LiveCodeBench version/window not printed. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:73.3。 V3.1-Think 输出 13,977 tokens,R1-0528 19,352。