← 模型目录

DeepSeek-V3.1

DeepSeek · 2025-08-21 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

DeepSeek-V3.1

DeepSeek-V3.1 发布文定位为首个混合推理架构模型:同一模型同时服务 deepseek-chat(非思考)与 deepseek-reasoner(思考),主打思考效率提升与 Agent 能力增强。评测分编码代理、搜索代理与思考效率三组:SWE-bench Verified 66、BrowseComp 30、AIME 2025(V3.1-Think)88.4。

输入模态
文本
上下文
128K
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

swebench 66 模型 deepseek-v3-1 · 版本 Verified · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 编程智能体 · row: SWE · figure: images/02.webp (archive of cdn.deepseek.com v3.1_benchmark_1.webp coding-agent 表) · quote_snippet: 在代码修复测评 SWE 与命令行终端环境下的复杂任务(Terminal-Bench)测试中,DeepSeek-V3.1 相比之前的 DeepSeek 系列模型有明显提高

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose claim verified (SWE named with explicit improvement claim vs prior DeepSeek series); absolute score lives only in image v3.1_benchmark_1.webp. Vision-assisted read (unconfirmed, per goal.md 12.5 not promoted to value): SWE-bench Verified DeepSeek-V3.1 66.0 vs DeepSeek-V3-0324 45.4 vs DeepSeek-R1-0528 44.6. 视觉转写自归档图 images/02.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:45.4 / 44.6。

打开官方来源

terminalbench 31.3 模型 deepseek-v3-1 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 编程智能体 · row: Terminal-Bench · figure: images/02.webp (archive of cdn.deepseek.com v3.1_benchmark_1.webp coding-agent 表) · quote_snippet: 在代码修复测评 SWE 与命令行终端环境下的复杂任务(Terminal-Bench)测试中,DeepSeek-V3.1 相比之前的 DeepSeek 系列模型有明显提高

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose claim verified; score in image only. Vision-assisted read of the archived chart image (confirmed 2026-09-01): Terminal-Bench V3.1 31.3 vs V3-0324 13.3 vs R1-0528 5.7. Version of Terminal-Bench not printed on page. 视觉转写自归档图 images/02.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:13.3 / 5.7。

打开官方来源

browsecomp 30 模型 deepseek-v3-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 搜索智能体 · row: browsecomp · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表) · quote_snippet: 在需要多步推理的复杂搜索测试(browsecomp)与多学科专家级难题测试(HLE)上,DeepSeek-V3.1 性能已大幅领先 R1-0528

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Prose claim verified; score in image only. Vision-assisted read of the archived chart image (confirmed 2026-09-01): Browsecomp V3.1 30.0 vs R1-0528 8.9. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:8.9。

打开官方来源

hlehle 29.8 模型 deepseek-v3-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 搜索智能体 · row: HLE · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表) · quote_snippet: 在需要多步推理的复杂搜索测试(browsecomp)与多学科专家级难题测试(HLE)上,DeepSeek-V3.1 性能已大幅领先 R1-0528

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose claim verified; score in image only. Vision-assisted read of the archived chart image (confirmed 2026-09-01): HLE V3.1 29.8 vs R1-0528 24.8. Tool condition for this row not stated in prose; do not compare against HLE no-tools rows in other releases. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:24.8。

打开官方来源

swebench-multilingual 54.5 模型 deepseek-v3-1 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 编程智能体 · table: benchmark chart image v3.1_benchmark_1.webp · row: SWE-bench Multilingual · figure: images/02.webp (archive of cdn.deepseek.com v3.1_benchmark_1.webp coding-agent 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-multilingual already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 54.5 vs V3-0324 29.3 vs R1-0528 30.5. 视觉转写自归档图 images/02.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:29.3 / 30.5。

打开官方来源

browsecomp-zh 49.2 模型 deepseek-v3-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: Browsecomp_zh · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: browsecomp-zh not yet in data/benchmarks/ (also appears in kimi-k2-thinking.json this batch). Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 49.2 vs R1-0528 35.7. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:35.7。

打开官方来源

xbench-deepsearch 71.2 模型 deepseek-v3-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: xbench-DeepSearch · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: xbench-deepsearch not yet in data/benchmarks/. Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 71.2 vs R1-0528 55.0. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:55.0。

打开官方来源

frames 83.7 模型 deepseek-v3-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: Frames · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: frames not yet in data/benchmarks/ (also in kimi-k2-thinking.json this batch). Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 83.7 vs R1-0528 82.0. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:82.0。

打开官方来源

simpleqa 93.4 模型 deepseek-v3-1 · 版本 未说明 · 指标 correct_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: SimpleQA · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 93.4 vs R1-0528 92.3. Note SimpleQA 'correct' here is much higher than Kimi K2's 31.0 'Correct' - tool/agentic conditions differ across vendors; not comparable. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:92.3。

打开官方来源

seal-0 42.6 模型 deepseek-v3-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 工具调用/智能体支持增强 / 搜索智能体 · table: benchmark chart image v3.1_benchmark_2.webp · row: Seal0 · figure: images/03.webp (archive of cdn.deepseek.com v3.1_benchmark_2.webp search-agent 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: seal-0 not yet in data/benchmarks/ (also in kimi-k2-thinking.json this batch). Vision-assisted read of the archived chart image (confirmed 2026-09-01): V3.1 42.6 vs R1-0528 29.7. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:29.7。

打开官方来源

aime-25 88.4 模型 deepseek-v3-1 · 版本 V3.1-Think, thinking efficiency chart · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 思考效率提升 · table: thinking efficiency chart image v3.1_benchmark_3.webp · row: AIME 2025 · figure: images/04.webp (archive of cdn.deepseek.com v3.1_benchmark_3.webp thinking-efficiency 图) · quote_snippet: V3.1-Think 在输出 token 数减少 20%-50% 的情况下,各项任务的平均表现与 R1-0528 持平

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "output-token comparison: V3.1-Think 15,889 vs R1-0528 22,615 tokens (AIME 2025)",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read of the archived chart image (confirmed 2026-09-01): AIME 2025 V3.1-Think 88.4 at 15,889 output tokens vs R1-0528 87.5 at 22,615 tokens. The chart pairs accuracy with token cost - both dimensions belong to the claim. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:87.5。 V3.1-Think 输出 15,889 tokens,R1-0528 22,615。

打开官方来源

gpqa 80.1 模型 deepseek-v3-1 · 版本 Diamond, V3.1-Think, thinking efficiency chart · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 思考效率提升 · table: thinking efficiency chart image v3.1_benchmark_3.webp · row: GPQA Diamond · figure: images/04.webp (archive of cdn.deepseek.com v3.1_benchmark_3.webp thinking-efficiency 图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "output-token comparison: V3.1-Think 4,122 vs R1-0528 7,678 tokens (GPQA Diamond)",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read of the archived chart image (confirmed 2026-09-01): GPQA Diamond V3.1-Think 80.1 at 4,122 tokens vs R1-0528 81.0 at 7,678 tokens - slightly lower accuracy at ~46% fewer tokens. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:81.0。 V3.1-Think 输出 4,122 tokens,R1-0528 7,678。

打开官方来源

lcb 74.8 模型 deepseek-v3-1 · 版本 V3.1-Think, thinking efficiency chart · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 思考效率提升 · table: thinking efficiency chart image v3.1_benchmark_3.webp · row: LiveCodeBench · figure: images/04.webp (archive of cdn.deepseek.com v3.1_benchmark_3.webp thinking-efficiency 图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": "output-token comparison: V3.1-Think 13,977 vs R1-0528 19,352 tokens (LiveCodeBench)",
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read of the archived chart image (confirmed 2026-09-01): LiveCodeBench V3.1-Think 74.8 at 13,977 tokens vs R1-0528 73.3 at 19,352 tokens. LiveCodeBench version/window not printed. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。同表对照列值:73.3。 V3.1-Think 输出 13,977 tokens,R1-0528 19,352。

打开官方来源