DeepSeek-V3.2-Exp
DeepSeek · 2025-09-29 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
DeepSeek-V3.2-Exp
DeepSeek 将 V3.2-Exp 定位为引入 DeepSeek 稀疏注意力(DSA)、聚焦长上下文训练与推理提效的实验性版本,训练设置与 V3.1-Terminus 严格对齐、公开评测表现基本持平,API 价格同步下调超 50%(页面未给出具体单价)。已收录评测覆盖通用知识、数学推理、代码与中英文检索等领域,Codeforces Div1 2121、AIME 2025 89.3。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp (V3.1-Terminus vs V3.2-Exp) · row: MMLU-Pro · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表) · quote_snippet: DeepSeek-V3.2-Exp 的表现与 V3.1-Terminus 基本持平
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page states training setup strictly aligned with V3.1-Terminus to isolate the sparse-attention effect. Vision-assisted read of the archived image (confirmed 2026-09-01): MMLU-Pro V3.2-Exp 85.0 vs V3.1-Terminus 85.0. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 85。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: GPQA-Diamond · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): GPQA-Diamond V3.2-Exp 79.9 vs V3.1-Terminus 80.7. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 80.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: Humanity's Last Exam · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): HLE V3.2-Exp 19.8 vs V3.1-Terminus 21.7. Tool condition not printed; GLM-4.6's chart lists DeepSeek-V3.2-Exp HLE at 19.8 (cross-vendor consistency check passed at value level). 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 21.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: BrowseComp · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read of the archived image (confirmed 2026-09-01): BrowseComp V3.2-Exp 40.1 vs V3.1-Terminus 38.5. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 38.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: BrowseComp-zh · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp-zh not yet in data/benchmarks/. Vision-assisted read of the archived image (confirmed 2026-09-01): BrowseComp-zh V3.2-Exp 47.9 vs V3.1-Terminus 45.0. Matches the 47.9 printed in Kimi K2 Thinking's DeepSeek-V3.2 column. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 45。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: SimpleQA · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): SimpleQA V3.2-Exp 97.1 vs V3.1-Terminus 96.8. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 96.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: LiveCodeBench · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): LiveCodeBench V3.2-Exp 74.1 vs V3.1-Terminus 74.9. Version window not printed. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 74.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: Codeforces-Div1 · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}codeforces id already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read of the archived image (confirmed 2026-09-01): Codeforces-Div1 Elo V3.2-Exp 2121 vs V3.1-Terminus 2046. Elo rating, not a percentage - never aggregate with percent rows. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 2046。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: Aider-Polyglot · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): Aider-Polyglot V3.2-Exp 74.5 vs V3.1-Terminus 76.1. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 76.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: SWE Verified · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): SWE Verified V3.2-Exp 67.8 vs V3.1-Terminus 68.4. Matches the 67.8 printed in Kimi K2 Thinking and GLM-4.6 (image) DeepSeek-V3.2 columns. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 68.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: SWE-bench Multilingual · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read of the archived image (confirmed 2026-09-01): V3.2-Exp 57.9 vs V3.1-Terminus 57.8. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 57.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: Terminal-bench · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): Terminal-bench V3.2-Exp 37.7 vs V3.1-Terminus 36.7. Matches the 37.7 printed in Kimi K2 Thinking's DeepSeek-V3.2 column. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 36.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: AIME 2025 · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): AIME 2025 V3.2-Exp 89.3 vs V3.1-Terminus 88.4. Matches the 89.3 printed in Kimi K2 Thinking's DeepSeek-V3.2 column. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 88.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: DeepSeek Sparse Attention(DSA)稀疏注意力机制 · table: benchmark comparison chart image v3_2_benchmark.webp · row: HMMT 2025 · figure: images/03.webp (archive of api-docs v3_2_benchmark.webp 对比表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hmmt-25 already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read of the archived image (confirmed 2026-09-01): HMMT 2025 V3.2-Exp 83.6 vs V3.1-Terminus 86.1. 视觉转写自归档图 images/03.webp(api-docs v3_2_benchmark.webp 对比表,2026-09-01 复核,与先前读数一致)。同表 V3.1-Terminus 对照列 86.1。