DeepSeek-V3.2 / DeepSeek-V3.2-Speciale
DeepSeek · 2025-12-01 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
DeepSeek-V3.2
DeepSeek 官方将 V3.2 正式版定位为「强化 Agent 能力、融入思考推理」,是 DeepSeek 首个将思考与工具调用融合的模型。已收录 26 项评测覆盖数学竞赛、代码、知识与工具/Agent 使用:亮点 AIME 2025(Thinking)93.1、Codeforces(Thinking)2386。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: AIME 2025 · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read of the archived image (confirmed 2026-09-01): V3.2 Thinking 93.1 (~16k tokens); Speciale 96.0 (~23k). Columns: GPT-5 High 94.6 (13k), Gemini-3.0 Pro 95.0 (15k), Kimi-K2 Thinking 94.5 (24k). Token cost printed per cell in parentheses - part of the claim. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 16k。 同表 Speciale 列 96(23k)。 竞品列:GPT-5 High 94.6(13k), Gemini-3.0 Pro 95.0(15k), Kimi-K2 Thinking 94.5(24k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: HMMT Feb 2025 · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 92.5 (19k); Speciale 99.2 (27k); GPT-5 88.3, Gemini 97.5, K2 Thinking 89.4. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 19k。 同表 Speciale 列 99.2(27k)。 竞品列:GPT-5 High 88.3(16k), Gemini-3.0 Pro 97.5(16k), Kimi-K2 Thinking 89.4(31k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: HMMT Nov 2025 · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 90.2 (18k); Speciale 94.4 (25k). Matches GLM-5/GLM-5.1 competitor columns (90.2). 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 18k。 同表 Speciale 列 94.4(25k)。 竞品列:GPT-5 High 89.2(20k), Gemini-3.0 Pro 93.3(15k), Kimi-K2 Thinking 89.2(29k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: IMOAnswerBench · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 78.3 (27k); Speciale 84.5 (45k). Matches GLM-5.1 competitor column (78.3). 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 27k。 同表 Speciale 列 84.5(45k)。 竞品列:GPT-5 High 76.0(31k), Gemini-3.0 Pro 83.3(18k), Kimi-K2 Thinking 78.6(37k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: LiveCodeBench · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 83.3 (16k); Speciale 88.7 (27k). Version window not printed. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 16k。 同表 Speciale 列 88.7(27k)。 竞品列:GPT-5 High 84.5(13k), Gemini-3.0 Pro 90.7(13k), Kimi-K2 Thinking 82.6(29k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: CodeForces · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 2386 (42k tokens); Speciale 2701 (77k); Gemini-3.0 Pro 2708. Elo rating, not a percentage. 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 42k。 同表 Speciale 列 2701(77k)。 竞品列:GPT-5 High 2537(29k), Gemini-3.0 Pro 2708(22k), Kimi-K2 Thinking '-'(Kimi-K2-Thinking 列在表 1 该行为 '-')。 汇总图 images/02.webp 同读数 2386 / 2701,两图一致。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: GPQA Diamond · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 82.4 (7k); Speciale 85.7 (16k). Matches GLM-5/5.1 competitor columns (82.4). 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 7k。 同表 Speciale 列 85.7(16k)。 竞品列:GPT-5 High 85.7(8k), Gemini-3.0 Pro 91.9(8k), Kimi-K2 Thinking 84.5(12k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: HLE · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 25.1 (21k); Speciale 30.6 (35k). Matches GLM-4.7/5/5.1 competitor columns (25.1). 视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。V3.2 单元格含 token 成本 21k。 同表 Speciale 列 30.6(35k)。 竞品列:GPT-5 High 26.3(15k), Gemini-3.0 Pro 37.7(15k), Kimi-K2 Thinking 23.9(24k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: tau2-Bench (Table 2 agentic) · figure: images/04.webp (archive of api-docs v3.2_251201 Table 2 ToolUse/agentic)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision read of the archived image (confirmed 2026-09-01): V3.2 Thinking 80.3. Columns: Sonnet-4.5 84.7, GPT-5 80.2, Gemini-3.0 Pro 85.4, K2 Thinking 74.3, MiniMax M2 76.9. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。 竞品列:Claude-4.5-Sonnet 84.7, GPT-5 High 80.2, Gemini-3.0 Pro 85.4, Kimi-K2 Thinking 74.3, MiniMax M2 76.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: MCP-Universe · figure: images/04.webp (archive of api-docs v3.2_251201 Table 2 ToolUse/agentic)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mcp-universe not yet in data/benchmarks/. Vision read of the archived image (confirmed 2026-09-01): V3.2 45.9. Columns: Sonnet-4.5 46.5, GPT-5 47.9, Gemini 50.7, K2 Thinking 35.6, M2 29.4. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。 竞品列:Claude-4.5-Sonnet 46.5, GPT-5 High 47.9, Gemini-3.0 Pro 50.7, Kimi-K2 Thinking 35.6, MiniMax M2 29.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: MCP-Mark · figure: images/04.webp (archive of api-docs v3.2_251201 Table 2 ToolUse/agentic)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mcp-mark not yet in data/benchmarks/. Vision read of the archived image (confirmed 2026-09-01): V3.2 38.0. Columns: Sonnet-4.5 33.3, GPT-5 50.9, Gemini 43.1, K2 Thinking 20.4, M2 24.4. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。 竞品列:Claude-4.5-Sonnet 33.3, GPT-5 High 50.9, Gemini-3.0 Pro 43.1, Kimi-K2 Thinking 20.4, MiniMax M2 24.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: Tool-Decathlon (Table 2) · figure: images/04.webp (archive of api-docs v3.2_251201 Table 2 ToolUse/agentic)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tool-decathlon. Vision read of the archived image (confirmed 2026-09-01): V3.2 35.2 (matches GLM-5/5.1 competitor columns). Columns: Sonnet-4.5 38.6, GPT-5 29.0, Gemini 36.4, K2 Thinking 17.6, M2 16.0. 视觉转写自归档图 images/04.webp(2026-09-01 复核,与先前读数一致)。 竞品列:Claude-4.5-Sonnet 38.6, GPT-5 High 29.0, Gemini-3.0 Pro 36.4, Kimi-K2 Thinking 17.6, MiniMax M2 16.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 智能体能力对比(hero chart Agentic Capabilities 组) · row: SWE Verified (Resolved) · figure: images/02.webp (archive of api-docs v3.2_251201 benchmark.webp hero chart, Agentic Capabilities group)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}仅出现在 hero chart(images/02.webp),Table 1/Table 2 未含此行。同一数值被 kimi-k2-6 发布页附录表转引(DeepSeek V3.2 SWE-Bench Verified 73.1),跨厂商一致,佐证读数。同组柱值:GPT-5 High 74.9, Claude-4.5-Sonnet 77.2, Gemini-3.0 Pro 76.2;Speciale 无 SWE 柱(研究型 API 无工具调用)。视觉转写自归档图(2026-09-01)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 智能体能力对比(hero chart Agentic Capabilities 组) · row: Terminal Bench 2.0 (Acc) · figure: images/02.webp (archive of api-docs v3.2_251201 benchmark.webp hero chart, Agentic Capabilities group)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}仅出现在 hero chart(images/02.webp)。同一数值被 kimi-k2-6 发布页附录表转引(DeepSeek V3.2 Terminal-Bench 2.0 46.4),跨厂商一致,佐证读数。同组柱值:GPT-5 High 35.2, Claude-4.5-Sonnet 42.8, Gemini-3.0 Pro 54.2(柱序按图例,灰阶柱读数)。视觉转写自归档图(2026-09-01)。
DeepSeek-V3.2-Speciale
DeepSeek-V3.2-Speciale 定位为长思考增强叠加 DeepSeek-Math-V2 定理证明的研究用临时 API(无工具调用)。已收录 Speciale 行集中于数学与竞赛推理:亮点 HMMT 2025 年 2 月卷 99.2、Codeforces 2701,另在 IMO/CMO/ICPC/IOI 2025 取得金牌级结论。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: IMO 2025 · quote_snippet: V3.2-Speciale 模型成功斩获 IMO 2025(国际数学奥林匹克)…金牌
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: imo-2025 already introduced by prior batches (M3), still not in data/benchmarks/. Medal-level outcome (not a percentage); official-prose claim verified.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: CMO 2025 · quote_snippet: CMO 2025(中国数学奥林匹克)…金牌
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cmo-2025 not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: ICPC World Finals 2025 · quote_snippet: ICPC 与 IOI 成绩分别达到了人类选手第二名与第十名的水平
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: icpc-world-finals-2025 not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 推理能力 / 思考融入工具调用 · row: IOI 2025
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ioi-2025 not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: AIME 2025 · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~23k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。竞品列:GPT-5 High 94.6(13k), Gemini-3.0 Pro 95.0(15k), Kimi-K2 Thinking 94.5(24k)。V3.2 Thinking 同行 93.1(16k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: HMMT Feb 2025 · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~27k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。竞品列:GPT-5 High 88.3(16k), Gemini-3.0 Pro 97.5(16k), Kimi-K2 Thinking 89.4(31k)。V3.2 Thinking 同行 92.5(19k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: HMMT Nov 2025 · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~25k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。竞品列:GPT-5 High 89.2(20k), Gemini-3.0 Pro 93.3(15k), Kimi-K2 Thinking 89.2(29k)。V3.2 Thinking 同行 90.2(18k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: IMOAnswerBench · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~45k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。竞品列:GPT-5 High 76.0(31k), Gemini-3.0 Pro 83.3(18k), Kimi-K2 Thinking 78.6(37k)。V3.2 Thinking 同行 78.3(27k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: LiveCodeBench · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~27k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。竞品列:GPT-5 High 84.5(13k), Gemini-3.0 Pro 90.7(13k), Kimi-K2 Thinking 82.6(29k)。V3.2 Thinking 同行 83.3(16k)。版本窗口未印出。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: CodeForces · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~77k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。Elo rating,非百分比。竞品列:GPT-5 High 2537(29k), Gemini-3.0 Pro 2708(22k), Kimi-K2 Thinking 未报告('-')。V3.2 Thinking 同行 2386(42k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: GPQA Diamond · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~16k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。竞品列:GPT-5 High 85.7(8k), Gemini-3.0 Pro 91.9(8k), Kimi-K2 Thinking 84.5(12k)。V3.2 Thinking 同行 82.4(7k)。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 思考能力对比(Table 1) · table: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 reasoning; 单元格括号内为 token 成本) · row: HLE · figure: images/03.webp (archive of api-docs v3.2_251201_benchmark_table_cn.webp, Table 1 Speciale 列)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": "单元格括号内 token 成本约 ~35k",
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}V3.2-Speciale 列值(该发布第二个模型;此前仅记录在 V3.2 Thinking 行的 notes 中,2026-09-01 升级为独立 evidence 行)。视觉转写自归档图 images/03.webp(2026-09-01 复核,与先前读数一致)。竞品列:GPT-5 High 26.3(15k), Gemini-3.0 Pro 37.7(15k), Kimi-K2 Thinking 23.9(24k)。V3.2 Thinking 同行 25.1(21k)。