← 模型目录

DeepSeek-V4-Flash-Vision-Exp

DeepSeek · 2026-08-21 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

DeepSeek-V4-Flash-Vision-Exp

DeepSeek 将 V4-Flash-Vision-Exp 定位为开启多模态 API 服务的实验性多模态模型,官方称其纯文本能力与 V4-Flash 持平、向 Opus-4.8 收敛多模态 Agent 差距。已收录 11 项评测(文本 Agent 与多模态 Agent 两组):亮点 Terminal-Bench 2.1 83.9、CyberGym 75.3。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

terminalbench 83.9 模型 deepseek-v4-flash-vision-exp · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 · table: benchmark chart image v4_260821_benchmark_cn.png, section 文本 Agent 能力评测 · row: Terminal Bench 2.1 · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png) · quote_snippet: DeepSeek 系列模型使用 DeepSeek Harness 极简模式作为框架进行测试

{
  "harness": "DeepSeek Harness (minimal mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Score values exist only inside the chart image. Vision-assisted OCR of the figure read 83.9 for V4-Flash-Vision-Exp (vs 82.7 V4-Flash-0731, 85.0 Opus-4.8); Confirmed 2026-09-01 by independent re-read of the archived image. Protocol applies to public Code Agent text tasks per the page caption. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 82.7,Opus-4.8 85。

打开官方来源

nl2repo 57.7 模型 deepseek-v4-flash-vision-exp · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 · table: benchmark chart image v4_260821_benchmark_cn.png, section 文本 Agent 能力评测 · row: NL2Repo · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": "DeepSeek Harness (minimal mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: nl2repo not yet in data/benchmarks.json. OCR of figure read 57.7 (vs 54.2 V4-Flash-0731, 69.7 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 54.2,Opus-4.8 69.7。

打开官方来源

cybergym 75.3 模型 deepseek-v4-flash-vision-exp · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 · table: benchmark chart image v4_260821_benchmark_cn.png, section 文本 Agent 能力评测 · row: Cybergym · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": "DeepSeek Harness (minimal mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cybergym not yet in data/benchmarks.json. OCR of figure read 75.3 (vs 76.7 V4-Flash-0731, 78.3 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 76.7,Opus-4.8 78.3。

打开官方来源

deepswe 59.3 模型 deepseek-v4-flash-vision-exp · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 · table: benchmark chart image v4_260821_benchmark_cn.png, section 文本 Agent 能力评测 · row: DeepSWE · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": "DeepSeek Harness (minimal mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepswe not yet in data/benchmarks/. OCR of figure read 59.3 (vs 54.4 V4-Flash-0731, 58.0 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. DeepSeek's harness for this row differs from the mini-SWE-agent protocol used by the DeepSWE leaderboard and from Kimi/G(Code) harness runs; not cross-comparable. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 54.4,Opus-4.8 58。

打开官方来源

toolathlon 75.9 模型 deepseek-v4-flash-vision-exp · 版本 Verified · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 · table: benchmark chart image v4_260821_benchmark_cn.png, section 文本 Agent 能力评测 · row: Toolathlon-Verified · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": "DeepSeek Harness (minimal mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: toolathlon not yet in data/benchmarks.json. OCR of figure read 75.9 (vs 70.3 V4-Flash-0731, 76.2 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 70.3,Opus-4.8 76.2。

打开官方来源

dsbench-hard 63.6 模型 deepseek-v4-flash-vision-exp · 版本 Hard · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 · table: benchmark chart image v4_260821_benchmark_cn.png, section 文本 Agent 能力评测 · row: DSBench-Hard · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": "DeepSeek Harness (minimal mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: dsbench-hard not yet in data/benchmarks.json. OCR of figure read 63.6 (vs 59.6 V4-Flash-0731, 71.7 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 59.6,Opus-4.8 71.7。

打开官方来源

automationbench 25.7 模型 deepseek-v4-flash-vision-exp · 版本 public subset · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 · table: benchmark chart image v4_260821_benchmark_cn.png, section 文本 Agent 能力评测 · row: AutomationBench (Public) · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": "DeepSeek Harness (minimal mode)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": 1,
  "top_p": 0.95,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: automationbench not yet in data/benchmarks/. OCR of figure read 25.7 (vs 25.1 V4-Flash-0731, 27.2 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. Subset version not stated on this page (other releases cite v1.0.6 or 600-task public subset); do not compare across subset versions. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 25.1,Opus-4.8 27.2。

打开官方来源

apexbench 36.5 模型 deepseek-v4-flash-vision-exp · 版本 Pass@1 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 / 多模态 Agent 能力评测 · table: benchmark chart image v4_260821_benchmark_cn.png, section 多模态 Agent 能力评测 · row: ApexBench (Pass@1) · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png) · quote_snippet: 文本模型 DeepSeek-V4-Flash 会忽略其中的多模态元素

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@1",
  "judge": null
}

new-benchmark: apexbench not yet in data/benchmarks.json. Page text confirms the benchmark name and notes that the text-only model ignores multimodal elements on ApexBench; score itself is image-only (OCR read 36.5, confirmed 2026-09-01). 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 26.2(带 ** 注),Opus-4.8 39.4。

打开官方来源

agents-last-exam 27.3 模型 deepseek-v4-flash-vision-exp · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 平衡文本与多模态能力 / 多模态 Agent 能力评测 · table: benchmark chart image v4_260821_benchmark_cn.png, section 多模态 Agent 能力评测 · row: Agents' Last Exam · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png) · quote_snippet: 在 ApexBench 与 Agents' Last Exam 测评中,文本模型 ... 会忽略其中的多模态元素

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: agents-last-exam not yet in data/benchmarks.json. OCR of figure read 27.3 (vs 25.2 V4-Flash-0731, 25.7 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 25.2(带 ** 注),Opus-4.8 25.7。

打开官方来源

chartography 64.3 模型 deepseek-v4-flash-vision-exp · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 多模态 Agent 能力评测 · table: benchmark chart image v4_260821_benchmark_cn.png, section 多模态 Agent 能力评测 · row: Chartography · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: chartography not yet in data/benchmarks.json. OCR of figure read 64.3 (V4-Flash-0731 not run, 65.0 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 -,Opus-4.8 65。

打开官方来源

zerobench 35 模型 deepseek-v4-flash-vision-exp · 版本 Pass@5 · 指标 pass@5 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 多模态 Agent 能力评测 · table: benchmark chart image v4_260821_benchmark_cn.png, section 多模态 Agent 能力评测 · row: ZeroBench (Pass@5) · figure: images/02.png (archive of api-docs.deepseek.com/zh-cn/img/v4_260821_benchmark_cn.png)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass@5",
  "judge": null
}

new-benchmark: zerobench not yet in data/benchmarks/. OCR of figure read 35.0 (V4-Flash-0731 not run, 34.0 Opus-4.8); confirmed 2026-09-01 by independent re-read of the archived image. pass@5 aggregation; not comparable to pass@1 rows. 视觉转写自归档图 images/02.png(api-docs v4_260821_benchmark_cn.png 表格,2026-09-01 复核,与先前 OCR 一致)。同表 DeepSeek V4-Flash-0731 -,Opus-4.8 34。

打开官方来源