← 模型目录

DeepSeek-V4.1-Flash

DeepSeek · 2026-09-10 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

DeepSeek-V4.1-Flash

DeepSeek 新一代非对称 Causal-Encoder-Decoder (CED) 架构的首款多模态 MoE 模型,总参 552B 仅激活 8B 输入与 16B 输出,KV Cache 显存缩减至 1/4;在 Terminal-Bench 2.1 达到 90.6%、GPQA Diamond 90.9%、DeepSWE v1.1 74.2% 与 CyberGym 88.1%,全面替代并下线 V4-Pro。

输入模态
文本 / 图像 / 代码
上下文
1M
参数
552B (8B prefill / 16B decode active)
价格(每百万 tokens)
USD 输入 0.15 / 输出 0.3 · 高峰时段 $0.15/$0.30 每百万 tokens,闲时(00:30-08:30 UTC)减半至 $0.075/$0.15,缓存命中 $0.003-$0.006

本变体的评测证据

terminalbench 90.6% 模型 deepseek-v4-1-flash · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Benchmark Results · row: Terminal-Bench 2.1

{
  "harness": "Terminus-2",
  "tools": [
    "bash"
  ],
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Hugging Face 官方模型卡明文大表转录,Terminal-Bench 2.1 报告 90.6%。

打开官方来源

gpqa 90.9% 模型 deepseek-v4-1-flash · 版本 Diamond · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Benchmark Results · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Hugging Face 官方模型卡明文大表转录,GPQA Diamond 达到 90.9%。

打开官方来源

deepswe 74.2% 模型 deepseek-v4-1-flash · 版本 1.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Benchmark Results · row: DeepSWE v1.1

{
  "harness": "mini-swe-agent",
  "tools": [
    "bash",
    "editor"
  ],
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Hugging Face 官方模型卡明文大表转录,DeepSWE v1.1 报告 74.2%,仓库级代码编辑能力显著提高。

打开官方来源

cybergym 88.1% 模型 deepseek-v4-1-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Benchmark Results · row: CyberGym

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Hugging Face 官方模型卡明文大表转录,网络安全攻防沙箱 CyberGym 达到 88.1%。

打开官方来源

apex 65.6% 模型 deepseek-v4-1-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Benchmark Results · row: MathArena Apex

{
  "harness": "MathArena-aligned",
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: matharena-apex not yet in data/benchmarks/. Hugging Face 模型卡明文大表转录,Apex 竞赛高难度数学基准得分 65.6%。

打开官方来源

hlehle 63.9% 模型 deepseek-v4-1-flash · 版本 with-tools · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Benchmark Results · row: HLE (w/ tools)

{
  "harness": null,
  "tools": [
    "python",
    "search"
  ],
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

HLE 人类最后考试带工具评测,得分 63.9%。

打开官方来源

automationbench 54.8% 模型 deepseek-v4-1-flash · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation Results · table: Benchmark Results · row: Automation-Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Automation-Bench 长程自动化基准得分 54.8%。

打开官方来源