DeepSeek-V4.1-Flash
DeepSeek · 2026-09-10 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
DeepSeek-V4.1-Flash
DeepSeek 新一代非对称 Causal-Encoder-Decoder (CED) 架构的首款多模态 MoE 模型,总参 552B 仅激活 8B 输入与 16B 输出,KV Cache 显存缩减至 1/4;在 Terminal-Bench 2.1 达到 90.6%、GPQA Diamond 90.9%、DeepSWE v1.1 74.2% 与 CyberGym 88.1%,全面替代并下线 V4-Pro。
- 输入模态
- 文本 / 图像 / 代码
- 上下文
- 1M
- 参数
- 552B (8B prefill / 16B decode active)
- 价格(每百万 tokens)
- USD 输入 0.15 / 输出 0.3 · 高峰时段 $0.15/$0.30 每百万 tokens,闲时(00:30-08:30 UTC)减半至 $0.075/$0.15,缓存命中 $0.003-$0.006
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Benchmark Results · row: Terminal-Bench 2.1
{
"harness": "Terminus-2",
"tools": [
"bash"
],
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Hugging Face 官方模型卡明文大表转录,Terminal-Bench 2.1 报告 90.6%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Benchmark Results · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Hugging Face 官方模型卡明文大表转录,GPQA Diamond 达到 90.9%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Benchmark Results · row: DeepSWE v1.1
{
"harness": "mini-swe-agent",
"tools": [
"bash",
"editor"
],
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Hugging Face 官方模型卡明文大表转录,DeepSWE v1.1 报告 74.2%,仓库级代码编辑能力显著提高。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Benchmark Results · row: CyberGym
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Hugging Face 官方模型卡明文大表转录,网络安全攻防沙箱 CyberGym 达到 88.1%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Benchmark Results · row: MathArena Apex
{
"harness": "MathArena-aligned",
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: matharena-apex not yet in data/benchmarks/. Hugging Face 模型卡明文大表转录,Apex 竞赛高难度数学基准得分 65.6%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Benchmark Results · row: HLE (w/ tools)
{
"harness": null,
"tools": [
"python",
"search"
],
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}HLE 人类最后考试带工具评测,得分 63.9%。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results · table: Benchmark Results · row: Automation-Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Automation-Bench 长程自动化基准得分 54.8%。