Devstral 2 / Devstral Small 2
Mistral AI · 2025-12-09 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Devstral 2
Mistral 将 Devstral 2 定位为面向智能体编码的下一代开源 SOTA 编码模型(modified MIT 协议),并随发布配套推出 Mistral Vibe CLI。已收录评测为软件工程与人评胜率两项,SWE-bench Verified 72.2%、人评 42.8% 胜 / 28.6% 负。
- 输入模态
- 文本
- 上下文
- 256K
- 参数
- 123B(稠密架构)
- 价格(每百万 tokens)
- USD 输入 0.4 / 输出 2 · API 当前处于免费期;官方页明示免费期后为 $0.40/$2.00 每百万 tokens(输入/输出)
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Devstral: the next generation of SOTA coding. · quote_snippet: It reaches 72.2% on SWE-bench Verified—establishing it as one of the best open-weight models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark swebench (canonical name is SWE-bench Verified). Also stated in the Highlights list with the same value. Competitor SWE-bench values appear only in the chart image (open weights vs proprietary), not in text, so no comparison_cited rows were created.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Built for production-grade workflows. · quote_snippet: Devstral 2 shows a clear advantage over DeepSeek V3.2, with a 42.8% win rate versus 28.6% loss rate
{
"harness": "Cline",
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "human preference win/loss rates",
"judge": "human evaluations conducted by an independent annotation provider"
}new-benchmark: devstral-2-human-winrate not yet in data/benchmarks/. Vendor-commissioned human preference evaluation (42.8% win and 28.6% loss are both Devstral 2's own rates against DeepSeek V3.2, not DeepSeek scores, so no separate comparison_cited row). Same section states Claude Sonnet 4.5 'remains significantly preferred' with no numbers printed. Task set, sample size, and eval duration are not stated on the page.
Devstral Small 2
Mistral 将 Devstral Small 2 定位为与 Devstral 2 同为 256K 上下文、以 Apache 2.0 开源的紧凑型编码模型,规模使其可在消费级硬件上本地快速推理。已收录评测为 SWE-bench Verified 68.0%。
- 输入模态
- 文本 / 图像
- 上下文
- 256K
- 参数
- 24B
- 价格(每百万 tokens)
- USD 输入 0.1 / 输出 0.3 · 官方页明示 API 免费期后为 $0.10/$0.30 每百万 tokens(输入/输出)
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Devstral: the next generation of SOTA coding. · quote_snippet: Devstral Small 2 scores 68.0% on SWE-bench Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark swebench. Companion claim 'places firmly among models up to five times its size' is a size-relative statement, not a printed competitor score.