← 模型目录

Devstral 2 / Devstral Small 2

Mistral AI · 2025-12-09 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Devstral 2

Mistral 将 Devstral 2 定位为面向智能体编码的下一代开源 SOTA 编码模型(modified MIT 协议),并随发布配套推出 Mistral Vibe CLI。已收录评测为软件工程与人评胜率两项,SWE-bench Verified 72.2%、人评 42.8% 胜 / 28.6% 负。

输入模态
文本
上下文
256K
参数
123B(稠密架构)
价格(每百万 tokens)
USD 输入 0.4 / 输出 2 · API 当前处于免费期;官方页明示免费期后为 $0.40/$2.00 每百万 tokens(输入/输出)

本变体的评测证据

swebench 72.2% 模型 devstral-2 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Devstral: the next generation of SOTA coding. · quote_snippet: It reaches 72.2% on SWE-bench Verified—establishing it as one of the best open-weight models

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark swebench (canonical name is SWE-bench Verified). Also stated in the Highlights list with the same value. Competitor SWE-bench values appear only in the chart image (open weights vs proprietary), not in text, so no comparison_cited rows were created.

打开官方来源

devstral-2-human-winrate 42.8% win / 28.6% loss 模型 devstral-2 · 版本 未说明 · 指标 human_win_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Built for production-grade workflows. · quote_snippet: Devstral 2 shows a clear advantage over DeepSeek V3.2, with a 42.8% win rate versus 28.6% loss rate

{
  "harness": "Cline",
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "human preference win/loss rates",
  "judge": "human evaluations conducted by an independent annotation provider"
}

new-benchmark: devstral-2-human-winrate not yet in data/benchmarks/. Vendor-commissioned human preference evaluation (42.8% win and 28.6% loss are both Devstral 2's own rates against DeepSeek V3.2, not DeepSeek scores, so no separate comparison_cited row). Same section states Claude Sonnet 4.5 'remains significantly preferred' with no numbers printed. Task set, sample size, and eval duration are not stated on the page.

打开官方来源

Devstral Small 2

Mistral 将 Devstral Small 2 定位为与 Devstral 2 同为 256K 上下文、以 Apache 2.0 开源的紧凑型编码模型,规模使其可在消费级硬件上本地快速推理。已收录评测为 SWE-bench Verified 68.0%。

输入模态
文本 / 图像
上下文
256K
参数
24B
价格(每百万 tokens)
USD 输入 0.1 / 输出 0.3 · 官方页明示 API 免费期后为 $0.10/$0.30 每百万 tokens(输入/输出)

本变体的评测证据

swebench 68.0% 模型 devstral-small-2 · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Devstral: the next generation of SOTA coding. · quote_snippet: Devstral Small 2 scores 68.0% on SWE-bench Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark swebench. Companion claim 'places firmly among models up to five times its size' is a size-relative statement, not a printed competitor score.

打开官方来源