← 模型目录

Qwen2.5-72B-Instruct / Qwen2.5-Coder-7B-Instruct / Qwen2.5-Math-72B-Instruct / Qwen2.5-Math-7B-Instruct / Qwen2.5-Math-1.5B-Instruct

Alibaba / Qwen · 2024-09-19 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Qwen2.5-72B-Instruct

Qwen2.5 发布文以「A Party of Foundation Models」呈现全家族,72B-Instruct 为其中最大开源指令模型(稠密解码器,18T 预训练 tokens)。已收录 14 项评测覆盖知识、数学、代码与对齐:亮点 GSM8K 95.8、MMLU-Pro 71.1。

输入模态
文本
上下文
官方资料未说明
参数
72B 稠密
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu-pro 71.1 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: MMLU-Pro · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表) · quote_snippet: We benchmark our largest open-source model, Qwen2.5-72B ... against leading open-source models like Llama-3.1-70B and Mistral-Large-V2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed, per goal.md 12.5): MMLU-Pro Qwen2.5-72B 71.1; Qwen2-72B 64.4, Mistral-Large2 69.4, Llama3.1-70B 66.4, Llama3.1-405B 73.3. No protocol footnotes on this page; shots undisclosed. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 64.4, Mistral-Large2 69.4, Llama3.1-70B 66.4, Llama3.1-405B 73.3。

打开官方来源

mmlu-redux 86.8 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: MMLU-redux · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): MMLU-redux Qwen2.5-72B 86.8; Qwen2-72B 81.6, Mistral-Large2 83.0, Llama3.1-70B 83.0, Llama3.1-405B 86.2. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 81.6, Mistral-Large2 83, Llama3.1-70B 83, Llama3.1-405B 86.2。

打开官方来源

gpqa 49 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: GPQA · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): GPQA Qwen2.5-72B 49.0; Qwen2-72B 42.4, Mistral-Large2 52.0, Llama3.1-70B 46.7, Llama3.1-405B 51.1. GPQA variant (main vs Diamond) not printed. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 42.4, Mistral-Large2 52, Llama3.1-70B 46.7, Llama3.1-405B 51.1。

打开官方来源

math 83.1 模型 qwen2-5-72b-instruct · 版本 Hendrycks MATH · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: MATH · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表) · quote_snippet: greatly improved capabilities in coding (HumanEval 85+) and mathematics (MATH 80+)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: math (Hendrycks MATH full test set, distinct from math500 subset id) not yet in data/benchmarks/. Vision-assisted read (unconfirmed): MATH Qwen2.5-72B 83.1; Qwen2-72B 69.0, Mistral-Large2 69.9, Llama3.1-70B 68.0, Llama3.1-405B 73.8. Prose family-level claim 'MATH 80+' cross-checks at value level. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 69, Mistral-Large2 69.9, Llama3.1-70B 68, Llama3.1-405B 73.8。

打开官方来源

gsm8k 95.8 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: GSM8K · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): GSM8K Qwen2.5-72B 95.8; Qwen2-72B 93.2, Mistral-Large2 92.7, Llama3.1-70B 95.1, Llama3.1-405B 96.8. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 93.2, Mistral-Large2 92.7, Llama3.1-70B 95.1, Llama3.1-405B 96.8。

打开官方来源

humaneval 86.6 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: HumanEval · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表) · quote_snippet: greatly improved capabilities in coding (HumanEval 85+)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): HumanEval Qwen2.5-72B 86.6; Qwen2-72B 86.0, Mistral-Large2 92.1, Llama3.1-70B 80.5, Llama3.1-405B 89.0. Prose family-level claim 'HumanEval 85+' cross-checks at value level. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 86, Mistral-Large2 92.1, Llama3.1-70B 80.5, Llama3.1-405B 89。

打开官方来源

mbpp 88.2 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: MBPP · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mbpp not yet in data/benchmarks/. Vision-assisted read (unconfirmed): MBPP Qwen2.5-72B 88.2; Qwen2-72B 80.2, Mistral-Large2 80.0, Llama3.1-70B 84.2, Llama3.1-405B 84.5. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 80.2, Mistral-Large2 80, Llama3.1-70B 84.2, Llama3.1-405B 84.5。

打开官方来源

multipl-e 75.1 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: MultiPL-E · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multipl-e not yet in data/benchmarks/ (also introduced in kimi-k2.json this batch). Vision-assisted read (unconfirmed): MultiPL-E Qwen2.5-72B 75.1; Qwen2-72B 69.2, Mistral-Large2 76.9, Llama3.1-70B 68.2, Llama3.1-405B 73.5. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 69.2, Mistral-Large2 76.9, Llama3.1-70B 68.2, Llama3.1-405B 73.5。

打开官方来源

lcb 55.5 模型 qwen2-5-72b-instruct · 版本 2305-2409 window · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: LiveCodeBench 2305-2409 · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): LiveCodeBench (2023-05 to 2024-09 window) Qwen2.5-72B 55.5; Qwen2-72B 32.2, Mistral-Large2 42.2, Llama3.1-70B 32.1, Llama3.1-405B 41.6. Time window differs from later releases' v6 rows - not comparable. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 32.2, Mistral-Large2 42.2, Llama3.1-70B 32.1, Llama3.1-405B 41.6。

打开官方来源

livebench 52.3 模型 qwen2-5-72b-instruct · 版本 0831 (2024-08-31 snapshot) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: LiveBench 0831 · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): LiveBench 0831 Qwen2.5-72B 52.3; Qwen2-72B 41.5, Mistral-Large2 48.5, Llama3.1-70B 46.6, Llama3.1-405B 53.2. Date-pinned snapshot is part of the protocol. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 41.5, Mistral-Large2 48.5, Llama3.1-70B 46.6, Llama3.1-405B 53.2。

打开官方来源

ifeval 84.1 模型 qwen2-5-72b-instruct · 版本 strict-prompt · 指标 prompt_strict_acc · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: IFEval strict-prompt · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表) · quote_snippet: significant improvements in instruction following

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "strict-prompt",
  "judge": null
}

Vision-assisted read (unconfirmed): IFEval strict-prompt Qwen2.5-72B 84.1; Qwen2-72B 77.6, Mistral-Large2 64.1, Llama3.1-70B 83.6, Llama3.1-405B 86.0. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 77.6, Mistral-Large2 64.1, Llama3.1-70B 83.6, Llama3.1-405B 86。

打开官方来源

arenahard 81.2 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 win_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: Arena-Hard · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "win rate (LLM judge)",
  "judge": null
}

Vision-assisted read (unconfirmed): Arena-Hard Qwen2.5-72B 81.2; Qwen2-72B 48.1, Mistral-Large2 73.1, Llama3.1-70B 55.7, Llama3.1-405B 69.3. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 48.1, Mistral-Large2 73.1, Llama3.1-70B 55.7, Llama3.1-405B 69.3。

打开官方来源

alignbench 8.16 模型 qwen2-5-72b-instruct · 版本 v1.1 · 指标 alignbench_score · 单位 points 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: AlignBench v1.1 · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: alignbench (Chinese alignment benchmark) not yet in data/benchmarks/. Unit is a 1-10 score, not percent - vision-assisted read (unconfirmed): AlignBench v1.1 Qwen2.5-72B 8.16; Qwen2-72B 8.15, Mistral-Large2 7.69, Llama3.1-70B 5.94, Llama3.1-405B 5.95. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 8.15, Mistral-Large2 7.69, Llama3.1-70B 5.94, Llama3.1-405B 5.95。

打开官方来源

mtbench 9.35 模型 qwen2-5-72b-instruct · 版本 未说明 · 指标 mtbench_score · 单位 points 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Performance / Qwen2.5 · table: benchmark comparison table image Qwen2.5-72B-Instruct-Score.jpg · row: MT-bench · figure: images/03.jpg (archive of qianwen-res Qwen2.5-72B-Instruct-Score.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Unit is a 1-10 score, not percent - vision-assisted read (unconfirmed): MT-bench Qwen2.5-72B 9.35; Qwen2-72B 9.12, Mistral-Large2 8.61, Llama3.1-70B 8.79, Llama3.1-405B 9.08. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表对照列:Qwen2-72B 9.12, Mistral-Large2 8.61, Llama3.1-70B 8.79, Llama3.1-405B 9.08。

打开官方来源

Qwen2.5-Coder-7B-Instruct

Qwen2.5-Coder-7B-Instruct 为家族代码专精模型(5.5T 代码 tokens 训练)。已收录 11 项代码评测:亮点 HumanEval 88.4、Spider 82。

输入模态
文本
上下文
官方资料未说明
参数
7B 稠密
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

humaneval 88.4 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: HumanEval · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图) · quote_snippet: we present the performance results of Qwen2.5-Coder-7B-Instruct, benchmarked against leading open-source models, including those with significantly larger parameter sizes

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): HumanEval Qwen2.5-Coder-7B-Instruct 88.4; DeepSeek-Coder-V2-Lite-Instruct 81.1, DeepSeek-Coder-33B-Instruct 79.3, CodeStral-22B 78.1, DeepSeek-Coder-6.7B-Instruct 78.6. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对,放大裁片确认:CodeStral 棕色=78.1,DS-Coder-6.7B 蓝色=78.6)。同图竞品:DS-Coder-V2-Lite 81.1, DS-Coder-33B 79.3, CodeStral-22B 78.1, DS-Coder-6.7B 78.6。

打开官方来源

evalplus 81.9 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: EvalPlus · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: evalplus not yet in data/benchmarks/. Vision-assisted read (unconfirmed): EvalPlus Qwen2.5-Coder-7B 81.9; DS-Coder-V2-Lite 76.8, DS-Coder-33B 74.9, CodeStral-22B 73.5, DS-Coder-6.7B 72.6. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 76.8, DS-Coder-33B 74.9, CodeStral-22B 73.5, DS-Coder-6.7B 72.6。

打开官方来源

aider 50.4 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: Aider · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Aider edit format (whole/diff) not printed; not directly comparable with later Aider-Polyglot rows. Vision-assisted read (unconfirmed): Aider Qwen2.5-Coder-7B 50.4; DS-Coder-V2-Lite 48.9, DS-Coder-33B 49.6, CodeStral-22B 35.3, DS-Coder-6.7B 34.6. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 48.9, DS-Coder-33B 49.6, CodeStral-22B 35.3, DS-Coder-6.7B 34.6。

打开官方来源

lcb 35.9 模型 qwen2-5-coder-7b-instruct · 版本 2305-2409 window · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: LiveCodeBench · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): LiveCodeBench Qwen2.5-Coder-7B 35.9; DS-Coder-V2-Lite 24.3, DS-Coder-33B 27.7, CodeStral-22B 32.9, DS-Coder-6.7B 20.5. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 24.3, DS-Coder-33B 27.7, CodeStral-22B 32.9, DS-Coder-6.7B 20.5。

打开官方来源

spider 82 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: Spider · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: spider (text-to-SQL) not yet in data/benchmarks/. Vision-assisted read (unconfirmed): Spider Qwen2.5-Coder-7B 82.0; DS-Coder-V2-Lite 74.6, DS-Coder-33B 73.8, CodeStral-22B 76.6, DS-Coder-6.7B 70.0. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 74.6, DS-Coder-33B 73.8, CodeStral-22B 76.6, DS-Coder-6.7B 70.0。 更正:先前转写把 33B 与 CodeStral 两格看反,放大核对图例颜色后 33B=73.8、CodeStral=76.6。

打开官方来源

bird-sql 51.1 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: BIRD-SQL · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: bird-sql not yet in data/benchmarks/. Vision-assisted read (unconfirmed): BIRD-SQL Qwen2.5-Coder-7B 51.1; DS-Coder-V2-Lite 41.6, DS-Coder-33B 45.6, CodeStral-22B 46.2, DS-Coder-6.7B 39.8. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对,放大裁片确认:33B 浅橙=45.6,CodeStral 棕=46.2)。同图竞品:DS-Coder-V2-Lite 41.6, DS-Coder-33B 45.6, CodeStral-22B 46.2, DS-Coder-6.7B 39.8。

打开官方来源

bigcodebench 33.1 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: BigCodeBench · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): BigCodeBench Qwen2.5-Coder-7B 33.1; DS-Coder-V2-Lite 28.1, DS-Coder-33B 32.5, CodeStral-22B 34.8, DS-Coder-6.7B 24.5. Full/instruct split not printed. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 28.1, DS-Coder-33B 32.5, CodeStral-22B 34.8, DS-Coder-6.7B 24.5。

打开官方来源

mceval 60.3 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: McEval · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mceval (multilingual coding eval) not yet in data/benchmarks/. Vision-assisted read (unconfirmed): McEval Qwen2.5-Coder-7B 60.3; DS-Coder-V2-Lite 54.7, DS-Coder-33B 54.3, CodeStral-22B 50.5, DS-Coder-6.7B 46.0. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 54.7, DS-Coder-33B 54.3, CodeStral-22B 50.5, DS-Coder-6.7B 46.0。

打开官方来源

multipl-e 76.5 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: MultiPL-E · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multipl-e (also in kimi-k2.json and the 72B rows above). Vision-assisted read (unconfirmed): MultiPL-E Qwen2.5-Coder-7B 76.5; DS-Coder-V2-Lite 73.2, DS-Coder-33B 69.2, CodeStral-22B 70.2, DS-Coder-6.7B 66.1. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 73.2, DS-Coder-33B 69.2, CodeStral-22B 70.2, DS-Coder-6.7B 66.1。

打开官方来源

cruxeval 65.9 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: CRUXEval · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cruxeval (code reasoning/output prediction) not yet in data/benchmarks/. Vision-assisted read (unconfirmed): CRUXEval Qwen2.5-Coder-7B 65.9; DS-Coder-V2-Lite 53.0, DS-Coder-33B 52.8, CodeStral-22B 62.2, DS-Coder-6.7B 43.9. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 53.0, DS-Coder-33B 52.8, CodeStral-22B 62.2, DS-Coder-6.7B 43.9。

打开官方来源

mbpp 83.5 模型 qwen2-5-coder-7b-instruct · 版本 未说明 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Coder · table: benchmark comparison table image coder-main.png · row: MBPP · figure: images/08.png (archive of qianwen-res coder-main.png 环形分组柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mbpp (also in the 72B rows above). Vision-assisted read (unconfirmed): MBPP Qwen2.5-Coder-7B 83.5; DS-Coder-V2-Lite 82.3, DS-Coder-33B 81.2, CodeStral-22B 73.3, DS-Coder-6.7B 75.1. 视觉转写自归档图 images/08.png(2026-09-01 复核,颜色经图例核对)。同图竞品:DS-Coder-V2-Lite 82.3, DS-Coder-33B 81.2, CodeStral-22B 73.3, DS-Coder-6.7B 75.1。

打开官方来源

Qwen2.5-Math-72B-Instruct

Qwen2.5-Math-72B-Instruct 为家族数学专精模型,支持 CoT/PoT/TIR 与中英双语。已收录评测为 MATH zero-shot@1 85.9。

输入模态
文本
上下文
官方资料未说明
参数
72B 稠密
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

math 85.9 模型 qwen2-5-math-72b-instruct · 版本 zero-shot@1 · 指标 zero_shot@1_accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Math · row: MATH · figure: images/09.png (archive of qianwen-res MATH Zero-shot@1 scatter) · quote_snippet: The general performance of Qwen2.5-Math-72B-Instruct surpasses both Qwen2-Math-72B-Instruct and GPT4-o

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "Zero-shot@1",
  "judge": null
}

new-benchmark: math (Hendrycks MATH) not yet in data/benchmarks/. Prose claim (surpasses Qwen2-Math-72B-Instruct and GPT4-o) verified; the absolute value 85.9 is printed as a labeled data point on the scatter (vision-confirmed 2026-09-01), so score_status is reported. Chart also labels Qwen2-Math-72B-Instruct 84.0. 视觉转写自归档图 images/09.png(MATH Zero-shot@1 散点,标签数值清晰,2026-09-01 复核)。同图对照点:Qwen2-Math-72B-Instruct 84.0。

打开官方来源

Qwen2.5-Math-7B-Instruct

Qwen2.5-Math-7B-Instruct 为家族数学专精 7B 档。已收录评测为 MATH zero-shot@1 83.6。

输入模态
文本
上下文
官方资料未说明
参数
7B 稠密
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

math 83.6 模型 qwen2-5-math-7b-instruct · 版本 zero-shot@1 · 指标 zero_shot@1_accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Math · table: scatter chart image 2024-08-qwen2.5-math-allsize.png (MATH accuracy vs parameters) · row: Qwen2.5-Math-7B-Instruct · figure: images/09.png (archive of qianwen-res MATH Zero-shot@1 scatter) · quote_snippet: even very small expert model like Qwen2.5-Math-1.5B-Instruct can achieve highly competitive performance

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "Zero-shot@1",
  "judge": null
}

new-benchmark: math. Vision-assisted read of the labeled scatter point (unconfirmed): MATH zero-shot@1 Qwen2.5-Math-7B-Instruct 83.6; Qwen2-Math-7B-Instruct 75.1. 视觉转写自归档图 images/09.png(MATH Zero-shot@1 散点,标签数值清晰,2026-09-01 复核)。同图对照点:Qwen2-Math-7B-Instruct 75.1。

打开官方来源

Qwen2.5-Math-1.5B-Instruct

Qwen2.5-Math-1.5B-Instruct 为家族数学专精最小档(1.5B)。已收录评测为 MATH zero-shot@1 75.8。

输入模态
文本
上下文
官方资料未说明
参数
1.5B 稠密
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

math 75.8 模型 qwen2-5-math-1-5b-instruct · 版本 zero-shot@1 · 指标 zero_shot@1_accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen2.5-Math · table: scatter chart image 2024-08-qwen2.5-math-allsize.png (MATH accuracy vs parameters) · row: Qwen2.5-Math-1.5B-Instruct · figure: images/09.png (archive of qianwen-res MATH Zero-shot@1 scatter) · quote_snippet: even very small expert model like Qwen2.5-Math-1.5B-Instruct can achieve highly competitive performance

{
  "harness": null,
  "tools": null,
  "shots": 0,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 1,
  "aggregation": "Zero-shot@1",
  "judge": null
}

new-benchmark: math. Vision-assisted read of the labeled scatter point (unconfirmed): MATH zero-shot@1 Qwen2.5-Math-1.5B-Instruct 75.8; Qwen2-Math-1.5B-Instruct 69.4. 视觉转写自归档图 images/09.png(MATH Zero-shot@1 散点,标签数值清晰,2026-09-01 复核)。同图对照点:Qwen2-Math-1.5B-Instruct 69.4。

打开官方来源