← 模型目录

Qwen3-235B-A22B / Qwen3-30B-A3B / Qwen3-32B / Qwen3-4B / Qwen3-14B / Qwen3-8B / Qwen3-1.7B / Qwen3-0.6B

Alibaba / Qwen · 2025-04-29 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Qwen3-235B-A22B

Qwen3 以「Think Deeper, Act Faster」发布,Qwen3-235B-A22B 为混合思考 MoE 旗舰。评测覆盖推理、编码与多语言,亮点如 ArenaHard 95.6%、AIME'24 85.7%。

输入模态
文本
上下文
官方资料未说明
参数
235B-A22B MoE
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

arenahard 95.6 模型 qwen3-235b-a22b · 版本 未说明 · 指标 win_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: ArenaHard · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-235B-A22B=95.6; Qwen3-32B=93.8. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 93.8;竞品列:OpenAI-o1 92.1, Deepseek-R1 93.2, Grok3 Beta '-', Gemini2.5-Pro 96.4, OpenAI-o3-mini 89.0。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

aime24 85.7 模型 qwen3-235b-a22b · 版本 2024 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: AIME'24 · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 64,
  "aggregation": "mean accuracy over 64 samples",
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-235B-A22B=85.7; Qwen3-32B=81.4. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 81.4;竞品列:OpenAI-o1 74.3, Deepseek-R1 79.8, Grok3 Beta 83.9, Gemini2.5-Pro 92.0, OpenAI-o3-mini 79.6。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

aime-25 81.5 模型 qwen3-235b-a22b · 版本 2025 (Part I + II, 30 questions) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: AIME'25 · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 64,
  "aggregation": "mean accuracy over 64 samples",
  "judge": null
}

new-benchmark: aime-25 not yet in data/benchmarks.json. Vision-read values (unconfirmed): Qwen3-235B-A22B=81.5; Qwen3-32B=72.9. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 72.9;竞品列:OpenAI-o1 79.2, Deepseek-R1 70.0, Grok3 Beta 77.3, Gemini2.5-Pro 86.7, OpenAI-o3-mini 74.8。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

lcb 70.7 模型 qwen3-235b-a22b · 版本 v5 (2024.10-2025.02) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: LiveCodeBench · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-235B-A22B=70.7; Qwen3-32B=65.7. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 65.7;竞品列:OpenAI-o1 63.9, Deepseek-R1 64.3, Grok3 Beta 70.6, Gemini2.5-Pro 70.4, OpenAI-o3-mini 66.3。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

codeforces 2056 模型 qwen3-235b-a22b · 版本 未说明 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: CodeForces (Elo Rating) · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: codeforces not yet in data/benchmarks.json. Vision-read values (unconfirmed): Qwen3-235B-A22B=2056; Qwen3-32B=1977. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 1977;竞品列:OpenAI-o1 1891, Deepseek-R1 2029, Grok3 Beta '-', Gemini2.5-Pro 2001, OpenAI-o3-mini 2036。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

aider 61.8 模型 qwen3-235b-a22b · 版本 pass@2, non-thinking · 指标 pass@2 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: Aider (Pass@2) · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "non-thinking (think mode not activated)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-235B-A22B=61.8; Qwen3-32B=50.2. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 50.2;竞品列:OpenAI-o1 61.7, Deepseek-R1 56.9, Grok3 Beta 53.3, Gemini2.5-Pro 72.9, OpenAI-o3-mini 53.8。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

livebench 77.1 模型 qwen3-235b-a22b · 版本 2024-11-25 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: LiveBench · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-235B-A22B=77.1; Qwen3-32B=74.9. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 74.9;竞品列:OpenAI-o1 75.7, Deepseek-R1 71.6, Grok3 Beta '-', Gemini2.5-Pro 82.4, OpenAI-o3-mini 70.0。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

bfcl 70.8 模型 qwen3-235b-a22b · 版本 v3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: BFCL v3 · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-235B-A22B=70.8; Qwen3-32B=70.3. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. Qwen3 rows use FC format per in-chart footnote; baseline cells use best-of FC/prompt, so competitor cells are not same-protocol. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 70.3;竞品列:OpenAI-o1 67.8, Deepseek-R1 56.9, Grok3 Beta '-', Gemini2.5-Pro 62.9, OpenAI-o3-mini 64.6。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

multiif 71.9 模型 qwen3-235b-a22b · 版本 8 languages · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: MultiIF (8 Languages) · figure: images/03.jpg (archive of qianwen-res Qwen3/qwen3-235a22.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multiif not yet in data/benchmarks.json. Vision-read values (unconfirmed): Qwen3-235B-A22B=71.9; Qwen3-32B=73.0. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/03.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-32B 列 73.0;竞品列:OpenAI-o1 48.8, Deepseek-R1 67.7, Grok3 Beta '-', Gemini2.5-Pro 77.8, OpenAI-o3-mini 48.4。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

Qwen3-30B-A3B

Qwen3-30B-A3B 为该代混合思考小 MoE 档。亮点如 ArenaHard 91.0%、AIME'24 80.4%。

输入模态
文本
上下文
官方资料未说明
参数
30B-A3B MoE
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

arenahard 91 模型 qwen3-30b-a3b · 版本 未说明 · 指标 win_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: ArenaHard · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-30B-A3B=91.0; Qwen3-4B=76.6. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 76.6;竞品列:QwQ-32B 89.5, Qwen2.5-72B-Instruct 81.2, Gemma3-27B-IT 86.8, DeepSeek-V3 85.5, GPT-4o 85.3。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

aime24 80.4 模型 qwen3-30b-a3b · 版本 2024 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: AIME'24 · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 64,
  "aggregation": "mean accuracy over 64 samples",
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-30B-A3B=80.4; Qwen3-4B=73.8. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 73.8;竞品列:QwQ-32B 79.5, Qwen2.5-72B-Instruct 18.9, Gemma3-27B-IT 32.6, DeepSeek-V3 39.2, GPT-4o 11.1。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

aime-25 70.9 模型 qwen3-30b-a3b · 版本 2025 (Part I + II, 30 questions) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: AIME'25 · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 64,
  "aggregation": "mean accuracy over 64 samples",
  "judge": null
}

new-benchmark: aime-25 not yet in data/benchmarks.json. Vision-read values (unconfirmed): Qwen3-30B-A3B=70.9; Qwen3-4B=65.6. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 65.6;竞品列:QwQ-32B 69.5, Qwen2.5-72B-Instruct 15.0, Gemma3-27B-IT 24.0, DeepSeek-V3 28.8, GPT-4o 7.6。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

lcb 62.6 模型 qwen3-30b-a3b · 版本 v5 (2024.10-2025.02) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: LiveCodeBench · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-30B-A3B=62.6; Qwen3-4B=54.2. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 54.2;竞品列:QwQ-32B 62.7, Qwen2.5-72B-Instruct 30.7, Gemma3-27B-IT 26.9, DeepSeek-V3 33.1, GPT-4o 32.7。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

codeforces 1974 模型 qwen3-30b-a3b · 版本 未说明 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: CodeForces (Elo Rating) · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: codeforces not yet in data/benchmarks.json. Vision-read values (unconfirmed): Qwen3-30B-A3B=1974; Qwen3-4B=1671. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 1671;竞品列:QwQ-32B 1982, Qwen2.5-72B-Instruct 859, Gemma3-27B-IT 1063, DeepSeek-V3 1134, GPT-4o 864。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

gpqa 65.8 模型 qwen3-30b-a3b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: GPQA · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-30B-A3B=65.8; Qwen3-4B=55.9. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. Chart label reads only 'GPQA' without subset qualifier; Diamond assumption must be confirmed before merging with existing gpqa (GPQA Diamond) data. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 55.9;竞品列:QwQ-32B 65.6, Qwen2.5-72B-Instruct 49.0, Gemma3-27B-IT 42.4, DeepSeek-V3 59.1, GPT-4o 46.0。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

livebench 74.3 模型 qwen3-30b-a3b · 版本 2024-11-25 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: LiveBench · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-30B-A3B=74.3; Qwen3-4B=63.6. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 63.6;竞品列:QwQ-32B 72.0, Qwen2.5-72B-Instruct 51.4, Gemma3-27B-IT 49.2, DeepSeek-V3 60.5, GPT-4o 52.2。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

bfcl 69.1 模型 qwen3-30b-a3b · 版本 v3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: BFCL (v3) · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read values (unconfirmed): Qwen3-30B-A3B=69.1; Qwen3-4B=65.9. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. Qwen3 rows use FC format per in-chart footnote. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 65.9;竞品列:QwQ-32B 66.4, Qwen2.5-72B-Instruct 63.4, Gemma3-27B-IT 59.1, DeepSeek-V3 57.6, GPT-4o 72.5。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

multiif 72.2 模型 qwen3-30b-a3b · 版本 8 languages · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Introduction (two benchmark comparison charts follow the opening paragraph) · row: MultiIF (8 Languages) · figure: images/04.jpg (archive of qianwen-res Qwen3/qwen3-30a3.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multiif not yet in data/benchmarks.json. Vision-read values (unconfirmed): Qwen3-30B-A3B=72.2; Qwen3-4B=66.3. In-chart footnotes (vision-read): AIME 24/25 sampled 64 times per query, average accuracy reported; AIME'25 = Part I + Part II, 30 questions total; Aider run does not activate think mode; BFCL: Qwen3 uses FC format, baselines use best of FC/prompt formats. 视觉转写自归档图 images/04.jpg(2026-09-01 复核,与先前读数一致)。同表 Qwen3-4B 列 66.3;竞品列:QwQ-32B 68.3, Qwen2.5-72B-Instruct 65.3, Gemma3-27B-IT 69.8, DeepSeek-V3 55.6, GPT-4o 65.6。图表脚注(图内小字):1) AIME 24/25 每题采样 64 次取平均;AIME'25 为 Part I+II 共 30 题;2) Aider 行 Qwen3 未激活 think 模式;3) BFCL 行 Qwen3 用 FC 格式,基线取 FC/prompt 两种格式最高分。

打开官方来源

Qwen3-32B

Qwen3-32B 为该代混合思考 dense 档。发布图表另给出该档对比,如 ArenaHard 93.8、AIME'24 81.4(账本未单列证据行)。

输入模态
文本
上下文
官方资料未说明
参数
32B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

Qwen3-4B

Qwen3-4B 为该代混合思考 dense 小档。发布图表另给出该档对比,如 ArenaHard 76.6、AIME'24 73.8(账本未单列证据行)。

输入模态
文本
上下文
官方资料未说明
参数
4B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

Qwen3-14B

Qwen3-14B 为该代 dense 系列中坚。本次发布图表未含该档独立评测数值。

输入模态
文本
上下文
官方资料未说明
参数
14B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

Qwen3-8B

Qwen3-8B 为该代 dense 系列通用档。本次发布图表未含该档独立评测数值。

输入模态
文本
上下文
官方资料未说明
参数
8B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

Qwen3-1.7B

Qwen3-1.7B 为该代 dense 系列轻量档。本次发布图表未含该档独立评测数值。

输入模态
文本
上下文
官方资料未说明
参数
1.7B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

Qwen3-0.6B

Qwen3-0.6B 为该代 dense 系列最小档。本次发布图表未含该档独立评测数值。

输入模态
文本
上下文
官方资料未说明
参数
0.6B dense
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。