Qwen3-Max-Instruct / Qwen3-Max-Thinking
Alibaba / Qwen · 2025-09-24 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Qwen3-Max-Instruct
Qwen3-Max 发布以「大就是好(Just Scale it)」为题,Instruct 为超 1T 总参 MoE 旗舰(36T 预训练 tokens、全局 batch 负载均衡损失,ChunkFlow 支持 1M 长上下文训练)。已收录 9 项评测:亮点 SWE-bench Verified 69.6、LMArena 文本榜 1430。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- >1T(MoE,总参)
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Instruct · row: SWE-Bench Verified · figure: Chart image https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3-Max/Qwen3-Max-Instruct.png (main benchmark table); prose value cross-checks with vision read 69.6 · quote_snippet: 在专注于解决现实编程挑战的基准测试 SWE-Bench Verified 上,Qwen3-Max-Instruct 取得了高达69.6分的优异成绩
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose score; harness not disclosed. Chart image columns for the same row: Qwen3-235B-A22B-Instruct-2507 52.2, Claude Opus 4 Non-thinking 72.5, DeepSeek-V3.1 Non-thinking 66.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Instruct · row: Tau2-Bench · figure: Chart image Qwen3-Max-Instruct.png; prose value cross-checks with vision read 74.8 · quote_snippet: 在评估智能体工具调用能力的严苛基准 Tau2-Bench 上,Qwen3-Max-Instruct 更是实现了突破性表现,以74.8分超越 Claude Opus 4与 DeepSeek-V3.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "weighted",
"judge": null
}Prose names the surpassed competitors (Claude Opus 4, DeepSeek-V3.1) without printing their numbers; chart image adds Qwen3-235B-A22B-Instruct-2507 52.9, Claude Opus 4 Non-thinking 67.7, DeepSeek-V3.1 Non-thinking 46.4. Weighted aggregation over domains.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Instruct (main benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Instruct.png · row: SuperGPQA · figure: images/04.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct.png 柱状图)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: supergpqa already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): SuperGPQA Qwen3-Max-Instruct 65.1; Qwen3-235B-A22B-Instruct-2507 62.6, Claude Opus 4 Non-thinking 56.5, DeepSeek-V3.1 Non-thinking 59.8. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Instruct-2507 62.6, Claude Opus 4 Non-thinking 56.5, DeepSeek-V3.1 Non-thinking 59.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Instruct (main benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Instruct.png · row: AIME25 · figure: images/04.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct.png 柱状图)
{
"harness": null,
"tools": [],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): AIME25 Qwen3-Max-Instruct 81.6; Qwen3-235B-A22B-Instruct-2507 70.3, Claude Opus 4 Non-thinking 33.9, DeepSeek-V3.1 Non-thinking 49.8. Different model and protocol from the Thinking-Heavy 100.0 row; never merge. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Instruct-2507 70.3, Claude Opus 4 Non-thinking 33.9, DeepSeek-V3.1 Non-thinking 49.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Instruct (main benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Instruct.png · row: LiveCodeBench v6 (25.02-25.05) · figure: images/04.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct.png 柱状图)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): LiveCodeBench v6 Qwen3-Max-Instruct 69.0; Qwen3-235B-A22B-Instruct-2507 51.8, Claude Opus 4 Non-thinking 44.6, DeepSeek-V3.1 Non-thinking 52.3. Version window 2025-02 to 2025-05 printed in chart. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Instruct-2507 51.8, Claude Opus 4 Non-thinking 44.6, DeepSeek-V3.1 Non-thinking 52.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 引言 / Qwen3-Max-Instruct · row: LMArena text leaderboard · figure: images/03.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct-text_arena.png 截图) · quote_snippet: 目前,Qwen3-Max-Instruct 的预览版在 LMArena 文本排行榜上位列第三,超越了 GPT-5-Chat
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard (crowd-sourced human preference battles)",
"judge": "LMArena crowd voters"
}Prose claims rank 3 on LMArena text leaderboard for the preview version, surpassing GPT-5-Chat; the rank is a leaderboard snapshot, not a vendor-run score, so attribution is third_party_reported and stays out of vendor self-report counts. Elo values live in the leaderboard screenshot image; vision read not performed for the numeric Elo (pending). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:截图 Last Updated Sep 18, 2025:qwen3-max-preview 排名 3(UB),Elo 1430 ±7,11,851 票;gpt-5-chat 同分 1430 排 5;gemini-2.5-pro 1456 排 1。
Qwen3-Max-Thinking
Qwen3-Max-Thinking 为推理增强版,发布时仍在训练中(Heavy 配置:代码解释器 + 并行测试时计算)。已收录评测为预览披露:亮点 AIME 2025(Heavy,带 Python)100%、HMMT 25(Heavy,带 Python)100%。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- >1T(MoE,总参)
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Thinking (Heavy) + 引言 · row: AIME 25 · figure: Chart image https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3-Max/Qwen3-Max-Thinking-blog.jpg; vision read 100.0 cross-checks the prose claim of full marks · quote_snippet: 在结合工具使用并增加测试时计算资源的情况下,该"思考"版本已在 AIME 25、HMMT 等高难度推理基准测试中取得 100% 的准确率
{
"harness": null,
"tools": [
"python (code interpreter)"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Score applies to the still-in-training Thinking (Heavy) preview disclosed in the same post; chart columns: Qwen3-235B-A22B-Thinking-2507 92.3, Grok4 Heavy w/ Python 100.0, GPT-5 Pro w/ Python 100.0. Do not attribute this score to qwen3-max-instruct (81.6 no-tools is a separate pending row).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Thinking (Heavy) · row: HMMT · figure: Chart image Qwen3-Max-Thinking-blog.jpg; vision read 100.0 cross-checks the prose full-marks claim · quote_snippet: 在极具挑战性的数学推理基准测试 AIME 25 和 HMMT 上,均取得了满分
{
"harness": null,
"tools": [
"python (code interpreter)"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}HMMT edition not specified in prose (HMMT 2025 assumed from chart context); hmmt-25 id already introduced by prior batches, still not in data/benchmarks/. Chart columns: Qwen3-235B-A22B-Thinking-2507 83.9, Grok4 Heavy 96.7, GPT-5 Pro 100.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Qwen3-Max-Thinking (Heavy) (benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Thinking-blog.jpg · row: GPQA · figure: images/05.jpg (archive of qianwen-res Qwen3-Max/Qwen3-Max-Thinking-blog.jpg 柱状图)
{
"harness": null,
"tools": [
"python (code interpreter)"
],
"shots": null,
"reasoning_effort": "Heavy (parallel test-time compute)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-assisted read (unconfirmed): GPQA Qwen3-Max-Thinking Heavy w/ Python 85.4; Qwen3-235B-A22B-Thinking-2507 81.1, Grok4 Heavy 88.4, GPT-5 Pro 89.4. GPQA variant (Diamond vs main) not resolvable from chart. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Thinking-2507 81.1, Grok4 Heavy w/ Python 88.4, GPT-5 Pro w/ Python 89.4;同图 AIME25 100.0、HMMT25 100.0。