← 模型目录

Qwen3-Max-Instruct / Qwen3-Max-Thinking

Alibaba / Qwen · 2025-09-24 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Qwen3-Max-Instruct

Qwen3-Max 发布以「大就是好(Just Scale it)」为题,Instruct 为超 1T 总参 MoE 旗舰(36T 预训练 tokens、全局 batch 负载均衡损失,ChunkFlow 支持 1M 长上下文训练)。已收录 9 项评测:亮点 SWE-bench Verified 69.6、LMArena 文本榜 1430。

输入模态
文本
上下文
1M
参数
>1T(MoE,总参)
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

swebench 69.6 模型 qwen3-max-instruct · 版本 Verified · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Instruct · row: SWE-Bench Verified · figure: Chart image https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3-Max/Qwen3-Max-Instruct.png (main benchmark table); prose value cross-checks with vision read 69.6 · quote_snippet: 在专注于解决现实编程挑战的基准测试 SWE-Bench Verified 上,Qwen3-Max-Instruct 取得了高达69.6分的优异成绩

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose score; harness not disclosed. Chart image columns for the same row: Qwen3-235B-A22B-Instruct-2507 52.2, Claude Opus 4 Non-thinking 72.5, DeepSeek-V3.1 Non-thinking 66.0.

打开官方来源

tau-bench 74.8 模型 qwen3-max-instruct · 版本 2, weighted (Tau2-Bench) · 指标 weighted_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Instruct · row: Tau2-Bench · figure: Chart image Qwen3-Max-Instruct.png; prose value cross-checks with vision read 74.8 · quote_snippet: 在评估智能体工具调用能力的严苛基准 Tau2-Bench 上,Qwen3-Max-Instruct 更是实现了突破性表现,以74.8分超越 Claude Opus 4与 DeepSeek-V3.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "weighted",
  "judge": null
}

Prose names the surpassed competitors (Claude Opus 4, DeepSeek-V3.1) without printing their numbers; chart image adds Qwen3-235B-A22B-Instruct-2507 52.9, Claude Opus 4 Non-thinking 67.7, DeepSeek-V3.1 Non-thinking 46.4. Weighted aggregation over domains.

打开官方来源

supergpqa 65.1 模型 qwen3-max-instruct · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Instruct (main benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Instruct.png · row: SuperGPQA · figure: images/04.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct.png 柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: supergpqa already introduced by prior batches, still not in data/benchmarks/. Vision-assisted read (unconfirmed): SuperGPQA Qwen3-Max-Instruct 65.1; Qwen3-235B-A22B-Instruct-2507 62.6, Claude Opus 4 Non-thinking 56.5, DeepSeek-V3.1 Non-thinking 59.8. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Instruct-2507 62.6, Claude Opus 4 Non-thinking 56.5, DeepSeek-V3.1 Non-thinking 59.8。

打开官方来源

aime-25 81.6 模型 qwen3-max-instruct · 版本 Instruct, no tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Instruct (main benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Instruct.png · row: AIME25 · figure: images/04.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct.png 柱状图)

{
  "harness": null,
  "tools": [],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): AIME25 Qwen3-Max-Instruct 81.6; Qwen3-235B-A22B-Instruct-2507 70.3, Claude Opus 4 Non-thinking 33.9, DeepSeek-V3.1 Non-thinking 49.8. Different model and protocol from the Thinking-Heavy 100.0 row; never merge. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Instruct-2507 70.3, Claude Opus 4 Non-thinking 33.9, DeepSeek-V3.1 Non-thinking 49.8。

打开官方来源

lcb 69 模型 qwen3-max-instruct · 版本 v6 (25.02-25.05) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Instruct (main benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Instruct.png · row: LiveCodeBench v6 (25.02-25.05) · figure: images/04.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct.png 柱状图)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): LiveCodeBench v6 Qwen3-Max-Instruct 69.0; Qwen3-235B-A22B-Instruct-2507 51.8, Claude Opus 4 Non-thinking 44.6, DeepSeek-V3.1 Non-thinking 52.3. Version window 2025-02 to 2025-05 printed in chart. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Instruct-2507 51.8, Claude Opus 4 Non-thinking 44.6, DeepSeek-V3.1 Non-thinking 52.3。

打开官方来源

arena 1430 模型 qwen3-max-instruct · 版本 LMArena text leaderboard · 指标 elo_rating · 单位 elo 来源等级 A · third_party_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 引言 / Qwen3-Max-Instruct · row: LMArena text leaderboard · figure: images/03.png (archive of qianwen-res Qwen3-Max/Qwen3-Max-Instruct-text_arena.png 截图) · quote_snippet: 目前,Qwen3-Max-Instruct 的预览版在 LMArena 文本排行榜上位列第三,超越了 GPT-5-Chat

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard (crowd-sourced human preference battles)",
  "judge": "LMArena crowd voters"
}

Prose claims rank 3 on LMArena text leaderboard for the preview version, surpassing GPT-5-Chat; the rank is a leaderboard snapshot, not a vendor-run score, so attribution is third_party_reported and stays out of vendor self-report counts. Elo values live in the leaderboard screenshot image; vision read not performed for the numeric Elo (pending). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:截图 Last Updated Sep 18, 2025:qwen3-max-preview 排名 3(UB),Elo 1430 ±7,11,851 票;gpt-5-chat 同分 1430 排 5;gemini-2.5-pro 1456 排 1。

打开官方来源

Qwen3-Max-Thinking

Qwen3-Max-Thinking 为推理增强版,发布时仍在训练中(Heavy 配置:代码解释器 + 并行测试时计算)。已收录评测为预览披露:亮点 AIME 2025(Heavy,带 Python)100%、HMMT 25(Heavy,带 Python)100%。

输入模态
文本
上下文
1M
参数
>1T(MoE,总参)
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

aime-25 100.0 模型 qwen3-max-thinking · 版本 Thinking (Heavy), w/ python · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Thinking (Heavy) + 引言 · row: AIME 25 · figure: Chart image https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3-Max/Qwen3-Max-Thinking-blog.jpg; vision read 100.0 cross-checks the prose claim of full marks · quote_snippet: 在结合工具使用并增加测试时计算资源的情况下,该"思考"版本已在 AIME 25、HMMT 等高难度推理基准测试中取得 100% 的准确率

{
  "harness": null,
  "tools": [
    "python (code interpreter)"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Score applies to the still-in-training Thinking (Heavy) preview disclosed in the same post; chart columns: Qwen3-235B-A22B-Thinking-2507 92.3, Grok4 Heavy w/ Python 100.0, GPT-5 Pro w/ Python 100.0. Do not attribute this score to qwen3-max-instruct (81.6 no-tools is a separate pending row).

打开官方来源

hmmt25 100.0 模型 qwen3-max-thinking · 版本 Thinking (Heavy), w/ python · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Thinking (Heavy) · row: HMMT · figure: Chart image Qwen3-Max-Thinking-blog.jpg; vision read 100.0 cross-checks the prose full-marks claim · quote_snippet: 在极具挑战性的数学推理基准测试 AIME 25 和 HMMT 上,均取得了满分

{
  "harness": null,
  "tools": [
    "python (code interpreter)"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

HMMT edition not specified in prose (HMMT 2025 assumed from chart context); hmmt-25 id already introduced by prior batches, still not in data/benchmarks/. Chart columns: Qwen3-235B-A22B-Thinking-2507 83.9, Grok4 Heavy 96.7, GPT-5 Pro 100.0.

打开官方来源

gpqa 85.4 模型 qwen3-max-thinking · 版本 Thinking (Heavy), w/ python · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Qwen3-Max-Thinking (Heavy) (benchmark chart) · table: benchmark comparison chart image Qwen3-Max-Thinking-blog.jpg · row: GPQA · figure: images/05.jpg (archive of qianwen-res Qwen3-Max/Qwen3-Max-Thinking-blog.jpg 柱状图)

{
  "harness": null,
  "tools": [
    "python (code interpreter)"
  ],
  "shots": null,
  "reasoning_effort": "Heavy (parallel test-time compute)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-assisted read (unconfirmed): GPQA Qwen3-Max-Thinking Heavy w/ Python 85.4; Qwen3-235B-A22B-Thinking-2507 81.1, Grok4 Heavy 88.4, GPT-5 Pro 89.4. GPQA variant (Diamond vs main) not resolvable from chart. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同图/同表:Qwen3-235B-A22B-Thinking-2507 81.1, Grok4 Heavy w/ Python 88.4, GPT-5 Pro w/ Python 89.4;同图 AIME25 100.0、HMMT25 100.0。

打开官方来源