← 模型目录

Qwen3-Coder-480B-A35B-Instruct

Alibaba / Qwen · 2025-07-22 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Qwen3-Coder-480B-A35B-Instruct

Qwen3-Coder-480B-A35B-Instruct 以「Agentic Coding in the World」为定位发布,480B MoE/35B 激活,原生 256K 上下文(YaRN 可扩至 1M)。已收录 15 项评测集中于代码仓库、终端与浏览器/工具 Agent:亮点 SWE-bench Verified(OpenHands 500 turns)69.6、Terminal-Bench 37.5。

输入模态
文本
上下文
256K(YaRN 可扩至 1M)
参数
480B-A35B MoE
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

terminalbench 37.5 模型 qwen3-coder-480b-a35b-instruct · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Terminal-Bench · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=37.5. Chart label has no version qualifier ('Terminal-Bench'); do not merge with Terminal-Bench 2.x rows of other vendors without variant confirmation. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 30.0, DeepSeek-V3-0324 2.5, Claude Sonnet-4 35.5, OpenAI GPT-4.1 25.3。

打开官方来源

swebench 69.6 模型 qwen3-coder-480b-a35b-instruct · 版本 OpenHands harness, 500 turns · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Verified w/ OpenHands, 500 turns · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=69.6. Competitor cells (vision): Claude-Sonnet-4 70.4. Prose: 'Qwen3-Coder achieves state-of-the-art performance among open-source models on SWE-Bench Verified without test-time scaling' (section Post-Training / Scaling Long-Horizon RL); scatter image swe.jpg repeats the 69.6/67.0 points. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2 '-', DeepSeek-V3 '-', Claude Sonnet-4 70.4, GPT-4.1 '-'。

打开官方来源

swebench 67 模型 qwen3-coder-480b-a35b-instruct · 版本 OpenHands harness, 100 turns · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Verified w/ OpenHands, 100 turns · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=67.0. Competitor cells (vision): Kimi-K2-Instruct 65.4, DeepSeek-V3-0324 38.8, Claude-Sonnet-4 68.0, GPT-4.1 48.6. Prose: 'Qwen3-Coder achieves state-of-the-art performance among open-source models on SWE-Bench Verified without test-time scaling' (section Post-Training / Scaling Long-Horizon RL); scatter image swe.jpg repeats the 69.6/67.0 points. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 65.4, DeepSeek-V3-0324 38.8, Claude Sonnet-4 68.0, OpenAI GPT-4.1 48.6。

打开官方来源

swebench 官方未公布数值 模型 qwen3-coder-480b-a35b-instruct · 版本 private scaffolding · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Verified w/ Private Scaffolding · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=None. Chart shows no value for Qwen3-Coder in this row (dash); competitor cells (vision): Kimi-K2-Instruct 65.8, Claude-Sonnet-4 72.7, GPT-4.1 63.8. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Qwen3-Coder 该格为 '-';Kimi-K2-Instruct 65.8, DeepSeek-V3 '-', Claude Sonnet-4 72.7, OpenAI GPT-4.1 63.8。

打开官方来源

swebench-live 26.3 模型 qwen3-coder-480b-a35b-instruct · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Live · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-live not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=26.3. Competitor cells (vision): Kimi-K2-Instruct 22.3, DeepSeek-V3-0324 13.0, Claude-Sonnet-4 27.7. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 22.3, DeepSeek-V3-0324 13.0, Claude Sonnet-4 27.7, GPT-4.1 '-'。

打开官方来源

swebench-multilingual 54.7 模型 qwen3-coder-480b-a35b-instruct · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Multilingual · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-multilingual not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=54.7. Competitor cells (vision): Kimi-K2-Instruct 47.3, DeepSeek-V3-0324 13.0, Claude-Sonnet-4 53.3, GPT-4.1 31.5. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 47.3, DeepSeek-V3-0324 13.0, Claude Sonnet-4 53.3, OpenAI GPT-4.1 31.5。

打开官方来源

multi-swe-bench 25.8 模型 qwen3-coder-480b-a35b-instruct · 版本 mini · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Multi-SWE-bench mini · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multi-swebench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=25.8. Competitor cells (vision): Kimi-K2-Instruct 19.8, DeepSeek-V3-0324 7.5, Claude-Sonnet-4 24.8. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 19.8, DeepSeek-V3-0324 7.5, Claude Sonnet-4 24.8, GPT-4.1 '-'。

打开官方来源

multi-swe-bench 27 模型 qwen3-coder-480b-a35b-instruct · 版本 flash · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Multi-SWE-bench flash · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multi-swebench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=27.0. Competitor cells (vision): Kimi-K2-Instruct 20.7, Claude-Sonnet-4 25.0. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 20.7, DeepSeek-V3 '-', Claude Sonnet-4 25.0, GPT-4.1 '-'。

打开官方来源

aider 61.8 模型 qwen3-coder-480b-a35b-instruct · 版本 polyglot · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Aider-Polyglot · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=61.8. Competitor cells (vision): Kimi-K2-Instruct 60.0, DeepSeek-V3-0324 56.9, Claude-Sonnet-4 56.4, GPT-4.1 52.4. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 60.0, DeepSeek-V3-0324 56.9, Claude Sonnet-4 56.4, OpenAI GPT-4.1 52.4。

打开官方来源

spider2 31.1 模型 qwen3-coder-480b-a35b-instruct · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Spider2 · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: spider2 not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=31.1. Competitor cells (vision): Kimi-K2-Instruct 25.2, DeepSeek-V3-0324 12.8, Claude-Sonnet-4 31.1, GPT-4.1 16.5. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 25.2, DeepSeek-V3-0324 12.8, Claude Sonnet-4 31.1, OpenAI GPT-4.1 16.5。

打开官方来源

webarena 49.9 模型 qwen3-coder-480b-a35b-instruct · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: WebArena · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=49.9. Agentic Browser Use section. Competitor cells (vision): Kimi-K2-Instruct 47.4, DeepSeek-V3-0324 40.0, Claude-Sonnet-4 51.1, GPT-4.1 44.3. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 47.4, DeepSeek-V3-0324 40.0, Claude Sonnet-4 51.1, OpenAI GPT-4.1 44.3。

打开官方来源

mind2web 55.8 模型 qwen3-coder-480b-a35b-instruct · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Mind2Web · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mind2web not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=55.8. Agentic Browser Use section. Competitor cells (vision): Kimi-K2-Instruct 42.7, DeepSeek-V3-0324 36.0, Claude-Sonnet-4 47.4, GPT-4.1 49.6. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 42.7, DeepSeek-V3-0324 36.0, Claude Sonnet-4 47.4, OpenAI GPT-4.1 49.6。

打开官方来源

bfcl 68.7 模型 qwen3-coder-480b-a35b-instruct · 版本 v3 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: BFCL-v3 · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=68.7. Agentic Tool Use section. Competitor cells (vision): Kimi-K2-Instruct 65.2, DeepSeek-V3-0324 64.7, Claude-Sonnet-4 73.3, GPT-4.1 62.9. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 65.2, DeepSeek-V3-0324 64.7, Claude Sonnet-4 73.3, OpenAI GPT-4.1 62.9。

打开官方来源

tau-bench 77.5 模型 qwen3-coder-480b-a35b-instruct · 版本 retail · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: TAU-Bench Retail · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=77.5. Agentic Tool Use section. Competitor cells (vision): Kimi-K2-Instruct 70.7, DeepSeek-V3-0324 59.1, Claude-Sonnet-4 80.5. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 70.7, DeepSeek-V3-0324 59.1, Claude Sonnet-4 80.5, GPT-4.1 '-'。

打开官方来源

tau-bench 60 模型 qwen3-coder-480b-a35b-instruct · 版本 airline · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: TAU-Bench Airline · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=60.0. Agentic Tool Use section. Competitor cells (vision): Kimi-K2-Instruct 53.5, DeepSeek-V3-0324 40.0, Claude-Sonnet-4 60.0. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 53.5, DeepSeek-V3-0324 40.0, Claude Sonnet-4 60.0, GPT-4.1 '-'。

打开官方来源