Qwen3-Coder-480B-A35B-Instruct
Alibaba / Qwen · 2025-07-22 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Qwen3-Coder-480B-A35B-Instruct
Qwen3-Coder-480B-A35B-Instruct 以「Agentic Coding in the World」为定位发布,480B MoE/35B 激活,原生 256K 上下文(YaRN 可扩至 1M)。已收录 15 项评测集中于代码仓库、终端与浏览器/工具 Agent:亮点 SWE-bench Verified(OpenHands 500 turns)69.6、Terminal-Bench 37.5。
- 输入模态
- 文本
- 上下文
- 256K(YaRN 可扩至 1M)
- 参数
- 480B-A35B MoE
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Terminal-Bench · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=37.5. Chart label has no version qualifier ('Terminal-Bench'); do not merge with Terminal-Bench 2.x rows of other vendors without variant confirmation. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 30.0, DeepSeek-V3-0324 2.5, Claude Sonnet-4 35.5, OpenAI GPT-4.1 25.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Verified w/ OpenHands, 500 turns · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=69.6. Competitor cells (vision): Claude-Sonnet-4 70.4. Prose: 'Qwen3-Coder achieves state-of-the-art performance among open-source models on SWE-Bench Verified without test-time scaling' (section Post-Training / Scaling Long-Horizon RL); scatter image swe.jpg repeats the 69.6/67.0 points. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2 '-', DeepSeek-V3 '-', Claude Sonnet-4 70.4, GPT-4.1 '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Verified w/ OpenHands, 100 turns · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=67.0. Competitor cells (vision): Kimi-K2-Instruct 65.4, DeepSeek-V3-0324 38.8, Claude-Sonnet-4 68.0, GPT-4.1 48.6. Prose: 'Qwen3-Coder achieves state-of-the-art performance among open-source models on SWE-Bench Verified without test-time scaling' (section Post-Training / Scaling Long-Horizon RL); scatter image swe.jpg repeats the 69.6/67.0 points. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 65.4, DeepSeek-V3-0324 38.8, Claude Sonnet-4 68.0, OpenAI GPT-4.1 48.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Verified w/ Private Scaffolding · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=None. Chart shows no value for Qwen3-Coder in this row (dash); competitor cells (vision): Kimi-K2-Instruct 65.8, Claude-Sonnet-4 72.7, GPT-4.1 63.8. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Qwen3-Coder 该格为 '-';Kimi-K2-Instruct 65.8, DeepSeek-V3 '-', Claude Sonnet-4 72.7, OpenAI GPT-4.1 63.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Live · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-live not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=26.3. Competitor cells (vision): Kimi-K2-Instruct 22.3, DeepSeek-V3-0324 13.0, Claude-Sonnet-4 27.7. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 22.3, DeepSeek-V3-0324 13.0, Claude Sonnet-4 27.7, GPT-4.1 '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: SWE-bench Multilingual · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=54.7. Competitor cells (vision): Kimi-K2-Instruct 47.3, DeepSeek-V3-0324 13.0, Claude-Sonnet-4 53.3, GPT-4.1 31.5. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 47.3, DeepSeek-V3-0324 13.0, Claude Sonnet-4 53.3, OpenAI GPT-4.1 31.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Multi-SWE-bench mini · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multi-swebench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=25.8. Competitor cells (vision): Kimi-K2-Instruct 19.8, DeepSeek-V3-0324 7.5, Claude-Sonnet-4 24.8. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 19.8, DeepSeek-V3-0324 7.5, Claude Sonnet-4 24.8, GPT-4.1 '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Multi-SWE-bench flash · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multi-swebench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=27.0. Competitor cells (vision): Kimi-K2-Instruct 20.7, Claude-Sonnet-4 25.0. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 20.7, DeepSeek-V3 '-', Claude Sonnet-4 25.0, GPT-4.1 '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Aider-Polyglot · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=61.8. Competitor cells (vision): Kimi-K2-Instruct 60.0, DeepSeek-V3-0324 56.9, Claude-Sonnet-4 56.4, GPT-4.1 52.4. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 60.0, DeepSeek-V3-0324 56.9, Claude Sonnet-4 56.4, OpenAI GPT-4.1 52.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Spider2 · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: spider2 not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=31.1. Competitor cells (vision): Kimi-K2-Instruct 25.2, DeepSeek-V3-0324 12.8, Claude-Sonnet-4 31.1, GPT-4.1 16.5. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 25.2, DeepSeek-V3-0324 12.8, Claude Sonnet-4 31.1, OpenAI GPT-4.1 16.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: WebArena · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=49.9. Agentic Browser Use section. Competitor cells (vision): Kimi-K2-Instruct 47.4, DeepSeek-V3-0324 40.0, Claude-Sonnet-4 51.1, GPT-4.1 44.3. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 47.4, DeepSeek-V3-0324 40.0, Claude Sonnet-4 51.1, OpenAI GPT-4.1 44.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: Mind2Web · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mind2web not yet in data/benchmarks.json. Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=55.8. Agentic Browser Use section. Competitor cells (vision): Kimi-K2-Instruct 42.7, DeepSeek-V3-0324 36.0, Claude-Sonnet-4 47.4, GPT-4.1 49.6. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 42.7, DeepSeek-V3-0324 36.0, Claude Sonnet-4 47.4, OpenAI GPT-4.1 49.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: BFCL-v3 · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=68.7. Agentic Tool Use section. Competitor cells (vision): Kimi-K2-Instruct 65.2, DeepSeek-V3-0324 64.7, Claude-Sonnet-4 73.3, GPT-4.1 62.9. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 65.2, DeepSeek-V3-0324 64.7, Claude Sonnet-4 73.3, OpenAI GPT-4.1 62.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: TAU-Bench Retail · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=77.5. Agentic Tool Use section. Competitor cells (vision): Kimi-K2-Instruct 70.7, DeepSeek-V3-0324 59.1, Claude-Sonnet-4 80.5. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 70.7, DeepSeek-V3-0324 59.1, Claude Sonnet-4 80.5, GPT-4.1 '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Page hero benchmark chart (between the intro paragraph and the 'Qwen3-Coder' heading) · row: TAU-Bench Airline · figure: images/02.jpg (archive of qianwen-res Qwen3-Coder/qwen3-coder-main.jpg 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Qwen3-Coder-480B-A35B-Instruct=60.0. Agentic Tool Use section. Competitor cells (vision): Kimi-K2-Instruct 53.5, DeepSeek-V3-0324 40.0, Claude-Sonnet-4 60.0. 视觉转写自归档图 images/02.jpg(2026-09-01 复核,与先前读数一致)。同表竞品列(Kimi-K2 / DeepSeek-V3-0324 / Claude Sonnet-4 / GPT-4.1):Kimi-K2-Instruct 53.5, DeepSeek-V3-0324 40.0, Claude Sonnet-4 60.0, GPT-4.1 '-'。