← 模型目录

Qwen3.7-Max

Alibaba / Qwen · 2026-05(仅精确到月) · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Qwen3.7-Max

发布文以“智能体新前沿”为题,将 Qwen3.7-Max 定位为面向智能体时代的新一代旗舰(经阿里云百炼 API 提供,在 Claude Code/OpenClaw/Qwen Code 间泛化)。40 余项评测覆盖智能体编码、工具与办公场景及数理推理,亮点为 HMMT 2026 二月场 97.1 与 SWE-bench Verified 80.4。

输入模态
文本
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

terminalbench 69.7 模型 qwen3-7-max · 版本 2.0-Terminus · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Terminal Bench 2.0-Terminus

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 65.4, K2.6 66.7, GLM-5.1 63.5, DS-V4-Pro 67.9, Qwen3.6-Plus 61.6.

打开官方来源

swebench 80.4 模型 qwen3-7-max · 版本 Verified · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SWE-Verified

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 80.8, K2.6 80.2, DS-V4-Pro 80.6, Qwen3.6-Plus 78.8.

打开官方来源

swebench-pro 60.6 模型 qwen3-7-max · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SWE-Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 57.3, K2.6 59.5, GLM-5.1 58.8, DS-V4-Pro 59.0, Qwen3.6-Plus 56.6. Matches GLM-5.2's Qwen3.7-Max column (60.6). Discrepancy: the GLM-5.1 cell here is 58.8 while GLM-5.1's own page prints 58.4 - unresolved cross-vendor difference.

打开官方来源

swebench-multilingual 78.3 模型 qwen3-7-max · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SWE-Multilingual

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-multilingual. Competitor cells: Opus-4.6 77.5, K2.6 76.7, DS-V4-Pro 76.2, Qwen3.6-Plus 73.8.

打开官方来源

nl2repo 47.2 模型 qwen3-7-max · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: NL2repo

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 47.6, K2.6 42.8, GLM-5.1 41.0, DS-V4-Pro 35.5, Qwen3.6-Plus 34.4. Matches GLM-5.2's column (47.2). Discrepancy: the GLM-5.1 cell here is 41.0 while GLM-5.1's own page prints 42.7 - unresolved cross-vendor difference.

打开官方来源

scicode 53.5 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SciCode

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

scicode id already introduced by prior batches. Competitor cells: Opus-4.6 51.9, K2.6 52.2, GLM-5.1 45.1, Qwen3.6-Plus 41.4.

打开官方来源

qwen-webdev-bench 1568 模型 qwen3-7-max · 版本 未说明 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: QwenWebDev

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-webdev-bench (Qwen in-house web-dev generation, Elo-scored; likely the same instrument printed as QwenReactBench 1538 in qwen3-8-max - naming varies across posts, values consistent in magnitude). Elo unit. Competitor cells: Opus-4.6 1617, GLM-5.1 1564, DS-V4-Pro 1570, Qwen3.6-Plus 1500.

打开官方来源

qwen-svg-bench 1608 模型 qwen3-7-max · 版本 未说明 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: QwenSVG

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-svg-bench (also in qwen3-8-max.json). Competitor cells: Opus-4.6 1541, K2.6 1325, GLM-5.1 1605, DS-V4-Pro 1506, Qwen3.6-Plus 1432.

打开官方来源

qwen-claw 64.3 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Qwenclaw

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-claw (Qwen in-house agentic claw suite) not yet in data/benchmarks/. Competitor cells: Opus-4.6 65.5, K2.6 54.7, GLM-5.1 58.7, DS-V4-Pro 59.2, Qwen3.6-Plus 57.2.

打开官方来源

coworkbench 67.2 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: CoWorkBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: coworkbench. Competitor cells: Opus-4.6 68.2, K2.6 58.2, GLM-5.1 66.0, DS-V4-Pro 66.3, Qwen3.6-Plus 64.5.

打开官方来源

claw-eval 65.2 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: ClawEval

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: claweval (text-mode sibling of claweval-mm) not yet in data/benchmarks/. Competitor cells: Opus-4.6 70.4, K2.6 61.5, GLM-5.1 62.7, DS-V4-Pro 58.4, Qwen3.6-Plus 57.1.

打开官方来源

skillsbench 59.2 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Skillsbench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: skillsbench (also in qwen3-8-max.json). Competitor cells: K2.6 56.2, GLM-5.1 53.1, DS-V4-Pro 52.3, Qwen3.6-Plus 45.7.

打开官方来源

bfcl 75 模型 qwen3-7-max · 版本 V4 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: BFCL-V4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

bfcl id exists in data/benchmarks/ (V3 in prior batches); V4 recorded as variant. Competitor cells: Opus-4.6 76.7, K2.6 71.3, GLM-5.1 70.9, DS-V4-Pro 70.6, Qwen3.6-Plus 68.9.

打开官方来源

mcp-mark 60.8 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MCP-Mark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mcp-mark (also in deepseek-v3-2.json this batch). Competitor cells: Opus-4.6 56.7, K2.6 55.9, GLM-5.1 57.5, DS-V4-Pro 57.1, Qwen3.6-Plus 48.2.

打开官方来源

mcp-atlas 76.4 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MCP-Atlas

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 75.8, K2.6 66.6, GLM-5.1 71.8, DS-V4-Pro 73.6, Qwen3.6-Plus 74.1. Matches GLM-5.2's column (76.4).

打开官方来源

vitabench 47.9 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Vitabench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: vitabench not yet in data/benchmarks/. Competitor cells: K2.6 39.1, GLM-5.1 45.1, DS-V4-Pro 51.9, Qwen3.6-Plus 42.8.

打开官方来源

spreadsheetbench 87 模型 qwen3-7-max · 版本 v1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SpreadSheetBench-v1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: spreadsheetbench already introduced by prior batches (v2 in Kimi K3). Competitor cells: Opus-4.6 89.3, K2.6 84.5, GLM-5.1 85.2, DS-V4-Pro 84.9, Qwen3.6-Plus 80.2.

打开官方来源

korb 1.98 / 96% 模型 qwen3-7-max · 版本 L3, speedup/correctness dual · 指标 geometric_mean_speedup · 单位 speedup_x 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Kernel Bench L3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: kernelbench not yet in data/benchmarks/. Dual cell: speedup 1.98x AND correctness 96%. Competitor cells: Opus-4.6 2.63/98%, K2.6 1.41/80%, GLM-5.1 2.00/78%, DS-V4-Pro 1.07/54%, Qwen3.6-Plus 1.03/48%. Speedup unit is x, not percent - never aggregate with percent rows.

打开官方来源

hlehle 53.5 模型 qwen3-7-max · 版本 w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: HLE w/ tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 53.0, K2.6 54.0, GLM-5.1 52.3, DS-V4-Pro 48.2, Qwen3.6-Plus 50.2. Matches GLM-5.2's column (53.5).

打开官方来源

qwen-world-bench 57.3 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: QwenWorldBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-world-bench (Qwen in-house world/agentic suite) not yet in data/benchmarks/. Competitor cells: Opus-4.6 56.1, K2.6 50.9, GLM-5.1 50.2, DS-V4-Pro 52.3, Qwen3.6-Plus 47.6.

打开官方来源

gpqa 92.4 模型 qwen3-7-max · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 91.3, K2.6 90.5, GLM-5.1 86.2, DS-V4-Pro 90.1, Qwen3.6-Plus 90.4. Matches GLM-5.2's column (92.4).

打开官方来源

hlehle 41.4 模型 qwen3-7-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: HLE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 40.0, K2.6 36.4, GLM-5.1 34.7, DS-V4-Pro 37.7, Qwen3.6-Plus 28.8. Matches GLM-5.2's column (41.4). Discrepancy: the GLM-5.1 cell here is 34.7 while GLM-5.1's own page (and GLM-5.2's GLM-5.1 column) prints 31.0 - unresolved cross-vendor difference.

打开官方来源

lcb 91.6 模型 qwen3-7-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: LiveCodeBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 88.8, K2.6 89.6, DS-V4-Pro 93.5, Qwen3.6-Plus 87.1.

打开官方来源

hmmt-26 97.1 模型 qwen3-7-max · 版本 February 2026 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: HMMT 2026 Feb

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hmmt-26. Competitor cells: Opus-4.6 96.2, K2.6 92.7, GLM-5.1 89.4, DS-V4-Pro 95.2, Qwen3.6-Plus 87.8. Matches GLM-5.2's column (97.1). Discrepancy: the GLM-5.1 cell here is 89.4 while GLM-5.1's own page (and GLM-5.2's GLM-5.1 column) prints 82.6 - unresolved cross-vendor difference.

打开官方来源

imo-answerbench 90 模型 qwen3-7-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: IMOAnswerBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: K2.6 86.0, GLM-5.1 83.8, DS-V4-Pro 89.8, Opus-4.6 75.3. Matches GLM-5.2's column (90.0).

打开官方来源

critpt 11.4 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: CritPT

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: critpt (also in glm-5-2.json this batch). Competitor cells: Opus-4.6 12.6, K2.6 8.0, GLM-5.1 4.6, DS-V4-Pro 12.9, Qwen3.6-Plus 2.9.

打开官方来源

apex 44.5 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Apex

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: apex (also in deepseek-v4.json). Competitor cells: Opus-4.6 34.5, K2.6 24.0, GLM-5.1 11.5, DS-V4-Pro 38.3, Qwen3.6-Plus 8.8.

打开官方来源

mmlu-pro 89.6 模型 qwen3-7-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMLU-Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 89.7, K2.6 87.1, GLM-5.1 86.3, DS-V4-Pro 87.5, Qwen3.6-Plus 88.5. Minor discrepancy: the GLM-5.1 cell here is 86.3 while GLM-5.2's GLM-5.1 column prints 86.2.

打开官方来源

mmlu-redux 95 模型 qwen3-7-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMLU-Redux

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 95.2, K2.6 95.3, GLM-5.1 94.3, DS-V4-Pro 94.8, Qwen3.6-Plus 94.5.

打开官方来源

supergpqa 73.6 模型 qwen3-7-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SuperGPQA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: supergpqa. Competitor cells: Opus-4.6 72.5, K2.6 71.3, GLM-5.1 68.0, DS-V4-Pro 69.9, Qwen3.6-Plus 71.6.

打开官方来源

ifeval 94.3 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: IFEval

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus-4.6 91.9, K2.6 94.5, GLM-5.1 94.5, DS-V4-Pro 91.9, Qwen3.6-Plus 94.3.

打开官方来源

ifbench 79.1 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: IFBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ifbench (also in qwen3-8-max.json). Competitor cells: Opus-4.6 62.5, K2.6 76.0, GLM-5.1 76.0, DS-V4-Pro 77.0, Qwen3.6-Plus 74.2.

打开官方来源

mrcr 90.4 模型 qwen3-7-max · 版本 128k · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MRCR-v2 128k

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mrcr-v2 (256K variant in qwen3-8-max.json). Competitor cells: Opus-4.6 84.0, K2.6 63.1, GLM-5.1 62.0, DS-V4-Pro 74.4, Qwen3.6-Plus 85.9.

打开官方来源

wmt24 85.8 模型 qwen3-7-max · 版本 ++ · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: WMT24++

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: wmt24++ (machine translation) not yet in data/benchmarks/. Competitor cells: Opus-4.6 82.7, K2.6 81.6, GLM-5.1 81.8, DS-V4-Pro 82.2, Qwen3.6-Plus 84.3.

打开官方来源

maxife 89.2 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MAXIFE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: maxife (multilingual instruction following) not yet in data/benchmarks/. Competitor cells: Opus-4.6 81.3, K2.6 87.7, GLM-5.1 87.7, DS-V4-Pro 88.9, Qwen3.6-Plus 88.2.

打开官方来源

mmmlu 90.3 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMMLU

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmlu (also in qwen3-8-flash.json base table). Competitor cells: Opus-4.6 90.6, K2.6 87.5, GLM-5.1 87.2, DS-V4-Pro 87.9, Qwen3.6-Plus 89.5.

打开官方来源

mmlu-prox 87 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMLU-ProX

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmlu-prox (multilingual MMLU-Pro) not yet in data/benchmarks/. Competitor cells: Opus-4.6 86.1, K2.6 83.7, GLM-5.1 83.9, DS-V4-Pro 83.9, Qwen3.6-Plus 84.7.

打开官方来源

nova-63 59 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: NOVA-63

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: nova-63 not yet in data/benchmarks/. Competitor cells: Opus-4.6 59.1, K2.6 56.7, GLM-5.1 54.6, DS-V4-Pro 52.8, Qwen3.6-Plus 57.9.

打开官方来源

include 86.2 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: INCLUDE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: include (also in qwen3-8-flash.json base table). Competitor cells: Opus-4.6 87.4, K2.6 84.2, GLM-5.1 84.3, DS-V4-Pro 86.1, Qwen3.6-Plus 85.1.

打开官方来源

global-piqa 91.4 模型 qwen3-7-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Global PIQA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: global-piqa not yet in data/benchmarks/. Competitor cells: Opus-4.6 91.2, K2.6 89.2, GLM-5.1 89.5, DS-V4-Pro 90.5, Qwen3.6-Plus 89.8.

打开官方来源

polymath-en 86.5 模型 qwen3-7-max · 版本 multilingual edition · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: PolyMATH

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

polymath-en id already introduced by prior batches (Kimi K2's PolyMath-en); this page's PolyMATH may be the multilingual expansion - mapped to polymath-en with variant note pending confirmation. Competitor cells: Opus-4.6 80.2, K2.6 82.7, GLM-5.1 67.6, DS-V4-Pro 72.0, Qwen3.6-Plus 77.4.

打开官方来源