Qwen3.7-Max
Alibaba / Qwen · 2026-05(仅精确到月) · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Qwen3.7-Max
发布文以“智能体新前沿”为题,将 Qwen3.7-Max 定位为面向智能体时代的新一代旗舰(经阿里云百炼 API 提供,在 Claude Code/OpenClaw/Qwen Code 间泛化)。40 余项评测覆盖智能体编码、工具与办公场景及数理推理,亮点为 HMMT 2026 二月场 97.1 与 SWE-bench Verified 80.4。
- 输入模态
- 文本
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Terminal Bench 2.0-Terminus
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 65.4, K2.6 66.7, GLM-5.1 63.5, DS-V4-Pro 67.9, Qwen3.6-Plus 61.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SWE-Verified
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 80.8, K2.6 80.2, DS-V4-Pro 80.6, Qwen3.6-Plus 78.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SWE-Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 57.3, K2.6 59.5, GLM-5.1 58.8, DS-V4-Pro 59.0, Qwen3.6-Plus 56.6. Matches GLM-5.2's Qwen3.7-Max column (60.6). Discrepancy: the GLM-5.1 cell here is 58.8 while GLM-5.1's own page prints 58.4 - unresolved cross-vendor difference.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SWE-Multilingual
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual. Competitor cells: Opus-4.6 77.5, K2.6 76.7, DS-V4-Pro 76.2, Qwen3.6-Plus 73.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: NL2repo
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 47.6, K2.6 42.8, GLM-5.1 41.0, DS-V4-Pro 35.5, Qwen3.6-Plus 34.4. Matches GLM-5.2's column (47.2). Discrepancy: the GLM-5.1 cell here is 41.0 while GLM-5.1's own page prints 42.7 - unresolved cross-vendor difference.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SciCode
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}scicode id already introduced by prior batches. Competitor cells: Opus-4.6 51.9, K2.6 52.2, GLM-5.1 45.1, Qwen3.6-Plus 41.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: QwenWebDev
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-webdev-bench (Qwen in-house web-dev generation, Elo-scored; likely the same instrument printed as QwenReactBench 1538 in qwen3-8-max - naming varies across posts, values consistent in magnitude). Elo unit. Competitor cells: Opus-4.6 1617, GLM-5.1 1564, DS-V4-Pro 1570, Qwen3.6-Plus 1500.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: QwenSVG
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-svg-bench (also in qwen3-8-max.json). Competitor cells: Opus-4.6 1541, K2.6 1325, GLM-5.1 1605, DS-V4-Pro 1506, Qwen3.6-Plus 1432.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Qwenclaw
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-claw (Qwen in-house agentic claw suite) not yet in data/benchmarks/. Competitor cells: Opus-4.6 65.5, K2.6 54.7, GLM-5.1 58.7, DS-V4-Pro 59.2, Qwen3.6-Plus 57.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: CoWorkBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: coworkbench. Competitor cells: Opus-4.6 68.2, K2.6 58.2, GLM-5.1 66.0, DS-V4-Pro 66.3, Qwen3.6-Plus 64.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: ClawEval
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: claweval (text-mode sibling of claweval-mm) not yet in data/benchmarks/. Competitor cells: Opus-4.6 70.4, K2.6 61.5, GLM-5.1 62.7, DS-V4-Pro 58.4, Qwen3.6-Plus 57.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Skillsbench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: skillsbench (also in qwen3-8-max.json). Competitor cells: K2.6 56.2, GLM-5.1 53.1, DS-V4-Pro 52.3, Qwen3.6-Plus 45.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: BFCL-V4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}bfcl id exists in data/benchmarks/ (V3 in prior batches); V4 recorded as variant. Competitor cells: Opus-4.6 76.7, K2.6 71.3, GLM-5.1 70.9, DS-V4-Pro 70.6, Qwen3.6-Plus 68.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MCP-Mark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mcp-mark (also in deepseek-v3-2.json this batch). Competitor cells: Opus-4.6 56.7, K2.6 55.9, GLM-5.1 57.5, DS-V4-Pro 57.1, Qwen3.6-Plus 48.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MCP-Atlas
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 75.8, K2.6 66.6, GLM-5.1 71.8, DS-V4-Pro 73.6, Qwen3.6-Plus 74.1. Matches GLM-5.2's column (76.4).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Vitabench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vitabench not yet in data/benchmarks/. Competitor cells: K2.6 39.1, GLM-5.1 45.1, DS-V4-Pro 51.9, Qwen3.6-Plus 42.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SpreadSheetBench-v1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: spreadsheetbench already introduced by prior batches (v2 in Kimi K3). Competitor cells: Opus-4.6 89.3, K2.6 84.5, GLM-5.1 85.2, DS-V4-Pro 84.9, Qwen3.6-Plus 80.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Kernel Bench L3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: kernelbench not yet in data/benchmarks/. Dual cell: speedup 1.98x AND correctness 96%. Competitor cells: Opus-4.6 2.63/98%, K2.6 1.41/80%, GLM-5.1 2.00/78%, DS-V4-Pro 1.07/54%, Qwen3.6-Plus 1.03/48%. Speedup unit is x, not percent - never aggregate with percent rows.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: HLE w/ tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 53.0, K2.6 54.0, GLM-5.1 52.3, DS-V4-Pro 48.2, Qwen3.6-Plus 50.2. Matches GLM-5.2's column (53.5).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: QwenWorldBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-world-bench (Qwen in-house world/agentic suite) not yet in data/benchmarks/. Competitor cells: Opus-4.6 56.1, K2.6 50.9, GLM-5.1 50.2, DS-V4-Pro 52.3, Qwen3.6-Plus 47.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 91.3, K2.6 90.5, GLM-5.1 86.2, DS-V4-Pro 90.1, Qwen3.6-Plus 90.4. Matches GLM-5.2's column (92.4).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: HLE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 40.0, K2.6 36.4, GLM-5.1 34.7, DS-V4-Pro 37.7, Qwen3.6-Plus 28.8. Matches GLM-5.2's column (41.4). Discrepancy: the GLM-5.1 cell here is 34.7 while GLM-5.1's own page (and GLM-5.2's GLM-5.1 column) prints 31.0 - unresolved cross-vendor difference.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: LiveCodeBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 88.8, K2.6 89.6, DS-V4-Pro 93.5, Qwen3.6-Plus 87.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: HMMT 2026 Feb
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hmmt-26. Competitor cells: Opus-4.6 96.2, K2.6 92.7, GLM-5.1 89.4, DS-V4-Pro 95.2, Qwen3.6-Plus 87.8. Matches GLM-5.2's column (97.1). Discrepancy: the GLM-5.1 cell here is 89.4 while GLM-5.1's own page (and GLM-5.2's GLM-5.1 column) prints 82.6 - unresolved cross-vendor difference.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: IMOAnswerBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: K2.6 86.0, GLM-5.1 83.8, DS-V4-Pro 89.8, Opus-4.6 75.3. Matches GLM-5.2's column (90.0).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: CritPT
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: critpt (also in glm-5-2.json this batch). Competitor cells: Opus-4.6 12.6, K2.6 8.0, GLM-5.1 4.6, DS-V4-Pro 12.9, Qwen3.6-Plus 2.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Apex
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: apex (also in deepseek-v4.json). Competitor cells: Opus-4.6 34.5, K2.6 24.0, GLM-5.1 11.5, DS-V4-Pro 38.3, Qwen3.6-Plus 8.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMLU-Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 89.7, K2.6 87.1, GLM-5.1 86.3, DS-V4-Pro 87.5, Qwen3.6-Plus 88.5. Minor discrepancy: the GLM-5.1 cell here is 86.3 while GLM-5.2's GLM-5.1 column prints 86.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMLU-Redux
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 95.2, K2.6 95.3, GLM-5.1 94.3, DS-V4-Pro 94.8, Qwen3.6-Plus 94.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: SuperGPQA
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: supergpqa. Competitor cells: Opus-4.6 72.5, K2.6 71.3, GLM-5.1 68.0, DS-V4-Pro 69.9, Qwen3.6-Plus 71.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: IFEval
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus-4.6 91.9, K2.6 94.5, GLM-5.1 94.5, DS-V4-Pro 91.9, Qwen3.6-Plus 94.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: IFBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ifbench (also in qwen3-8-max.json). Competitor cells: Opus-4.6 62.5, K2.6 76.0, GLM-5.1 76.0, DS-V4-Pro 77.0, Qwen3.6-Plus 74.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MRCR-v2 128k
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mrcr-v2 (256K variant in qwen3-8-max.json). Competitor cells: Opus-4.6 84.0, K2.6 63.1, GLM-5.1 62.0, DS-V4-Pro 74.4, Qwen3.6-Plus 85.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: WMT24++
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: wmt24++ (machine translation) not yet in data/benchmarks/. Competitor cells: Opus-4.6 82.7, K2.6 81.6, GLM-5.1 81.8, DS-V4-Pro 82.2, Qwen3.6-Plus 84.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MAXIFE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: maxife (multilingual instruction following) not yet in data/benchmarks/. Competitor cells: Opus-4.6 81.3, K2.6 87.7, GLM-5.1 87.7, DS-V4-Pro 88.9, Qwen3.6-Plus 88.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMMLU
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmlu (also in qwen3-8-flash.json base table). Competitor cells: Opus-4.6 90.6, K2.6 87.5, GLM-5.1 87.2, DS-V4-Pro 87.9, Qwen3.6-Plus 89.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: MMLU-ProX
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmlu-prox (multilingual MMLU-Pro) not yet in data/benchmarks/. Competitor cells: Opus-4.6 86.1, K2.6 83.7, GLM-5.1 83.9, DS-V4-Pro 83.9, Qwen3.6-Plus 84.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: NOVA-63
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: nova-63 not yet in data/benchmarks/. Competitor cells: Opus-4.6 59.1, K2.6 56.7, GLM-5.1 54.6, DS-V4-Pro 52.8, Qwen3.6-Plus 57.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: INCLUDE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: include (also in qwen3-8-flash.json base table). Competitor cells: Opus-4.6 87.4, K2.6 84.2, GLM-5.1 84.3, DS-V4-Pro 86.1, Qwen3.6-Plus 85.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: Global PIQA
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: global-piqa not yet in data/benchmarks/. Competitor cells: Opus-4.6 91.2, K2.6 89.2, GLM-5.1 89.5, DS-V4-Pro 90.5, Qwen3.6-Plus 89.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 (DOM table) · table: DOM table 'Coding Agent / General Agent / STEM & Reasoning / General Capability / Multilingualism' (41 benchmark rows x 6 models, machine-readable) · row: PolyMATH
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}polymath-en id already introduced by prior batches (Kimi K2's PolyMath-en); this page's PolyMATH may be the multilingual expansion - mapped to polymath-en with variant note pending confirmation. Competitor cells: Opus-4.6 80.2, K2.6 82.7, GLM-5.1 67.6, DS-V4-Pro 72.0, Qwen3.6-Plus 77.4.