← 模型目录

Qwen3.8-Max

Alibaba / Qwen · 2026-08-03 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Qwen3.8-Max

通义将 Qwen3.8-Max 定位为当前 Qwen 家族最强大的模型,基于 Qwen 3.5 架构将参数规模扩展至 2.4 万亿(激活参数 95B),在编程、办公、科研与长周期任务上全面提升,并成为首个开源权重的 Qwen-Max 级模型(权重于下周发布)。已收录评测覆盖智能体编码与办公、多模态理解与操作、视频理解及长上下文等领域,Terminal-Bench 2.1 86.6、PaperBench 93。

输入模态
文本 / 图像 / 视频
上下文
官方资料未说明
参数
2.4T(激活 95B)
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

terminalbench 86.6 模型 qwen3-8-max · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Terminal Bench 2.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus4.8 84.6, Fable5 84.6, GPT5.6 Sol 88.8, Qwen3.7-Max 74.5. Harness not printed on this page.

打开官方来源

swebench-pro 67.7 模型 qwen3-8-max · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: SWE-bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus4.8 69.2, Fable5 80.0, GPT5.6 64.6, Qwen3.7-Max 60.6.

打开官方来源

deepswe 56.6 模型 qwen3-8-max · 版本 1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: DeepSWE 1.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

deepswe id already introduced by prior batches. Competitor cells: Opus4.8 59.0, Fable5 70.0, GPT5.6 73.0, Qwen3.7-Max 21.6.

打开官方来源

nl2repo 55.9 模型 qwen3-8-max · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: NL2Repo-Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

nl2repo id already introduced by prior batches. Competitor cells: Opus4.8 69.4, Qwen3.7-Max 47.2.

打开官方来源

frontierswe 73.5 模型 qwen3-8-max · 版本 未说明 · 指标 dominance_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: FrontierSWE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

frontierswe id already introduced by prior batches. Competitor cells: Opus4.8 70.0, Fable5 88.8, Qwen3.7-Max 40.7. Snapshot date not printed (GLM-5.2's row was as-of 2026/06/16).

打开官方来源

mls-bench-lite 41 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: MLS-Bench-Lite

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mls-bench-lite already introduced by prior batches. Competitor cells: Opus4.8 42.8, Fable5 49.9, GPT5.6 46.2, Qwen3.7-Max 31.7.

打开官方来源

paperbench 93 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: PaperBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: paperbench already introduced by prior batches. Competitor cells: Opus4.8 80.3, Fable5 88.8, GPT5.6 90.5, Qwen3.7-Max 64.8.

打开官方来源

androidbench 75.1 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: AndroidBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: androidbench not yet in data/benchmarks/. Competitor cells: Opus4.8 69.8, Fable5 84.5, GPT5.6 74.0, Qwen3.7-Max 56.5.

打开官方来源

qwen-swe-bench 80.7 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenSWEBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-swe-bench (Qwen in-house SWE bench) not yet in data/benchmarks/. Vendor-owned bench; treat as in-house. Competitor cells: Opus4.8 84.0, Fable5 86.3, GPT5.6 73.5, Qwen3.7-Max 63.4.

打开官方来源

qwen-qoder-bench 58.4 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenQoderBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-qoder-bench (Qwen in-house IDE coding bench) not yet in data/benchmarks/. Competitor cells: Opus4.8 62.7, Fable5 63.1, GPT5.6 53.8, Qwen3.7-Max 36.8.

打开官方来源

qwen-react-bench 1724 模型 qwen3-8-max · 版本 未说明 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenReactBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-react-bench (React front-end generation, Elo-scored) not yet in data/benchmarks/. Elo unit. Competitor cells: Opus4.8 1694, Fable5 1770, GPT5.6 1564, Qwen3.7-Max 1538.

打开官方来源

qwen-svg-bench 1713 模型 qwen3-8-max · 版本 未说明 · 指标 elo_rating · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenSVGBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-svg-bench (SVG generation, Elo-scored) not yet in data/benchmarks/. Competitor cells: Opus4.8 1648, Fable5 1690, GPT5.6 1758, Qwen3.7-Max 1499.

打开官方来源

coworkbench 74.8 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: CoWorkBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: coworkbench already introduced by prior batches. Competitor cells: Opus4.8 72.3, Fable5 75.9, GPT5.6 71.5, Qwen3.7-Max 64.6.

打开官方来源

workspace-bench 67.7 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: WorkSpaceBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: workspace-bench already introduced by prior batches. Competitor cells: Opus4.8 66.8, Fable5 68.7, GPT5.6 65.6, Qwen3.7-Max 61.4.

打开官方来源

jobbench 53.4 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: JobBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: jobbench already introduced by prior batches. Competitor cells: Opus4.8 48.4, Fable5 57.4, GPT5.6 45.4, Qwen3.7-Max 31.3.

打开官方来源

skillsbench 70.2 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: SkillsBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: skillsbench not yet in data/benchmarks/. Competitor cells: Opus4.8 65.1, Fable5 70.9, GPT5.6 73.5, Qwen3.7-Max 61.2.

打开官方来源

agents-last-exam 27.0 / 52.4 模型 qwen3-8-max · 版本 Pass / Score dual metric · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Agents' Last Exam (Pass / Score)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

agents-last-exam id already introduced by prior batches. Dual metric printed in one cell: Pass@1 27.0 AND Score 52.4 - value field holds Score, display holds both. Competitor cells: Opus4.8 27.0/45.1, GPT5.6 30.6/53.6, Qwen3.7-Max 11.8/31.1.

打开官方来源

automationbench 27.3 模型 qwen3-8-max · 版本 Pass@1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Automation-Bench (Pass@1)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

automationbench id already introduced by prior batches. Competitor cells: Opus4.8 27.2, Fable5 29.1, GPT5.6 29.7, Qwen3.7-Max 14.2.

打开官方来源

toolathlon 72.5 模型 qwen3-8-max · 版本 Verified, Pass@1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Toolathlon Verified (Pass@1)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

toolathlon id already introduced by prior batches. Competitor cells: Opus4.8 76.2, Fable5 77.9, GPT5.6 74.9, Qwen3.7-Max 49.7.

打开官方来源

widesearch 81.9 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: WideSearch

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: widesearch. Competitor cells: Opus4.8 72.9, Fable5 81.2, Qwen3.7-Max 75.2.

打开官方来源

hlehle 56.2 模型 qwen3-8-max · 版本 w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: HLE w/ tools

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus4.8 57.9, Fable5 64.5, GPT5.6 58.0, Qwen3.7-Max 53.5.

打开官方来源

gpqa 92.6 模型 qwen3-8-max · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus4.8 92.0, Fable5 92.6, GPT5.6 94.1, Qwen3.7-Max 92.4.

打开官方来源

hlehle 43.6 模型 qwen3-8-max · 版本 no tools (table default) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: HLE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Opus4.8 45.7, Fable5 53.3, GPT5.6 47.2, Qwen3.7-Max 41.4.

打开官方来源

ifbench 82.8 模型 qwen3-8-max · 版本 IFBench · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: IFBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ifbench (instruction-following bench, distinct from ifeval) not yet in data/benchmarks/. Competitor cells: Opus4.8 62.2, Fable5 63.5, GPT5.6 72.7, Qwen3.7-Max 79.1.

打开官方来源

one-million-bench 52.5 模型 qwen3-8-max · 版本 expert score · 指标 expert_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: $OneMillion-Bench (expert score)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

one-million-bench id already introduced by prior batches. Competitor cells: Opus4.8 41.8, Fable5 55.9, GPT5.6 53.8, Qwen3.7-Max 44.4.

打开官方来源

healthbench 60.2 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: HealthBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: healthbench (also introduced in kimi-k2-thinking.json batch 1). Competitor cells: Opus4.8 52.4, GPT5.6 55.3, Qwen3.7-Max 54.5.

打开官方来源

plawbench 73.2 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: PLawBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: plawbench (legal) not yet in data/benchmarks/. Competitor cells: Opus4.8 69.6, Fable5 70.2, GPT5.6 72.3, Qwen3.7-Max 58.9.

打开官方来源

prbench-finance 58.3 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: PRBench-Finance

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: prbench-finance not yet in data/benchmarks/. Competitor cells: Opus4.8 51.9, Fable5 55.8, GPT5.6 55.5, Qwen3.7-Max 46.8.

打开官方来源

mrcr 92.9 模型 qwen3-8-max · 版本 256K, 8-needle · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: MRCR v2 256K (8-needle)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mrcr-v2 not yet in data/benchmarks/ (distinct from deepseek-v4's MRCR 1M row). Competitor cells: Opus4.8 83.2, GPT5.6 93.8, Qwen3.7-Max 86.7.

打开官方来源

longbench 66.3 模型 qwen3-8-max · 版本 v2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: LongBench v2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

longbench id exists (v1); v2 variant. Competitor cells: Opus4.8 69.1, GPT5.6 67.1, Qwen3.7-Max 65.3.

打开官方来源

mmmu-pro 82.3 模型 qwen3-8-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.6 | 81.2 | 80.5 | 83.0 | 79.0. Footnote 3: Gemini3.1-Pro and GPT5.6-Sol cells taken from official model reports / system cards; all other models internally evaluated.

打开官方来源

mathvision 95.2 / 97.7 模型 qwen3-8-max · 版本 Without CI / With CI dual · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark variant: mathvision exists in data/benchmarks/ (dual-condition variant recorded here). Dual cell per footnote 1: 无 CI / 有 CI (value stores With-CI). Footnote 2: our model uses a fixed prompt (reason step by step, final answer in \boxed{}); competitors reported as the higher of with/without \boxed{} runs. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 87.1/97.1 | 92.7/98.6 | 87.4/95.7 | 90.8/97.8 | 90.3/--.

打开官方来源

babyvision 82.0 / 91.3 模型 qwen3-8-max · 版本 Without CI / With CI dual (footnote 1) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 28.4/81.2 | 42.5/90.5 | 55.9/68.3 | 65.5/88.9 | 64.7/70.4. Footnote 1: reported as 无 CI / 有 CI (value stores With-CI).

打开官方来源

hle-vl 52.2 模型 qwen3-8-max · 版本 w/ Tools (code interpreter + search) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hle-vl (HLE vision-language variant) not yet in data/benchmarks/. Footnote 6: tools = code interpreter (CI) + search; Gemini3.1-Pro / GPT5.6-Sol tool scores measured end-to-end via their native tool-calling APIs. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): -- | -- | 43.9 | 51.2 | 25.6.

打开官方来源

zerobench 24.0 / 49.0 模型 qwen3-8-max · 版本 Pass@5, Without CI / With CI dual · 指标 pass@5 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

zerobench id exists in data/benchmarks/ (dual-condition Pass@5 variant recorded here). Footnote 1: reported as 无 CI / 有 CI (value stores With-CI). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 17.0/34.0 | 20.0/46.0 | 17.0/23.0 | 22.0/35.0 | 19.0/19.0.

打开官方来源

zerobench-sub 48.5 模型 qwen3-8-max · 版本 visual-subset · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: zerobench-sub (ZeroBench sub-question split printed as a separate row) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 31.1 | 37.1 | 36.5 | 46.7 | 41.0.

打开官方来源

logicvista 91.9 模型 qwen3-8-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 76.7 | 85.7 | 82.6 | 89.7 | 84.3.

打开官方来源

hipho 90.0 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hipho not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 69.3 | 78.6 | 85.4 | 86.8 | 84.1.

打开官方来源

phyx 83.5 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: phyx not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 54.2 | 71.7 | 79.4 | 79.1 | 80.0.

打开官方来源

slake 90.8 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: slake (biomedical VQA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.9 | 86.6 | 82.9 | 85.1 | 83.2.

打开官方来源

medxpertqa-mm 80.4 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: medxpertqa-mm not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 71.7 | 80.0 | 80.7 | 81.5 | 71.0.

打开官方来源

pmc-vqa 66.2 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: pmc-vqa not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 59.2 | 63.2 | 62.5 | 62.3 | 63.4.

打开官方来源

osworld 86.1 模型 qwen3-8-max · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

osworld id exists in data/benchmarks/; Verified split recorded as variant. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 83.4 | 85.0 | 76.2 | 83.2 | 73.3.

打开官方来源

osworld 19.4 / 46.7 模型 qwen3-8-max · 版本 2.0, Binary / Partial dual · 指标 partial_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

osworld id exists; 2.0 dual-condition variant. Footnote 7: 二元 / 部分 format (value stores Partial = share of full+partial reward). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 20.6/54.8 | --/66.1 | 7.8/30.6 | --/62.6 | 2.8/21.5.

打开官方来源

screenspot-pro 84.5 模型 qwen3-8-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

screenspot-pro id exists in data/benchmarks/. Footnote 8: Opus4.8 and Fable5 cells taken from official system cards; Fable5's figure refers to the corresponding Mythos Preview score; all other models internally evaluated. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 82.3 | 87.3 | 68.1 | 81.3 | 79.0.

打开官方来源

webarena 66.8 模型 qwen3-8-max · 版本 Verified · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

webarena id exists; Verified split. Footnote 9: scored with the official WebArena scorer inside the OSWorld framework. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.9 | 71.3 | 64.3 | 69.7 | 55.3.

打开官方来源

androidworld 85.3 模型 qwen3-8-max · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

androidworld id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.0 | 88.8 | 70.7 | 77.6 | 81.0.

打开官方来源

mobileworld 77.8 模型 qwen3-8-max · 版本 未说明 · 指标 success_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

mobileworld id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.5 | 85.5 | 58.1 | 76.9 | 51.2.

打开官方来源

claw-eval 77.2 / 74.8 模型 qwen3-8-max · 版本 Pass@3 / Average dual · 指标 average_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

claweval-mm id exists in data/benchmarks/ (dual-metric cell; value stores Average per footnote 4, display stores Pass@3 / Average). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 73.3/73.8 | 81.2/77.5 | 50.5/55.2 | 81.2/78.9 | 57.4/60.1.

打开官方来源

vision2web 69.0 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

vision2web id exists in data/benchmarks/. Footnote 5: score = mean over front-end / web / website categories, Claude Code harness with gpt-5.4-2026-03-05 as judge. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 62.4 | 70.5 | -- | 62.1 | 42.1.

打开官方来源

qwen-blender-bench 69.9 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-blender-bench (Qwen in-house Blender task bench, footnote 13) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 62.4 | 69.5 | 23.0 | 68.6 | 41.5.

打开官方来源

parametric-cad-bench 91.5 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: parametric-cad-bench not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 85.1 | 87.5 | 73.5 | 86.2 | 73.8.

打开官方来源

recreationbench 51.7 模型 qwen3-8-max · 版本 5 platforms (Ubuntu/macOS/Windows/Android/Web) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

recreationbench id exists in data/benchmarks/. Footnote 10: internal long-horizon app-recreation bench across five platforms. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 48.0 | 56.1 | 16.2 | 47.6 | 30.2.

打开官方来源

present-bench 79.6 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

present-bench id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 80.9 | 79.8 | 55.4 | 82.9 | 65.7.

打开官方来源

charxiv-reasoning 88.4 / 93.5 模型 qwen3-8-max · 版本 RQ, Without CI / With CI dual · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

charxiv-reasoning id exists; RQ dual-condition variant (value stores With-CI per footnote 1). Footnote 1 also notes a small number of wrong gold labels corrected with human verification for MathVision and CharXiv (RQ). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 78.5/89.9 | 87.9/93.5 | 84.4/89.9 | 85.1/89.1 | 85.8/85.9.

打开官方来源

omnidocbench 92.1 模型 qwen3-8-max · 版本 1.5 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

omnidocbench id exists in data/benchmarks/; version 1.5 recorded as variant. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 86.5 | 89.5 | 90.0 | 86.7 | 91.4.

打开官方来源

ocr-bench-v2 74.2 / 68.3 模型 qwen3-8-max · 版本 EN / ZH dual · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ocr-bench-v2 not yet in data/benchmarks/. Dual cell = EN / ZH (value stores ZH). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 53.9/55.3 | 65.3/58.1 | 64.6/58.2 | 69.0/57.3 | 70.7/67.1.

打开官方来源

cc-ocr-bench-v2 79.6 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cc-ocr-bench-v2 not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 60.3 | 72.4 | 68.9 | 68.0 | 72.7.

打开官方来源

mtvqa 56.6 模型 qwen3-8-max · 版本 Test · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mtvqa (multilingual text-VQA, Test split) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 48.1 | 41.6 | 54.3 | 52.7 | 51.2.

打开官方来源

madqa 91.8 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: madqa (multi-format document QA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 86.8 | 86.0 | 81.1 | 87.8 | 87.1.

打开官方来源

qwen-visual-office 44.6 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: qwen-visual-office (Qwen in-house, footnote 13) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 34.5 | 32.4 | 39.6 | 29.5 | 32.4.

打开官方来源

realworldqa 88.0 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

realworldqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 76.6 | 85.9 | 83.5 | 83.7 | 86.9.

打开官方来源

erqa 77.8 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

erqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 57.2 | 70.0 | 68.0 | 70.0 | 69.8.

打开官方来源

lingoqa 84.8 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: lingoqa (driving-scenario VQA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 73.8 | 77.4 | 66.8 | 72.6 | 83.4.

打开官方来源

surds 77.8 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: surds (screen/UI understanding) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 62.2 | 79.4 | 64.0 | 63.0 | 77.2.

打开官方来源

simplevqa 75.0 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

simplevqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.3 | 73.4 | 73.1 | 66.6 | 70.3.

打开官方来源

worldvqa 53.2 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

worldvqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 33.9 | 53.5 | 54.0 | 45.1 | 43.9.

打开官方来源

mmstar 85.9 模型 qwen3-8-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

mmstar id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 76.7 | 80.5 | 84.0 | 82.5 | 83.2.

打开官方来源

perception-bench 63.5 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

perception-bench id exists in data/benchmarks/. Footnote 11: competitor cells taken from the benchmark's official release report; our model internally evaluated. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 47.2 | 57.2 | 56.2 | 59.7 | 51.1.

打开官方来源

countqa 82.4 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: countqa (counting QA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 41.3 | 63.1 | 72.8 | 68.6 | 77.0.

打开官方来源

refadv-s 80.2 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: refadv-s (RefAdv spatial split printed as RefAdv-S) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 61.7 | 68.6 | 71.9 | 69.2 | 73.0.

打开官方来源

dense200 87.0 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: dense200 (dense-grounding 200-item set) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 20.8 | 31.1 | 69.7 | 55.3 | 60.7.

打开官方来源

coco 78.7 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: coco (detection/grounding split printed in this table) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 50.7 | 56.4 | 72.4 | 61.2 | 74.2.

打开官方来源

visfactor 60.8 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: visfactor (visual factor knowledge) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 30.1 | 54.5 | 39.8 | 62.8 | 42.8.

打开官方来源

vlmsarebiased 88.3 模型 qwen3-8-max · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

vlmsarebiased id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 43.8 | 61.2 | 74.1 | 59.8 | 36.6.

打开官方来源

video-mme 90.4 模型 qwen3-8-max · 版本 w/ Subtitles · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

video-mme id exists in data/benchmarks/; subtitles-enabled variant. Footnote 12: VideoMME (w/ Sub.) and VideoMME v2 (w/ Sub.) evaluated with subtitles enabled. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 85.4 | -- | 86.7 | 89.5 | 88.0.

打开官方来源

video-mme-v2 68.3 模型 qwen3-8-max · 版本 w/ Subtitles · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: video-mme-v2 not yet in data/benchmarks/. Footnote 12: evaluated with subtitles enabled. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 49.0 | 52.2 | 66.9 | 71.1 | 59.7.

打开官方来源

video-mmmu 88.7 模型 qwen3-8-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

video-mmmu id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.3 | 81.2 | 85.3 | 85.0 | 85.4.

打开官方来源

mmvu 82.4 模型 qwen3-8-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

mmvu id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.4 | 72.0 | 77.9 | 81.2 | 76.6.

打开官方来源

mlvu 90.8 模型 qwen3-8-max · 版本 M-Avg · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mlvu (multimodal long-video understanding, M-Avg) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 53.4 | -- | 84.7 | 87.6 | 87.4.

打开官方来源

tvbench 81.9 模型 qwen3-8-max · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

tvbench id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 61.5 | -- | 73.0 | 83.2 | 78.2.

打开官方来源

lvbench 81.8 模型 qwen3-8-max · 版本 default · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

lvbench id exists in data/benchmarks/; default protocol row (memory-augmented row recorded separately). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.3 | -- | 75.1 | 78.8 | 76.2.

打开官方来源

lvbench 85.6 模型 qwen3-8-max · 版本 w/ Mem. (Qwen-MM-Plugins video memory) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

lvbench id exists; memory-augmented variant. Footnote 14: LVBench (w/ Mem.) and EgoLife (w/ Mem.) use the memory system built on Qwen-MM-Plugins with fine-grained long-horizon video memory. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 84.3 | 90.1 | -- | 84.2 | 74.5.

打开官方来源

egolife 80.3 模型 qwen3-8-max · 版本 w/ Mem. · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: egolife (egocentric long-video life-logging bench) not yet in data/benchmarks/. Footnote 14: memory system built on Qwen-MM-Plugins. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 78.3 | 82.3 | -- | 70.8 | 68.8.

打开官方来源