Qwen3.8-Max
Alibaba / Qwen · 2026-08-03 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Qwen3.8-Max
通义将 Qwen3.8-Max 定位为当前 Qwen 家族最强大的模型,基于 Qwen 3.5 架构将参数规模扩展至 2.4 万亿(激活参数 95B),在编程、办公、科研与长周期任务上全面提升,并成为首个开源权重的 Qwen-Max 级模型(权重于下周发布)。已收录评测覆盖智能体编码与办公、多模态理解与操作、视频理解及长上下文等领域,Terminal-Bench 2.1 86.6、PaperBench 93。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 官方资料未说明
- 参数
- 2.4T(激活 95B)
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Terminal Bench 2.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus4.8 84.6, Fable5 84.6, GPT5.6 Sol 88.8, Qwen3.7-Max 74.5. Harness not printed on this page.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: SWE-bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus4.8 69.2, Fable5 80.0, GPT5.6 64.6, Qwen3.7-Max 60.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: DeepSWE 1.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}deepswe id already introduced by prior batches. Competitor cells: Opus4.8 59.0, Fable5 70.0, GPT5.6 73.0, Qwen3.7-Max 21.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: NL2Repo-Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}nl2repo id already introduced by prior batches. Competitor cells: Opus4.8 69.4, Qwen3.7-Max 47.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: FrontierSWE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}frontierswe id already introduced by prior batches. Competitor cells: Opus4.8 70.0, Fable5 88.8, Qwen3.7-Max 40.7. Snapshot date not printed (GLM-5.2's row was as-of 2026/06/16).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: MLS-Bench-Lite
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mls-bench-lite already introduced by prior batches. Competitor cells: Opus4.8 42.8, Fable5 49.9, GPT5.6 46.2, Qwen3.7-Max 31.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: PaperBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: paperbench already introduced by prior batches. Competitor cells: Opus4.8 80.3, Fable5 88.8, GPT5.6 90.5, Qwen3.7-Max 64.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: AndroidBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: androidbench not yet in data/benchmarks/. Competitor cells: Opus4.8 69.8, Fable5 84.5, GPT5.6 74.0, Qwen3.7-Max 56.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenSWEBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-swe-bench (Qwen in-house SWE bench) not yet in data/benchmarks/. Vendor-owned bench; treat as in-house. Competitor cells: Opus4.8 84.0, Fable5 86.3, GPT5.6 73.5, Qwen3.7-Max 63.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenQoderBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-qoder-bench (Qwen in-house IDE coding bench) not yet in data/benchmarks/. Competitor cells: Opus4.8 62.7, Fable5 63.1, GPT5.6 53.8, Qwen3.7-Max 36.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenReactBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-react-bench (React front-end generation, Elo-scored) not yet in data/benchmarks/. Elo unit. Competitor cells: Opus4.8 1694, Fable5 1770, GPT5.6 1564, Qwen3.7-Max 1538.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: QwenSVGBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-svg-bench (SVG generation, Elo-scored) not yet in data/benchmarks/. Competitor cells: Opus4.8 1648, Fable5 1690, GPT5.6 1758, Qwen3.7-Max 1499.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: CoWorkBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: coworkbench already introduced by prior batches. Competitor cells: Opus4.8 72.3, Fable5 75.9, GPT5.6 71.5, Qwen3.7-Max 64.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: WorkSpaceBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: workspace-bench already introduced by prior batches. Competitor cells: Opus4.8 66.8, Fable5 68.7, GPT5.6 65.6, Qwen3.7-Max 61.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: JobBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: jobbench already introduced by prior batches. Competitor cells: Opus4.8 48.4, Fable5 57.4, GPT5.6 45.4, Qwen3.7-Max 31.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: SkillsBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: skillsbench not yet in data/benchmarks/. Competitor cells: Opus4.8 65.1, Fable5 70.9, GPT5.6 73.5, Qwen3.7-Max 61.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Agents' Last Exam (Pass / Score)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}agents-last-exam id already introduced by prior batches. Dual metric printed in one cell: Pass@1 27.0 AND Score 52.4 - value field holds Score, display holds both. Competitor cells: Opus4.8 27.0/45.1, GPT5.6 30.6/53.6, Qwen3.7-Max 11.8/31.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Automation-Bench (Pass@1)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}automationbench id already introduced by prior batches. Competitor cells: Opus4.8 27.2, Fable5 29.1, GPT5.6 29.7, Qwen3.7-Max 14.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: Toolathlon Verified (Pass@1)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}toolathlon id already introduced by prior batches. Competitor cells: Opus4.8 76.2, Fable5 77.9, GPT5.6 74.9, Qwen3.7-Max 49.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: WideSearch
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: widesearch. Competitor cells: Opus4.8 72.9, Fable5 81.2, Qwen3.7-Max 75.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: HLE w/ tools
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus4.8 57.9, Fable5 64.5, GPT5.6 58.0, Qwen3.7-Max 53.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus4.8 92.0, Fable5 92.6, GPT5.6 94.1, Qwen3.7-Max 92.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: HLE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Opus4.8 45.7, Fable5 53.3, GPT5.6 47.2, Qwen3.7-Max 41.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: IFBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ifbench (instruction-following bench, distinct from ifeval) not yet in data/benchmarks/. Competitor cells: Opus4.8 62.2, Fable5 63.5, GPT5.6 72.7, Qwen3.7-Max 79.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: $OneMillion-Bench (expert score)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}one-million-bench id already introduced by prior batches. Competitor cells: Opus4.8 41.8, Fable5 55.9, GPT5.6 53.8, Qwen3.7-Max 44.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: HealthBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: healthbench (also introduced in kimi-k2-thinking.json batch 1). Competitor cells: Opus4.8 52.4, GPT5.6 55.3, Qwen3.7-Max 54.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: PLawBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: plawbench (legal) not yet in data/benchmarks/. Competitor cells: Opus4.8 69.6, Fable5 70.2, GPT5.6 72.3, Qwen3.7-Max 58.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: PRBench-Legal
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: prbench-legal not yet in data/benchmarks/. Competitor cells: Opus4.8 52.7, Fable5 57.6, GPT5.6 57.6, Qwen3.7-Max 48.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: PRBench-Finance
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: prbench-finance not yet in data/benchmarks/. Competitor cells: Opus4.8 51.9, Fable5 55.8, GPT5.6 55.5, Qwen3.7-Max 46.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: MRCR v2 256K (8-needle)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mrcr-v2 not yet in data/benchmarks/ (distinct from deepseek-v4's MRCR 1M row). Competitor cells: Opus4.8 83.2, GPT5.6 93.8, Qwen3.7-Max 86.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 1 (DOM) · table: DOM table 'Coding Agent / General Agent / General Capabilities' (33 benchmark rows x 5 models, machine-readable) · row: LongBench v2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}longbench id exists (v1); v2 variant. Competitor cells: Opus4.8 69.1, GPT5.6 67.1, Qwen3.7-Max 65.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.6 | 81.2 | 80.5 | 83.0 | 79.0. Footnote 3: Gemini3.1-Pro and GPT5.6-Sol cells taken from official model reports / system cards; all other models internally evaluated.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark variant: mathvision exists in data/benchmarks/ (dual-condition variant recorded here). Dual cell per footnote 1: 无 CI / 有 CI (value stores With-CI). Footnote 2: our model uses a fixed prompt (reason step by step, final answer in \boxed{}); competitors reported as the higher of with/without \boxed{} runs. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 87.1/97.1 | 92.7/98.6 | 87.4/95.7 | 90.8/97.8 | 90.3/--.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 28.4/81.2 | 42.5/90.5 | 55.9/68.3 | 65.5/88.9 | 64.7/70.4. Footnote 1: reported as 无 CI / 有 CI (value stores With-CI).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hle-vl (HLE vision-language variant) not yet in data/benchmarks/. Footnote 6: tools = code interpreter (CI) + search; Gemini3.1-Pro / GPT5.6-Sol tool scores measured end-to-end via their native tool-calling APIs. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): -- | -- | 43.9 | 51.2 | 25.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}zerobench id exists in data/benchmarks/ (dual-condition Pass@5 variant recorded here). Footnote 1: reported as 无 CI / 有 CI (value stores With-CI). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 17.0/34.0 | 20.0/46.0 | 17.0/23.0 | 22.0/35.0 | 19.0/19.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: zerobench-sub (ZeroBench sub-question split printed as a separate row) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 31.1 | 37.1 | 36.5 | 46.7 | 41.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 76.7 | 85.7 | 82.6 | 89.7 | 84.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hipho not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 69.3 | 78.6 | 85.4 | 86.8 | 84.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: phyx not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 54.2 | 71.7 | 79.4 | 79.1 | 80.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: slake (biomedical VQA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.9 | 86.6 | 82.9 | 85.1 | 83.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: medxpertqa-mm not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 71.7 | 80.0 | 80.7 | 81.5 | 71.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: pmc-vqa not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 59.2 | 63.2 | 62.5 | 62.3 | 63.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}osworld id exists in data/benchmarks/; Verified split recorded as variant. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 83.4 | 85.0 | 76.2 | 83.2 | 73.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}osworld id exists; 2.0 dual-condition variant. Footnote 7: 二元 / 部分 format (value stores Partial = share of full+partial reward). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 20.6/54.8 | --/66.1 | 7.8/30.6 | --/62.6 | 2.8/21.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}screenspot-pro id exists in data/benchmarks/. Footnote 8: Opus4.8 and Fable5 cells taken from official system cards; Fable5's figure refers to the corresponding Mythos Preview score; all other models internally evaluated. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 82.3 | 87.3 | 68.1 | 81.3 | 79.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}webarena id exists; Verified split. Footnote 9: scored with the official WebArena scorer inside the OSWorld framework. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.9 | 71.3 | 64.3 | 69.7 | 55.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}androidworld id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.0 | 88.8 | 70.7 | 77.6 | 81.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}mobileworld id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.5 | 85.5 | 58.1 | 76.9 | 51.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}claweval-mm id exists in data/benchmarks/ (dual-metric cell; value stores Average per footnote 4, display stores Pass@3 / Average). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 73.3/73.8 | 81.2/77.5 | 50.5/55.2 | 81.2/78.9 | 57.4/60.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}vision2web id exists in data/benchmarks/. Footnote 5: score = mean over front-end / web / website categories, Claude Code harness with gpt-5.4-2026-03-05 as judge. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 62.4 | 70.5 | -- | 62.1 | 42.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-blender-bench (Qwen in-house Blender task bench, footnote 13) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 62.4 | 69.5 | 23.0 | 68.6 | 41.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: parametric-cad-bench not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 85.1 | 87.5 | 73.5 | 86.2 | 73.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}recreationbench id exists in data/benchmarks/. Footnote 10: internal long-horizon app-recreation bench across five platforms. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 48.0 | 56.1 | 16.2 | 47.6 | 30.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}present-bench id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 80.9 | 79.8 | 55.4 | 82.9 | 65.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}charxiv-reasoning id exists; RQ dual-condition variant (value stores With-CI per footnote 1). Footnote 1 also notes a small number of wrong gold labels corrected with human verification for MathVision and CharXiv (RQ). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 78.5/89.9 | 87.9/93.5 | 84.4/89.9 | 85.1/89.1 | 85.8/85.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}omnidocbench id exists in data/benchmarks/; version 1.5 recorded as variant. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 86.5 | 89.5 | 90.0 | 86.7 | 91.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ocr-bench-v2 not yet in data/benchmarks/. Dual cell = EN / ZH (value stores ZH). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 53.9/55.3 | 65.3/58.1 | 64.6/58.2 | 69.0/57.3 | 70.7/67.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cc-ocr-bench-v2 not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 60.3 | 72.4 | 68.9 | 68.0 | 72.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mtvqa (multilingual text-VQA, Test split) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 48.1 | 41.6 | 54.3 | 52.7 | 51.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: madqa (multi-format document QA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 86.8 | 86.0 | 81.1 | 87.8 | 87.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: qwen-visual-office (Qwen in-house, footnote 13) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 34.5 | 32.4 | 39.6 | 29.5 | 32.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}realworldqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 76.6 | 85.9 | 83.5 | 83.7 | 86.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}erqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 57.2 | 70.0 | 68.0 | 70.0 | 69.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: lingoqa (driving-scenario VQA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 73.8 | 77.4 | 66.8 | 72.6 | 83.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: surds (screen/UI understanding) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 62.2 | 79.4 | 64.0 | 63.0 | 77.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}simplevqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.3 | 73.4 | 73.1 | 66.6 | 70.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}worldvqa id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 33.9 | 53.5 | 54.0 | 45.1 | 43.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}mmstar id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 76.7 | 80.5 | 84.0 | 82.5 | 83.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}perception-bench id exists in data/benchmarks/. Footnote 11: competitor cells taken from the benchmark's official release report; our model internally evaluated. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 47.2 | 57.2 | 56.2 | 59.7 | 51.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: countqa (counting QA) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 41.3 | 63.1 | 72.8 | 68.6 | 77.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: refadv-s (RefAdv spatial split printed as RefAdv-S) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 61.7 | 68.6 | 71.9 | 69.2 | 73.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: dense200 (dense-grounding 200-item set) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 20.8 | 31.1 | 69.7 | 55.3 | 60.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: coco (detection/grounding split printed in this table) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 50.7 | 56.4 | 72.4 | 61.2 | 74.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: visfactor (visual factor knowledge) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 30.1 | 54.5 | 39.8 | 62.8 | 42.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}vlmsarebiased id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 43.8 | 61.2 | 74.1 | 59.8 | 36.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}video-mme id exists in data/benchmarks/; subtitles-enabled variant. Footnote 12: VideoMME (w/ Sub.) and VideoMME v2 (w/ Sub.) evaluated with subtitles enabled. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 85.4 | -- | 86.7 | 89.5 | 88.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: video-mme-v2 not yet in data/benchmarks/. Footnote 12: evaluated with subtitles enabled. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 49.0 | 52.2 | 66.9 | 71.1 | 59.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}video-mmmu id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 75.3 | 81.2 | 85.3 | 85.0 | 85.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}mmvu id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.4 | 72.0 | 77.9 | 81.2 | 76.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mlvu (multimodal long-video understanding, M-Avg) not yet in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 53.4 | -- | 84.7 | 87.6 | 87.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}tvbench id exists in data/benchmarks/. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 61.5 | -- | 73.0 | 83.2 | 78.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}lvbench id exists in data/benchmarks/; default protocol row (memory-augmented row recorded separately). Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 67.3 | -- | 75.1 | 78.8 | 76.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}lvbench id exists; memory-augmented variant. Footnote 14: LVBench (w/ Mem.) and EgoLife (w/ Mem.) use the memory system built on Qwen-MM-Plugins with fine-grained long-horizon video memory. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 84.3 | 90.1 | -- | 84.2 | 74.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: egolife (egocentric long-video life-logging bench) not yet in data/benchmarks/. Footnote 14: memory system built on Qwen-MM-Plugins. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 78.3 | 82.3 | -- | 70.8 | 68.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 Table 2 (DOM) · table: DOM table 'Multimodal Reasoning / Visual Agent & Coding / Document & Office Intelligence / Real-World & Spatial Understanding / Visual Perception & Grounding / Video Intelligence & Agents' (55 data rows x 6 models, machine-readable)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: videodr not yet in data/benchmarks/. Footnote 15: evaluated with search tools enabled. Competitor cells (Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus): 65.6 | 77.1 | -- | 71.3 | 41.0.