← 模型目录

Seed2.1 Pro / Seed2.1 Turbo / Seed2.1 Preview

ByteDance Seed / 豆包 · 2026-06-23 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Seed2.1 Pro

字节 Seed2.1 发布将 Pro 定位为系列最高档,官方称评估优先考察真实工作流表现而非静态榜单分数。已收录评测横跨通用 Agent、编码、多模态与视频理解:亮点 GDPval 87.9、OSWorld 78.8。

输入模态
文本 / 图像 / 视频
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

workspace-bench 53 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Workspace Bench · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 delivers consistent performance on the Workspace Bench and Agent Startup Bench benchmarks

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: workspace-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=53.0; Seed2.1 Turbo=54.7. All page score panels are images; no DOM table.Prose claim is for 'Seed2.1' generally; chart columns are Pro/Turbo. Page describes Workspace Bench as evaluating information retrieval, contextual understanding and result generation for complex workplace documents. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 53, Seed2.1 Turbo 54.7. Same-table competitor cells: Opus 4.7 55.1 / GPT-5.5 58.7 / Gemini 3.1 Pro 32.8.

打开官方来源

present-bench 54.6 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: PresentBench · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: present-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=54.6; Seed2.1 Turbo=48.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 48.3;竞品列:Opus 61.8, GPT-5.5 68.9, Gemini 52.1。

打开官方来源

agent-startup-bench 68.8 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Agent Startup Bench · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: comprehensively assesses response quality through research and interviews with real AI-native startups, combined with expert review

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: agent-startup-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=68.8; Seed2.1 Turbo=54.0. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 68.8, Seed2.1 Turbo 54.0. Same-table competitor cells: Opus 4.7 62.3 / GPT-5.5 68.1 / Gemini 3.1 Pro 45.7.

打开官方来源

agents-last-exam 19.5 模型 seed-2.1-pro · 版本 full pass rate (left panel) · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Agents' Last Exam (pass rate) · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro ranks among the top tier of participating models on the Agents' Last Exam (ALE) benchmark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: agents-last-exam not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=19.5; Seed2.1 Turbo=None. All page score panels are images; no DOM table.Chart caption: 'the left side shows the full pass rate, and the right side shows the average overall score' - two panels kept as separate rows. Turbo cell not visible in pass-rate panel per vision read. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 19.5. Same-table competitor cells: cell prints "19.5 / 41.4" (pass rate / avg overall score); Opus 4.7 18.4/40.5, GPT-5.5 24.0/42.8, Gemini 3.1 Pro 15.8/32.0; Turbo cell "-".

打开官方来源

agents-last-exam 41.4 模型 seed-2.1-pro · 版本 average overall score (right panel) · 指标 overall_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Agents' Last Exam (avg overall score) · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: ranks among the top tier of participating models on the Agents' Last Exam (ALE) benchmark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: agents-last-exam not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=41.4; Seed2.1 Turbo=None. All page score panels are images; no DOM table.Right-panel metric. Page notes the benchmark 'was released only recently', limiting task-specific optimization. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 41.4. Same-table competitor cells: right-hand value of the "19.5 / 41.4" cell.

打开官方来源

one-million-bench 68.8 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OneMillion Bench · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: one-million-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=68.8; Seed2.1 Turbo=66.6. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 66.6;竞品列:Opus 73.0, GPT-5.5 69.6, Gemini 60.2。

打开官方来源

officeqa-pro 70.9 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OfficeQA Pro · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: officeqa-pro not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=70.9; Seed2.1 Turbo=62.8. All page score panels are images; no DOM table.Text-mode panel of Image 1; the separate multimodal 'OfficeQA Pro (MM)' panel of Image 3 is a distinct row. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 62.8;竞品列:Opus 76.5, GPT-5.5 62.9, Gemini 72.5。

打开官方来源

gdpval 87.9 模型 seed-2.1-pro · 版本 未说明 · 指标 quality_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: GDPval · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro achieves the highest score on GDPVal

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: gdpval not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=87.9; Seed2.1 Turbo=82.7. All page score panels are images; no DOM table.Prose 'highest score' cross-checks with vision read (Pro 87.9 vs Opus 4.7 82.7 / GPT-5.5 84.9). Page: 'GDPVal measures the completion quality and economic value of models on real-world work tasks.' 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 87.9, Seed2.1 Turbo 82.7. Same-table competitor cells: Opus 4.7 82.7 / GPT-5.5 84.9 / Gemini 3.1 Pro 67.3.

打开官方来源

finance-agent 60.7 模型 seed-2.1-pro · 版本 v1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Finance Agent v1.1 · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: finance-agent not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=60.7; Seed2.1 Turbo=56.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 56;竞品列:Opus 64.4, GPT-5.5 65.3, Gemini 59.7。

打开官方来源

apex-agents 33.8 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: APEX Agents · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: apex-agents not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=33.8; Seed2.1 Turbo=29.2. All page score panels are images; no DOM table.Row present only in chart; no prose claim. Same benchmark id as kimi-k3 / minimax-m3 apex-agents rows. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 29.2;竞品列:Opus 33.9, GPT-5.5 35.4, Gemini 33.5。

打开官方来源

xdailybench 61 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: xDailyBench · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 achieves steady performance on benchmarks such as xDailyBench and Doubao Multi-Turn Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: xdailybench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=61.0; Seed2.1 Turbo=56.4. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 61, Seed2.1 Turbo 56.4. Same-table competitor cells: Opus 4.7 69.0 / GPT-5.5 73.0 / Gemini 3.1 Pro 35.2.

打开官方来源

doubao-multi-turn-bench 52.5 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Doubao Multi-Turn Bench · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: steady performance on benchmarks such as xDailyBench and Doubao Multi-Turn Bench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: doubao-multi-turn-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=52.5; Seed2.1 Turbo=49.0. All page score panels are images; no DOM table.Vendor-named in-house benchmark. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 52.5, Seed2.1 Turbo 49.0. Same-table competitor cells: Opus 4.7 49.8 / GPT-5.5 62.5 / Gemini 3.1 Pro 52.0.

打开官方来源

mcp-atlas 83.8 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: MCP-Atlas · figure: images/02.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mcp-atlas not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=83.8; Seed2.1 Turbo=80.3. All page score panels are images; no DOM table.Row present only in chart (prose mentions Toolathlon and ClawBench, not MCP-Atlas). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 80.3;竞品列:Opus 79.1, GPT-5.5 81.6, Gemini 78.2。

打开官方来源

toolathlon 50.6 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Toolathlon · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: remains competitive on Toolathlon and ClawBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: toolathlon not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=50.6; Seed2.1 Turbo=49.1. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 50.6, Seed2.1 Turbo 49.1. Same-table competitor cells: Opus 4.7 52.8 / GPT-5.5 55.6 / Gemini 3.1 Pro 48.8.

打开官方来源

clawbench 66.6 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: SeedClawBench · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: remains competitive on Toolathlon and ClawBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: clawbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=66.6; Seed2.1 Turbo=63.8. All page score panels are images; no DOM table.Chart row label reads 'SeedClawBench'; prose calls it ClawBench. Caption defines it as Seed-internal, OpenClaw-style assistance scenarios. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 66.6, Seed2.1 Turbo 63.8. Same-table competitor cells: row label SeedClawBench; Opus 4.7 64.1 / GPT-5.5 66.4 / Gemini 3.1 Pro 57.1.

打开官方来源

claw-eval 51 模型 seed-2.1-pro · 版本 MM (metric Pass^3) · 指标 pass_cubed · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Claw-Eval (MM) · figure: images/03.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: strong overall competitiveness on Visual Agent benchmarks including Claw-Eval (MM)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: claw-eval not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=51.0; Seed2.1 Turbo=46.0. All page score panels are images; no DOM table.Vision read marks Pro 51.0 as the bolded best in row. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 51, Seed2.1 Turbo 46.0. Same-table competitor cells: Pass^3 row; Opus 4.7 44.0 / GPT-5.5 43.0 / Gemini 3.1 Pro 27.0.

打开官方来源

officeqa-pro 72.2 模型 seed-2.1-pro · 版本 MM (metric Avg Score) · 指标 avg_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OfficeQA Pro (MM) · figure: images/03.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: officeqa-pro not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=72.2; Seed2.1 Turbo=71.1. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 71.1;竞品列:Opus 76.5, GPT-5.5 69.5, Gemini 72.5。

打开官方来源

wildclawbench 61.7 模型 seed-2.1-pro · 版本 未说明 · 指标 avg_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: WildClawBench · figure: images/03.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: wildclawbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=61.7; Seed2.1 Turbo=62.8. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 62.8;竞品列:Opus 67.0, GPT-5.5 65.6, Gemini 61.1。

打开官方来源

image2floorplan 48 模型 seed-2.1-pro · 版本 未说明 · 指标 avg_score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Image2FloorPlan (Inhouse) · figure: images/03.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro) · quote_snippet: Image2FloorPlan is an internally developed evaluation set for assessing the task of understanding multiple real-world photos and drawing a floor plan

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: image2floorplan not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=48.0; Seed2.1 Turbo=35.9. All page score panels are images; no DOM table.No prose performance claim; caption is definitional. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 35.9;竞品列:Opus 50.2, GPT-5.5 50.7, Gemini 55.1。

打开官方来源

mobileworld 73.1 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: MobileWorld · figure: images/04.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 achieves the highest score on the MobileWorld benchmark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mobileworld not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=73.1; Seed2.1 Turbo=70.0. All page score panels are images; no DOM table.Prose 'highest score' cross-checks with vision read (Pro 73.1 vs Opus 4.7 57.1 / GPT-5.5 54.7). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 73.1, Seed2.1 Turbo 70.0. Same-table competitor cells: Opus 4.7 57.1 / GPT-5.5 54.7 / Gemini 3.1 Pro 48.4.

打开官方来源

osworld 78.8 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OSWorld · figure: images/04.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: the model remains competitive on OSWorld

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=78.8; Seed2.1 Turbo=76.4. All page score panels are images; no DOM table.Adjacent prose: RL training 'reduc[es] the average number of steps required to complete tasks by 16%' across GUI and non-GUI action spaces - a step-efficiency claim, not a completion-rate variant. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 78.8, Seed2.1 Turbo 76.4. Same-table competitor cells: Opus 4.7 82.8 / GPT-5.5 78.7 / Gemini 3.1 Pro 76.2.

打开官方来源

creativework 42.5 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: CreativeWork · figure: images/04.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 also delivers standout results on the CreativeWork benchmark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: creativework not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=42.5; Seed2.1 Turbo=34.5. All page score panels are images; no DOM table.Caption: in-house Seed benchmark for agents coordinating GUI and MCP tool usage (Notion / Canva / Figma environments). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 42.5, Seed2.1 Turbo 34.5. Same-table competitor cells: Opus 4.7 28.3 / GPT-5.5 30.5 / Gemini 3.1 Pro 27.4.

打开官方来源

gameworld 31.2 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: General agent capabilities significantly enhanced for reliable complex task execution · row: GameWorld · figure: images/04.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: gameworld not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=31.2; Seed2.1 Turbo=25.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 25.9;竞品列:Opus 26.5, GPT-5.5 34.1, Gemini 21.2。

打开官方来源

terminalbench 71 模型 seed-2.1-pro · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: Terminal-Bench 2.1 · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Seed2.1 Pro=71.0; Seed2.1 Turbo=67.6. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 71.7, GPT-5.5 73.8 (bold), Gemini 3.1 Pro 70.7. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 67.6;竞品列:Opus 71.7, GPT-5.5 73.8, Gemini 70.7。

打开官方来源

swebench-pro 57.5 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: SWE-Bench Pro · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-pro not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=57.5; Seed2.1 Turbo=57.0. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 64.3 (bold), GPT-5.5 58.6, Gemini 3.1 Pro 54.2. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 57;竞品列:Opus 64.3, GPT-5.5 58.6, Gemini 54.2。

打开官方来源

cybergym 68.7 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: CyberGym · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: cybergym not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=68.7; Seed2.1 Turbo=67.0. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 73.1, GPT-5.5 81.8 (bold), Gemini 3.1 Pro dash. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 67;竞品列:Opus 73.1, GPT-5.5 81.8, Gemini '-'。

打开官方来源

program-bench 50.3 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: ProgramBench · figure: images/05.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro remains competitive on ProgramBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: program-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=50.3; Seed2.1 Turbo=49.4. All page score panels are images; no DOM table.Prose: demonstrates 'ability to deliver system-level engineering from scratch, including independent software architecture design and full code implementation.' 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 50.3, Seed2.1 Turbo 49.4. Same-table competitor cells: Opus 4.7 52.1 / GPT-5.5 65.9 / Gemini 3.1 Pro 40.7.

打开官方来源

nl2repo 47 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: NL2Repo-Bench · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: nl2repo not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=47.0; Seed2.1 Turbo=43.7. All page score panels are images; no DOM table.Chart label 'NL2Repo-Bench'; same underlying benchmark family as minimax-m3 nl2repo rows. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 43.7;竞品列:Opus 58.2, GPT-5.5 45.1, Gemini 33.4。

打开官方来源

swe-atlas 35.2 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: SWE-Atlas · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swe-atlas not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=35.2; Seed2.1 Turbo=30.6. All page score panels are images; no DOM table.Seed's chart shows one SWE-Atlas row without the QnA / Test-Writing split used by minimax-m3; keep separate until sub-variant is confirmed. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 30.6;竞品列:Opus 38.7, GPT-5.5 44.7, Gemini 23.6。

打开官方来源

deepswe 32.7 模型 seed-2.1-pro · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: DeepSWE · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepswe not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=32.7; Seed2.1 Turbo=23.0. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 54.0, GPT-5.5 70.0 (bold), Gemini 3.1 Pro 10.0. Same benchmark id as kimi-k3 deepswe rows (those were v1.1; variant unspecified here). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 23;竞品列:Opus 54.0, GPT-5.5 70.0, Gemini 10.0。

打开官方来源

code-arena-frontend 1539 模型 seed-2.1-pro · 版本 未说明 · 指标 arena_score · 单位 elo 来源等级 A · benchmark_owner_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: Code Arena: Frontend rank 8 · figure: Leaderboard screenshot image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4xfa4mqq4xsvz.jpg (Arena.ai 'Code Arena: Frontend' leaderboard, Seed-2.1-Pro (Preview) highlighted at rank 8 with 1539, based on 107,962 votes) · quote_snippet: it ranks 8th with a score of 1539, and secures a top-10 position in 5 out of 7 frontend subcategories

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: code-arena-frontend not yet in data/benchmarks.json. Prose gives rank 8 / score 1539 explicitly; leaderboard screenshot (vision) shows 'Seed-2.1-Pro (Preview)' at rank 8 with 1,539 among 15 rows (top: Claude Fable 5 High 1,654; GLM-5.2 Max 1,593; MiniMax-M3 at 15 with 1,505), based on 107,962 votes. Score originates from the third-party Arena.ai leaderboard, hence benchmark_owner_reported rather than vendor-run; blog separately notes the 'Seed2.1 Preview' naming. Also prose: 'top-10 position in 5 out of 7 frontend subcategories.'

打开官方来源

charxiv-reasoning 85.4 模型 seed-2.1-pro · 版本 RQ (w. Tool) · 指标 reasoning_accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: CharXiv-RQ (w. Tool) · figure: images/08.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro achieves the highest scores on multiple benchmarks including CharXiv-RQ and MeasureBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: charxiv-reasoning not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=85.4 (86.4); Seed2.1 Turbo=82.5 (83.6). All page score panels are images; no DOM table.Vision read shows a parenthetical second value per Seed cell whose meaning (e.g. different prompt/pass setting) is not explained in page text; needs human confirmation. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 85.4, Seed2.1 Turbo 82.5. Same-table competitor cells: cell prints "85.4 (86.4)" / "82.5 (83.6)"; parenthetical value recorded in notes, primary value taken as the unparenthesized number; Opus 4.7 82.1 / GPT-5.5 83.2 / Gemini 3.1 Pro 83.5.

打开官方来源

measurebench 62.9 模型 seed-2.1-pro · 版本 avg. real & synthetic · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MeasureBench · figure: images/08.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: achieves the highest scores on multiple benchmarks including CharXiv-RQ and MeasureBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: measurebench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=62.9; Seed2.1 Turbo=58.9. All page score panels are images; no DOM table.Vision cross-check: Pro 62.9 is row best (Opus 4.7 29.7, GPT-5.5 High 49.9, Gemini 3.1 Pro 44.4). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 62.9, Seed2.1 Turbo 58.9. Same-table competitor cells: row "MeasureBench (avg. real & synthetic)"; Opus 4.7 29.7 / GPT-5.5 49.9 / Gemini 3.1 Pro 44.4.

打开官方来源

mathvision 92.6 模型 seed-2.1-pro · 版本 w. Tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MathVision (w. Tool) · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mathvision not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=92.6 (94.5); Seed2.1 Turbo=90.1 (92.7). All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 90.1;竞品列:括号内 w. Tool 第二值 Pro (94.5) / Turbo (92.7);Opus 83.1, GPT-5.5 92.2, Gemini 89.2。

打开官方来源

mmmu 81.6 模型 seed-2.1-pro · 版本 Pro (w. Tool) · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MMMU-Pro (w. Tool) · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Seed2.1 Pro=81.6 (82.7); Seed2.1 Turbo=80.1 (82.2). All page score panels are images; no DOM table.Maps to existing benchmark mmmu, variant Pro with tool use. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 80.1;竞品列:括号第二值 Pro (82.7) / Turbo (82.2);Opus 74.0, GPT-5.5 81.2, Gemini 80.5。

打开官方来源

zerobench 18 模型 seed-2.1-pro · 版本 w. Tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: ZEROBench (w. Tool) · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: zerobench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=18.0 (22.0); Seed2.1 Turbo=11.0 (20.0). All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 11;竞品列:括号第二值 Pro (22.0) / Turbo (20.0);Opus 8.0, GPT-5.5 13.0, Gemini 12.0。

打开官方来源

realworldqa 86.7 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: RealWorldQA · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: realworldqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=86.7; Seed2.1 Turbo=86.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 86.3;竞品列:Opus 75.6, GPT-5.5 82.2, Gemini 85.4。

打开官方来源

babyvision 73.7 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: BabyVision · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: babyvision not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=73.7; Seed2.1 Turbo=62.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 62.9;竞品列:Opus 22.2, GPT-5.5 55.9, Gemini 54.4。

打开官方来源

erqa 72 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: ERQA · figure: images/09.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 also delivers top results on the ERQA benchmark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: erqa not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=72.0; Seed2.1 Turbo=71.3. All page score panels are images; no DOM table.Vision cross-check: Pro 72.0 row best (Opus 4.7 52.5, GPT-5.5 High 64.5, Gemini 3.1 Pro 70.8). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 72, Seed2.1 Turbo 71.3. Same-table competitor cells: Opus 4.7 52.5 / GPT-5.5 64.5 / Gemini 3.1 Pro 70.8.

打开官方来源

embspatial-bench 83.4 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: EmbSpatial-Bench · figure: images/09.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: embspatial-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=83.4; Seed2.1 Turbo=82.5. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 82.5;竞品列:Opus 77.2, GPT-5.5 81.9, Gemini 84.2。

打开官方来源

mmlongbench 78.3 模型 seed-2.1-pro · 版本 128K · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MMLongBench-128K · figure: images/09.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: standout performance on the MMLongBench-128K long-context benchmark

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmlongbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=78.3; Seed2.1 Turbo=76.9. All page score panels are images; no DOM table.Opus 4.7 and GPT-5.5 High cells are dashes per vision read; Gemini 3.1 Pro 70.7. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 78.3, Seed2.1 Turbo 76.9. Same-table competitor cells: Opus 4.7 and GPT-5.5 cells print "-"; Gemini 3.1 Pro 70.7.

打开官方来源

tomato 79.5 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: TOMATO · figure: images/10.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro scores at industry-leading levels on TVBench and TOMATO

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tomato not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=79.5; Seed2.1 Turbo=56.8. All page score panels are images; no DOM table.Competitor columns here are Gemini 3.1 Pro 60.4 and Gemini 3.5 Flash 71.9. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 79.5, Seed2.1 Turbo 56.8. Same-table competitor cells: competitor columns Gemini 3.1 Pro 60.4 / Gemini 3.5 Flash 71.9.

打开官方来源

tvbench 80.5 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: TVBench · figure: images/10.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: scores at industry-leading levels on TVBench and TOMATO

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tvbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=80.5; Seed2.1 Turbo=77.2. All page score panels are images; no DOM table.Competitor columns: Gemini 3.1 Pro 71, Gemini 3.5 Flash 76.4. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 80.5, Seed2.1 Turbo 77.2. Same-table competitor cells: competitor columns Gemini 3.1 Pro 71 / Gemini 3.5 Flash 76.4.

打开官方来源

video-mme 89.2 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: VideoMME · figure: images/12.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: strong performance on benchmarks including Video MME and LVBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: video-mme not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=89.2; Seed2.1 Turbo=89.0. All page score panels are images; no DOM table.Same benchmark id as minimax-m3 / seed-1-8 video-mme rows; variant/subtitle condition not stated on this page. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 89.2, Seed2.1 Turbo 89. Same-table competitor cells: competitor columns Gemini 3.1 Pro 86.7 / Gemini 3.5 Flash 87.2.

打开官方来源

lvbench 78 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: LVBench · figure: images/12.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: strong performance on benchmarks including Video MME and LVBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: lvbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=78.0; Seed2.1 Turbo=76.8. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 78, Seed2.1 Turbo 76.8. Same-table competitor cells: competitor columns Gemini 3.1 Pro 75.1 / Gemini 3.5 Flash 76.3.

打开官方来源

ovobench 80.7 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: OVOBench · figure: images/11.png (官方评测表,Streaming 组:OVOBench / OVBench)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ovobench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=80.7; Seed2.1 Turbo=79.2. All page score panels are images; no DOM table.Streaming-group row; distinct benchmark from OVBench despite similar name. Prose mentions only OVBench. 视觉转写自归档图 images/11.png(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 79.2;竞品列 Gemini 3.1 Pro 64.1, Gemini 3.5 Flash 64.5。

打开官方来源

ovbench 70 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: OVBench · figure: images/12.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: standout results on OVBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ovbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=70.0; Seed2.1 Turbo=69.7. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 70, Seed2.1 Turbo 69.7. Same-table competitor cells: competitor columns Gemini 3.1 Pro 58.8 / Gemini 3.5 Flash 56.5.

打开官方来源

supergpqa 70.8 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: SuperGPQA · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: supergpqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=70.8; Seed2.1 Turbo=67.4. All page score panels are images; no DOM table.Same benchmark id as seed-1-8 supergpqa rows. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 67.4;竞品列:Opus 68.5, GPT-5.5 72.7, Gemini 76.6。

打开官方来源

kina 48.3 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: KINA · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: kina not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=48.3; Seed2.1 Turbo=46.6. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 46.6;竞品列:Opus 46.7, GPT-5.5 52.6, Gemini 53.2。

打开官方来源

hlehle 42.9 模型 seed-2.1-pro · 版本 Verified · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: HLE-Verified · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Vision-read value (unconfirmed): Seed2.1 Pro=42.9; Seed2.1 Turbo=42.4. All page score panels are images; no DOM table.Maps to existing benchmark hlehle (Humanity's Last Exam); 'Verified' qualifier is as printed on the chart. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 42.4;竞品列:Opus 46.9, GPT-5.5 50.4, Gemini 48.2。

打开官方来源

scicode 59.8 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: SciCode · figure: images/11.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: It performs well on benchmarks such as SciCode and FrontierScience-Olympiad

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: scicode not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=59.8; Seed2.1 Turbo=57.8. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 59.8, Seed2.1 Turbo 57.8. Same-table competitor cells: Opus 4.7 56.4 / GPT-5.5 58.4 / Gemini 3.1 Pro 62.3.

打开官方来源

frontier-science-olympiad 75 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: FrontierScience-Olympiad · figure: images/11.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: performs well on benchmarks such as SciCode and FrontierScience-Olympiad

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: frontier-science-olympiad not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=75.0; Seed2.1 Turbo=76.0. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 75, Seed2.1 Turbo 76.0. Same-table competitor cells: Opus 4.7 69.0 / GPT-5.5 69.0 / Gemini 3.1 Pro 79.0.

打开官方来源

msqa 50.2 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MSQA · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: msqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=50.2; Seed2.1 Turbo=42.0. All page score panels are images; no DOM table.Caption: 'MSQA is an in-house multilingual benchmark designed to assess culture-specific knowledge across 11 major languages.' No explicit prose score claim. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 42;竞品列:Opus 42.8, GPT-5.5 57.4, Gemini 69.7。

打开官方来源

posttrain-bench 16.5 模型 seed-2.1-pro · 版本 未说明 · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: PostTrainBench · figure: images/13.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: posttrain-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=16.5; Seed2.1 Turbo=18.3. All page score panels are images; no DOM table.Numerically incompatible with minimax-m3 PostTrainBench rows (M3 37.1/0.37) - different harness or scale; do not merge without protocol confirmation. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 18.3;竞品列:Opus 27.4, GPT-5.5 25.0, Gemini '-'。

打开官方来源

frontier-science-research 28.3 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: FrontierScience-Research · figure: images/13.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: It maintains competitive on frontier research benchmarks such as FrontierScience-Research

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: frontier-science-research not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=28.3; Seed2.1 Turbo=33.3. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 28.3, Seed2.1 Turbo 33.3. Same-table competitor cells: Opus 4.7 20.0 / GPT-5.5 33.9 / Gemini 3.1 Pro 16.7.

打开官方来源

frontiercs 46.3 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: FrontierCS · figure: images/13.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: frontiercs not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=46.3; Seed2.1 Turbo=50.8. All page score panels are images; no DOM table.Turbo cell bolded as row best per vision read. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 50.8;竞品列:Opus '-', GPT-5.5 58.6, Gemini 64.4。

打开官方来源

horizonmath 2 模型 seed-2.1-pro · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: HorizonMath · figure: images/13.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: horizonmath not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=2.0; Seed2.1 Turbo=2.0. All page score panels are images; no DOM table.Very low absolute values (Pro/Turbo 2.0; GPT-5.5 7.1 best per vision read); scale/normalization unexplained on page. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 2;竞品列:Opus 4.0, GPT-5.5 7.1, Gemini 4.0。

打开官方来源

Seed2.1 Turbo

Seed2.1 Turbo 为系列提速档,与 Pro 在官方对照表同表评测。账本评测行以 Pro 列为准记录,Turbo 同表对照值见归档(如 Workspace Bench 54.7、GDPval 82.7)。

输入模态
文本 / 图像 / 视频
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。

Seed2.1 Preview

Seed2.1 Preview 为参与 Code Arena: Frontend 第三方榜单的预览条目(榜单排名 8/1539)。本发布未为其单列正式评测行。

输入模态
官方资料未说明
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。