Seed2.1 Pro / Seed2.1 Turbo / Seed2.1 Preview
ByteDance Seed / 豆包 · 2026-06-23 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Seed2.1 Pro
字节 Seed2.1 发布将 Pro 定位为系列最高档,官方称评估优先考察真实工作流表现而非静态榜单分数。已收录评测横跨通用 Agent、编码、多模态与视频理解:亮点 GDPval 87.9、OSWorld 78.8。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Workspace Bench · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 delivers consistent performance on the Workspace Bench and Agent Startup Bench benchmarks
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: workspace-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=53.0; Seed2.1 Turbo=54.7. All page score panels are images; no DOM table.Prose claim is for 'Seed2.1' generally; chart columns are Pro/Turbo. Page describes Workspace Bench as evaluating information retrieval, contextual understanding and result generation for complex workplace documents. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 53, Seed2.1 Turbo 54.7. Same-table competitor cells: Opus 4.7 55.1 / GPT-5.5 58.7 / Gemini 3.1 Pro 32.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: PresentBench · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: present-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=54.6; Seed2.1 Turbo=48.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 48.3;竞品列:Opus 61.8, GPT-5.5 68.9, Gemini 52.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Agent Startup Bench · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: comprehensively assesses response quality through research and interviews with real AI-native startups, combined with expert review
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: agent-startup-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=68.8; Seed2.1 Turbo=54.0. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 68.8, Seed2.1 Turbo 54.0. Same-table competitor cells: Opus 4.7 62.3 / GPT-5.5 68.1 / Gemini 3.1 Pro 45.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Agents' Last Exam (pass rate) · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro ranks among the top tier of participating models on the Agents' Last Exam (ALE) benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: agents-last-exam not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=19.5; Seed2.1 Turbo=None. All page score panels are images; no DOM table.Chart caption: 'the left side shows the full pass rate, and the right side shows the average overall score' - two panels kept as separate rows. Turbo cell not visible in pass-rate panel per vision read. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 19.5. Same-table competitor cells: cell prints "19.5 / 41.4" (pass rate / avg overall score); Opus 4.7 18.4/40.5, GPT-5.5 24.0/42.8, Gemini 3.1 Pro 15.8/32.0; Turbo cell "-".
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Agents' Last Exam (avg overall score) · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: ranks among the top tier of participating models on the Agents' Last Exam (ALE) benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: agents-last-exam not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=41.4; Seed2.1 Turbo=None. All page score panels are images; no DOM table.Right-panel metric. Page notes the benchmark 'was released only recently', limiting task-specific optimization. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 41.4. Same-table competitor cells: right-hand value of the "19.5 / 41.4" cell.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OneMillion Bench · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: one-million-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=68.8; Seed2.1 Turbo=66.6. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 66.6;竞品列:Opus 73.0, GPT-5.5 69.6, Gemini 60.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OfficeQA Pro · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: officeqa-pro not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=70.9; Seed2.1 Turbo=62.8. All page score panels are images; no DOM table.Text-mode panel of Image 1; the separate multimodal 'OfficeQA Pro (MM)' panel of Image 3 is a distinct row. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 62.8;竞品列:Opus 76.5, GPT-5.5 62.9, Gemini 72.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: GDPval · figure: images/01.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro achieves the highest score on GDPVal
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gdpval not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=87.9; Seed2.1 Turbo=82.7. All page score panels are images; no DOM table.Prose 'highest score' cross-checks with vision read (Pro 87.9 vs Opus 4.7 82.7 / GPT-5.5 84.9). Page: 'GDPVal measures the completion quality and economic value of models on real-world work tasks.' 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 87.9, Seed2.1 Turbo 82.7. Same-table competitor cells: Opus 4.7 82.7 / GPT-5.5 84.9 / Gemini 3.1 Pro 67.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Finance Agent v1.1 · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: finance-agent not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=60.7; Seed2.1 Turbo=56.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 56;竞品列:Opus 64.4, GPT-5.5 65.3, Gemini 59.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: APEX Agents · figure: images/01.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: apex-agents not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=33.8; Seed2.1 Turbo=29.2. All page score panels are images; no DOM table.Row present only in chart; no prose claim. Same benchmark id as kimi-k3 / minimax-m3 apex-agents rows. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 29.2;竞品列:Opus 33.9, GPT-5.5 35.4, Gemini 33.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: xDailyBench · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 achieves steady performance on benchmarks such as xDailyBench and Doubao Multi-Turn Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: xdailybench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=61.0; Seed2.1 Turbo=56.4. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 61, Seed2.1 Turbo 56.4. Same-table competitor cells: Opus 4.7 69.0 / GPT-5.5 73.0 / Gemini 3.1 Pro 35.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Doubao Multi-Turn Bench · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: steady performance on benchmarks such as xDailyBench and Doubao Multi-Turn Bench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: doubao-multi-turn-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=52.5; Seed2.1 Turbo=49.0. All page score panels are images; no DOM table.Vendor-named in-house benchmark. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 52.5, Seed2.1 Turbo 49.0. Same-table competitor cells: Opus 4.7 49.8 / GPT-5.5 62.5 / Gemini 3.1 Pro 52.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: MCP-Atlas · figure: images/02.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mcp-atlas not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=83.8; Seed2.1 Turbo=80.3. All page score panels are images; no DOM table.Row present only in chart (prose mentions Toolathlon and ClawBench, not MCP-Atlas). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 80.3;竞品列:Opus 79.1, GPT-5.5 81.6, Gemini 78.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Toolathlon · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: remains competitive on Toolathlon and ClawBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: toolathlon not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=50.6; Seed2.1 Turbo=49.1. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 50.6, Seed2.1 Turbo 49.1. Same-table competitor cells: Opus 4.7 52.8 / GPT-5.5 55.6 / Gemini 3.1 Pro 48.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: SeedClawBench · figure: images/02.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: remains competitive on Toolathlon and ClawBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: clawbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=66.6; Seed2.1 Turbo=63.8. All page score panels are images; no DOM table.Chart row label reads 'SeedClawBench'; prose calls it ClawBench. Caption defines it as Seed-internal, OpenClaw-style assistance scenarios. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 66.6, Seed2.1 Turbo 63.8. Same-table competitor cells: row label SeedClawBench; Opus 4.7 64.1 / GPT-5.5 66.4 / Gemini 3.1 Pro 57.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Claw-Eval (MM) · figure: images/03.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: strong overall competitiveness on Visual Agent benchmarks including Claw-Eval (MM)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: claw-eval not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=51.0; Seed2.1 Turbo=46.0. All page score panels are images; no DOM table.Vision read marks Pro 51.0 as the bolded best in row. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 51, Seed2.1 Turbo 46.0. Same-table competitor cells: Pass^3 row; Opus 4.7 44.0 / GPT-5.5 43.0 / Gemini 3.1 Pro 27.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OfficeQA Pro (MM) · figure: images/03.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: officeqa-pro not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=72.2; Seed2.1 Turbo=71.1. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 71.1;竞品列:Opus 76.5, GPT-5.5 69.5, Gemini 72.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: WildClawBench · figure: images/03.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: wildclawbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=61.7; Seed2.1 Turbo=62.8. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 62.8;竞品列:Opus 67.0, GPT-5.5 65.6, Gemini 61.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: Image2FloorPlan (Inhouse) · figure: images/03.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro) · quote_snippet: Image2FloorPlan is an internally developed evaluation set for assessing the task of understanding multiple real-world photos and drawing a floor plan
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: image2floorplan not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=48.0; Seed2.1 Turbo=35.9. All page score panels are images; no DOM table.No prose performance claim; caption is definitional. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 35.9;竞品列:Opus 50.2, GPT-5.5 50.7, Gemini 55.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: MobileWorld · figure: images/04.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 achieves the highest score on the MobileWorld benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mobileworld not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=73.1; Seed2.1 Turbo=70.0. All page score panels are images; no DOM table.Prose 'highest score' cross-checks with vision read (Pro 73.1 vs Opus 4.7 57.1 / GPT-5.5 54.7). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 73.1, Seed2.1 Turbo 70.0. Same-table competitor cells: Opus 4.7 57.1 / GPT-5.5 54.7 / Gemini 3.1 Pro 48.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: OSWorld · figure: images/04.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: the model remains competitive on OSWorld
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=78.8; Seed2.1 Turbo=76.4. All page score panels are images; no DOM table.Adjacent prose: RL training 'reduc[es] the average number of steps required to complete tasks by 16%' across GUI and non-GUI action spaces - a step-efficiency claim, not a completion-rate variant. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 78.8, Seed2.1 Turbo 76.4. Same-table competitor cells: Opus 4.7 82.8 / GPT-5.5 78.7 / Gemini 3.1 Pro 76.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: CreativeWork · figure: images/04.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 also delivers standout results on the CreativeWork benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: creativework not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=42.5; Seed2.1 Turbo=34.5. All page score panels are images; no DOM table.Caption: in-house Seed benchmark for agents coordinating GUI and MCP tool usage (Notion / Canva / Figma environments). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 42.5, Seed2.1 Turbo 34.5. Same-table competitor cells: Opus 4.7 28.3 / GPT-5.5 30.5 / Gemini 3.1 Pro 27.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: General agent capabilities significantly enhanced for reliable complex task execution · row: GameWorld · figure: images/04.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gameworld not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=31.2; Seed2.1 Turbo=25.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 25.9;竞品列:Opus 26.5, GPT-5.5 34.1, Gemini 21.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: Terminal-Bench 2.1 · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Seed2.1 Pro=71.0; Seed2.1 Turbo=67.6. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 71.7, GPT-5.5 73.8 (bold), Gemini 3.1 Pro 70.7. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 67.6;竞品列:Opus 71.7, GPT-5.5 73.8, Gemini 70.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: SWE-Bench Pro · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-pro not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=57.5; Seed2.1 Turbo=57.0. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 64.3 (bold), GPT-5.5 58.6, Gemini 3.1 Pro 54.2. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 57;竞品列:Opus 64.3, GPT-5.5 58.6, Gemini 54.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: CyberGym · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cybergym not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=68.7; Seed2.1 Turbo=67.0. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 73.1, GPT-5.5 81.8 (bold), Gemini 3.1 Pro dash. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 67;竞品列:Opus 73.1, GPT-5.5 81.8, Gemini '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: ProgramBench · figure: images/05.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro remains competitive on ProgramBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: program-bench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=50.3; Seed2.1 Turbo=49.4. All page score panels are images; no DOM table.Prose: demonstrates 'ability to deliver system-level engineering from scratch, including independent software architecture design and full code implementation.' 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 50.3, Seed2.1 Turbo 49.4. Same-table competitor cells: Opus 4.7 52.1 / GPT-5.5 65.9 / Gemini 3.1 Pro 40.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: NL2Repo-Bench · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: nl2repo not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=47.0; Seed2.1 Turbo=43.7. All page score panels are images; no DOM table.Chart label 'NL2Repo-Bench'; same underlying benchmark family as minimax-m3 nl2repo rows. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 43.7;竞品列:Opus 58.2, GPT-5.5 45.1, Gemini 33.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: SWE-Atlas · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swe-atlas not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=35.2; Seed2.1 Turbo=30.6. All page score panels are images; no DOM table.Seed's chart shows one SWE-Atlas row without the QnA / Test-Writing split used by minimax-m3; keep separate until sub-variant is confirmed. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 30.6;竞品列:Opus 38.7, GPT-5.5 44.7, Gemini 23.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: DeepSWE · figure: images/05.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepswe not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=32.7; Seed2.1 Turbo=23.0. All page score panels are images; no DOM table.Competitor cells (vision): Opus 4.7 54.0, GPT-5.5 70.0 (bold), Gemini 3.1 Pro 10.0. Same benchmark id as kimi-k3 deepswe rows (those were v1.1; variant unspecified here). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 23;竞品列:Opus 54.0, GPT-5.5 70.0, Gemini 10.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: End-to-end coding capabilities significantly enhanced for reliable delivery in enterprise production scenarios · row: Code Arena: Frontend rank 8 · figure: Leaderboard screenshot image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4xfa4mqq4xsvz.jpg (Arena.ai 'Code Arena: Frontend' leaderboard, Seed-2.1-Pro (Preview) highlighted at rank 8 with 1539, based on 107,962 votes) · quote_snippet: it ranks 8th with a score of 1539, and secures a top-10 position in 5 out of 7 frontend subcategories
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: code-arena-frontend not yet in data/benchmarks.json. Prose gives rank 8 / score 1539 explicitly; leaderboard screenshot (vision) shows 'Seed-2.1-Pro (Preview)' at rank 8 with 1,539 among 15 rows (top: Claude Fable 5 High 1,654; GLM-5.2 Max 1,593; MiniMax-M3 at 15 with 1,505), based on 107,962 votes. Score originates from the third-party Arena.ai leaderboard, hence benchmark_owner_reported rather than vendor-run; blog separately notes the 'Seed2.1 Preview' naming. Also prose: 'top-10 position in 5 out of 7 frontend subcategories.'
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: CharXiv-RQ (w. Tool) · figure: images/08.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro achieves the highest scores on multiple benchmarks including CharXiv-RQ and MeasureBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: charxiv-reasoning not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=85.4 (86.4); Seed2.1 Turbo=82.5 (83.6). All page score panels are images; no DOM table.Vision read shows a parenthetical second value per Seed cell whose meaning (e.g. different prompt/pass setting) is not explained in page text; needs human confirmation. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 85.4, Seed2.1 Turbo 82.5. Same-table competitor cells: cell prints "85.4 (86.4)" / "82.5 (83.6)"; parenthetical value recorded in notes, primary value taken as the unparenthesized number; Opus 4.7 82.1 / GPT-5.5 83.2 / Gemini 3.1 Pro 83.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MeasureBench · figure: images/08.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: achieves the highest scores on multiple benchmarks including CharXiv-RQ and MeasureBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: measurebench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=62.9; Seed2.1 Turbo=58.9. All page score panels are images; no DOM table.Vision cross-check: Pro 62.9 is row best (Opus 4.7 29.7, GPT-5.5 High 49.9, Gemini 3.1 Pro 44.4). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 62.9, Seed2.1 Turbo 58.9. Same-table competitor cells: row "MeasureBench (avg. real & synthetic)"; Opus 4.7 29.7 / GPT-5.5 49.9 / Gemini 3.1 Pro 44.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MathVision (w. Tool) · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mathvision not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=92.6 (94.5); Seed2.1 Turbo=90.1 (92.7). All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 90.1;竞品列:括号内 w. Tool 第二值 Pro (94.5) / Turbo (92.7);Opus 83.1, GPT-5.5 92.2, Gemini 89.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MMMU-Pro (w. Tool) · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Seed2.1 Pro=81.6 (82.7); Seed2.1 Turbo=80.1 (82.2). All page score panels are images; no DOM table.Maps to existing benchmark mmmu, variant Pro with tool use. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 80.1;竞品列:括号第二值 Pro (82.7) / Turbo (82.2);Opus 74.0, GPT-5.5 81.2, Gemini 80.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: ZEROBench (w. Tool) · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: zerobench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=18.0 (22.0); Seed2.1 Turbo=11.0 (20.0). All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 11;竞品列:括号第二值 Pro (22.0) / Turbo (20.0);Opus 8.0, GPT-5.5 13.0, Gemini 12.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: RealWorldQA · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: realworldqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=86.7; Seed2.1 Turbo=86.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 86.3;竞品列:Opus 75.6, GPT-5.5 82.2, Gemini 85.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: BabyVision · figure: images/08.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: babyvision not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=73.7; Seed2.1 Turbo=62.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 62.9;竞品列:Opus 22.2, GPT-5.5 55.9, Gemini 54.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: ERQA · figure: images/09.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 also delivers top results on the ERQA benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: erqa not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=72.0; Seed2.1 Turbo=71.3. All page score panels are images; no DOM table.Vision cross-check: Pro 72.0 row best (Opus 4.7 52.5, GPT-5.5 High 64.5, Gemini 3.1 Pro 70.8). 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 72, Seed2.1 Turbo 71.3. Same-table competitor cells: Opus 4.7 52.5 / GPT-5.5 64.5 / Gemini 3.1 Pro 70.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: EmbSpatial-Bench · figure: images/09.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: embspatial-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=83.4; Seed2.1 Turbo=82.5. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 82.5;竞品列:Opus 77.2, GPT-5.5 81.9, Gemini 84.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MMLongBench-128K · figure: images/09.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: standout performance on the MMLongBench-128K long-context benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmlongbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=78.3; Seed2.1 Turbo=76.9. All page score panels are images; no DOM table.Opus 4.7 and GPT-5.5 High cells are dashes per vision read; Gemini 3.1 Pro 70.7. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 78.3, Seed2.1 Turbo 76.9. Same-table competitor cells: Opus 4.7 and GPT-5.5 cells print "-"; Gemini 3.1 Pro 70.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: TOMATO · figure: images/10.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: Seed2.1 Pro scores at industry-leading levels on TVBench and TOMATO
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tomato not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=79.5; Seed2.1 Turbo=56.8. All page score panels are images; no DOM table.Competitor columns here are Gemini 3.1 Pro 60.4 and Gemini 3.5 Flash 71.9. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 79.5, Seed2.1 Turbo 56.8. Same-table competitor cells: competitor columns Gemini 3.1 Pro 60.4 / Gemini 3.5 Flash 71.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: TVBench · figure: images/10.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: scores at industry-leading levels on TVBench and TOMATO
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tvbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=80.5; Seed2.1 Turbo=77.2. All page score panels are images; no DOM table.Competitor columns: Gemini 3.1 Pro 71, Gemini 3.5 Flash 76.4. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 80.5, Seed2.1 Turbo 77.2. Same-table competitor cells: competitor columns Gemini 3.1 Pro 71 / Gemini 3.5 Flash 76.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: VideoMME · figure: images/12.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: strong performance on benchmarks including Video MME and LVBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: video-mme not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=89.2; Seed2.1 Turbo=89.0. All page score panels are images; no DOM table.Same benchmark id as minimax-m3 / seed-1-8 video-mme rows; variant/subtitle condition not stated on this page. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 89.2, Seed2.1 Turbo 89. Same-table competitor cells: competitor columns Gemini 3.1 Pro 86.7 / Gemini 3.5 Flash 87.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: LVBench · figure: images/12.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: strong performance on benchmarks including Video MME and LVBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: lvbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=78.0; Seed2.1 Turbo=76.8. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 78, Seed2.1 Turbo 76.8. Same-table competitor cells: competitor columns Gemini 3.1 Pro 75.1 / Gemini 3.5 Flash 76.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: OVOBench · figure: images/11.png (官方评测表,Streaming 组:OVOBench / OVBench)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ovobench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=80.7; Seed2.1 Turbo=79.2. All page score panels are images; no DOM table.Streaming-group row; distinct benchmark from OVBench despite similar name. Prose mentions only OVBench. 视觉转写自归档图 images/11.png(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 79.2;竞品列 Gemini 3.1 Pro 64.1, Gemini 3.5 Flash 64.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: OVBench · figure: images/12.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: standout results on OVBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ovbench not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=70.0; Seed2.1 Turbo=69.7. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 70, Seed2.1 Turbo 69.7. Same-table competitor cells: competitor columns Gemini 3.1 Pro 58.8 / Gemini 3.5 Flash 56.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: SuperGPQA · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: supergpqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=70.8; Seed2.1 Turbo=67.4. All page score panels are images; no DOM table.Same benchmark id as seed-1-8 supergpqa rows. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 67.4;竞品列:Opus 68.5, GPT-5.5 72.7, Gemini 76.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: KINA · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: kina not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=48.3; Seed2.1 Turbo=46.6. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 46.6;竞品列:Opus 46.7, GPT-5.5 52.6, Gemini 53.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: HLE-Verified · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Seed2.1 Pro=42.9; Seed2.1 Turbo=42.4. All page score panels are images; no DOM table.Maps to existing benchmark hlehle (Humanity's Last Exam); 'Verified' qualifier is as printed on the chart. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 42.4;竞品列:Opus 46.9, GPT-5.5 50.4, Gemini 48.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: SciCode · figure: images/11.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: It performs well on benchmarks such as SciCode and FrontierScience-Olympiad
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: scicode not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=59.8; Seed2.1 Turbo=57.8. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 59.8, Seed2.1 Turbo 57.8. Same-table competitor cells: Opus 4.7 56.4 / GPT-5.5 58.4 / Gemini 3.1 Pro 62.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: FrontierScience-Olympiad · figure: images/11.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: performs well on benchmarks such as SciCode and FrontierScience-Olympiad
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: frontier-science-olympiad not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=75.0; Seed2.1 Turbo=76.0. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 75, Seed2.1 Turbo 76.0. Same-table competitor cells: Opus 4.7 69.0 / GPT-5.5 69.0 / Gemini 3.1 Pro 79.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: MSQA · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: msqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=50.2; Seed2.1 Turbo=42.0. All page score panels are images; no DOM table.Caption: 'MSQA is an in-house multilingual benchmark designed to assess culture-specific knowledge across 11 major languages.' No explicit prose score claim. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 42;竞品列:Opus 42.8, GPT-5.5 57.4, Gemini 69.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: HLE-textonly (with Search) · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): Seed2.1 Pro=55.7; Seed2.1 Turbo=54.6. All page score panels are images; no DOM table.Search-augmented condition; Pro column bolded as row best. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 54.6;竞品列:Opus 54.7, GPT-5.5 52.2, Gemini 51.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: BrowseComp (with Search) · figure: images/12.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=86.2; Seed2.1 Turbo=84.9. All page score panels are images; no DOM table.Search-augmented condition; same benchmark id as kimi-k3 / minimax-m3 browsecomp rows. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 84.9;竞品列:Opus 79.3, GPT-5.5 84.4, Gemini 85.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: PostTrainBench · figure: images/13.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: posttrain-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=16.5; Seed2.1 Turbo=18.3. All page score panels are images; no DOM table.Numerically incompatible with minimax-m3 PostTrainBench rows (M3 37.1/0.37) - different harness or scale; do not merge without protocol confirmation. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 18.3;竞品列:Opus 27.4, GPT-5.5 25.0, Gemini '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: FrontierScience-Research · figure: images/13.png (panel table; Seed2.1 Pro / Turbo columns) · quote_snippet: It maintains competitive on frontier research benchmarks such as FrontierScience-Research
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: frontier-science-research not yet in data/benchmarks.json. Vision-read CONFIRMED (2026-09-01, native-resolution archive panel): Seed2.1 Pro=28.3; Seed2.1 Turbo=33.3. All page score panels are images; no DOM table. 2026-09-01 audit: upgraded to reported - Seed2.1 Pro 28.3, Seed2.1 Turbo 33.3. Same-table competitor cells: Opus 4.7 20.0 / GPT-5.5 33.9 / Gemini 3.1 Pro 16.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: FrontierCS · figure: images/13.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: frontiercs not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=46.3; Seed2.1 Turbo=50.8. All page score panels are images; no DOM table.Turbo cell bolded as row best per vision read. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 50.8;竞品列:Opus '-', GPT-5.5 58.6, Gemini 64.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Foundational capabilities including multimodal understanding maintain industry leadership, further empowering agentic scenarios · row: HorizonMath · figure: images/13.png (官方评测表,Seed2.1 Pro / Seed2.1 Turbo vs Claude Opus 4.7 / GPT-5.5 / Gemini 3.1 Pro)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: horizonmath not yet in data/benchmarks.json. Vision-read value (unconfirmed): Seed2.1 Pro=2.0; Seed2.1 Turbo=2.0. All page score panels are images; no DOM table.Very low absolute values (Pro/Turbo 2.0; GPT-5.5 7.1 best per vision read); scale/normalization unexplained on page. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表 Seed2.1 Turbo 列 2;竞品列:Opus 4.0, GPT-5.5 7.1, Gemini 4.0。
Seed2.1 Turbo
Seed2.1 Turbo 为系列提速档,与 Pro 在官方对照表同表评测。账本评测行以 Pro 列为准记录,Turbo 同表对照值见归档(如 Workspace Bench 54.7、GDPval 82.7)。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。
Seed2.1 Preview
Seed2.1 Preview 为参与 Code Arena: Frontend 第三方榜单的预览条目(榜单排名 8/1539)。本发布未为其单列正式评测行。
- 输入模态
- 官方资料未说明
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。