Seed1.8 / Seed1.8 Thinking / Seed1.5-VL
ByteDance Seed / 豆包 · 2025-12(仅精确到月) · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Seed1.8
Seed1.8 被发布文定位为通用化 Agentic 模型,覆盖 GUI/浏览器/移动操作、代理搜索、代理编码与经济价值领域任务,并提供三种思考模式。评测横跨代理、数学、STEM、知识、多模态与长视频等约 13 组:GAIA 87.4、BrowseComp-zh 81.3、AIME-25 94.3。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: OSWorld · figure: images/01.png (Device Use 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=61.9. All page score panels are images; no DOM table.Competitor cells (vision): Seed1.5-VL 36.7, Claude-Sonnet-4.5 62.9, Gemini-2.5-pro 13.3, GPT-O3-CUA 38.1. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:36.7 / 62.9 / 13.3 / 38.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: Realbench · figure: images/01.png (Device Use 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: realbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=49.1. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:46.0 / 39.3 / 38.4 / 34.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: Online-Mind2web · figure: images/01.png (Device Use 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: online-mind2web not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=85.9. All page score panels are images; no DOM table.Distinct live-web variant from static Mind2Web (see qwen3-coder mind2web row); kept as separate id. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:76.4 / '-' / 69.0 / 61.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: AndroidWorld · figure: images/01.png (Device Use 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=70.7. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:62.1 / 56.0 / 69.7 / '-'。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: BrowseComp-en · figure: Chart image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4og2ymjazm9e9.png (Seed1.8 vs GPT-5-high, Claude-Sonnet-4.5, Gemini-2.5-pro, Gemini-3-pro; General Agentic Search + Visual Search; asterisk = 'sourced from public technical reports', superscript 1 = 'official full-set scores') · quote_snippet: it achieves a high score of 67.6 in the BrowseComp-en benchmark, surpassing other leading models such as Gemini-3-Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks.json. Prose value; competitor cells carry the public-technical-report asterisk (GPT-5-high 54.9*, Claude-Sonnet-4.5 24.1*, Gemini-2.5-pro 9.9*, Gemini-3-pro 37.8).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: BrowseComp-zh · figure: images/02.png (General Agentic Search 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: browsecomp not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=81.3. All page score panels are images; no DOM table.Competitor cells (vision): GPT-5-high 63.0*, Claude-Sonnet-4.5 42.4*, Gemini-2.5-pro 34.6, Gemini-3-pro 51.6. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 63.0*, Claude-Sonnet-4.5 42.4*, Gemini-2.5-pro 34.6, Gemini-3-pro 51.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: GAIA · figure: images/02.png (General Agentic Search 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=87.4. All page score panels are images; no DOM table.Agentic-search condition; harness/tools not stated on page. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 76.7, Claude-Sonnet-4.5 66.0, Gemini-2.5-pro 57.3, Gemini-3-pro 74.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: WideSearch · figure: images/02.png (General Agentic Search 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: widesearch not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=63.8. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 62.2, Claude-Sonnet-4.5 65.7, Gemini-2.5-pro 52.6, Gemini-3-pro 57.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: HLE (text-only) · figure: images/02.png (General Agentic Search 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=40.9. All page score panels are images; no DOM table.Competitor cells: GPT-5-high 41.7*, Claude-Sonnet-4.5 32.0*, Gemini-2.5-pro 19.8, Gemini-3-pro 45.8 with superscript-1 (official full-set scores) - mixed sourcing in one row. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 41.7*, Claude-Sonnet-4.5 32.0*, Gemini-2.5-pro 19.8, Gemini-3-pro 45.8*1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: MM-BrowseComp · figure: images/02.png (Visual Search 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mm-browsecomp not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=46.3. All page score panels are images; no DOM table.Visual-search group; Claude-Sonnet-4.5 cell is a dash. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 27.7, Claude '-', Gemini-2.5-pro 7.2, Gemini-3-pro 25.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: HLE-VL · figure: images/02.png (Visual Search 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=31.5. All page score panels are images; no DOM table.Vision-search group; Claude-Sonnet-4.5 cell is a dash. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 24.6, Claude '-', Gemini-2.5-pro 19.0, Gemini-3-pro 36.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: SWE-Bench Verified · figure: images/03.png (Agentic Coding 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=72.9. All page score panels are images; no DOM table.Competitor cells carry the public-technical-report asterisk (GPT-5-high 74.9*, Claude-Sonnet-4.5 77.2*, Gemini-2.5-pro 59.6*, Gemini-3-pro 76.2*). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 74.9*, Claude-Sonnet-4.5 77.2*, Gemini-2.5-pro 59.6*, Gemini-3-pro 76.2*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: Multi-SWE-Bench · figure: images/03.png (Agentic Coding 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multi-swebench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=42.0. All page score panels are images; no DOM table.Chart does not split mini/flash sub-variants (qwen3-coder chart does); variant left unspecified until confirmed. Claude-Sonnet-4.5 44.3* row best. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 41.7, Claude-Sonnet-4.5 44.3*, Gemini-2.5-pro 20.7, Gemini-3-pro 42.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: AInstein-SWE-Bench · figure: images/03.png (Agentic Coding 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ainstein-swe-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=36.7. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 35.4, Claude-Sonnet-4.5 33.7, Gemini-2.5-pro 19.3, Gemini-3-pro 42.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: Terminal Bench 2.0 · figure: images/03.png (Agentic Coding 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=45.2. All page score panels are images; no DOM table.Variant 2.0 - not mergeable with Terminal-Bench 2.1 rows elsewhere. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 35.2*, Claude-Sonnet-4.5 42.8*, Gemini-2.5-pro 32.6*, Gemini-3-pro 54.2*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: U-Artifacts · figure: images/03.png (Agentic Coding 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: u-artifacts not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=49.2. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 56.8, Claude-Sonnet-4.5 37.3, Gemini-2.5-pro 33.4, Gemini-3-pro 57.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: FinSearchComp(T2&T3) · figure: Chart image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4og2ymjazrscc.png (Seed1.8 vs GPT-5-high, Claude-Sonnet-4.5, Gemini-2.5-pro, Gemini-3-pro; FinSearchComp(T2&T3), XpertBench by field, WorldTravel multi-modal/text; caption 'Scores related to WorldTravel are based on the best of five attempts.') · quote_snippet: The evaluations on FinSearchComp and XpertBench show that the model delivers relatively consistent and efficient performance
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: finsearchcomp not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=62.8. All page score panels are images; no DOM table.Prose names the benchmark with a qualitative claim; value in image only. 2026-09-01 audit: value upgraded from not_extracted to reported - archived chart images/04.png re-read at native resolution, cell reads 62.8. Same-table competitor cells (GPT-5-high / Claude-Sonnet-4.5 / Gemini-2.5-pro / Gemini-3-pro): 64.5 / 58.6 / 34.0 / 49.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: XpertBench (by field) · figure: images/04.png (archive of the five-panel results table; XpertBench block) · quote_snippet: The evaluations on FinSearchComp and XpertBench show that the model delivers relatively consistent and efficient performance
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: xpertbench not yet in data/benchmarks.json. Per-field vision reads (unconfirmed): Law 55.2, Fin 62.0, Edu 47.9, Research 31.4, Humanities 60.2; GPT-5-high leads most fields. 2026-09-01 audit: archived chart images/04.png re-read at native resolution - per-field Seed1.8 cells CONFIRMED as printed: Law 55.2, Fin 62.0, Edu 47.9, Research 31.4, Humanities 60.2. Kept not_extracted because the row aggregates five separately-printed sub-fields with no single printed total; sub-field values are now confirmed rather than unconfirmed. Same-table competitor cells (GPT-5-high / Claude-Sonnet-4.5 / Gemini-2.5-pro / Gemini-3-pro): Law 54.7/58.7/47.3/52.3, Fin 64.5/44.5/30.3/56.1, Edu 56.9/44.5/47.9/49.2, Research 48.2/27.5/25.5/34.9, Humanities 68.5/54.9/52.3/68.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: WorldTravel (multi-modal) · figure: Chart image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4og2ymjazrscc.png (Seed1.8 vs GPT-5-high, Claude-Sonnet-4.5, Gemini-2.5-pro, Gemini-3-pro; FinSearchComp(T2&T3), XpertBench by field, WorldTravel multi-modal/text; caption 'Scores related to WorldTravel are based on the best of five attempts.') · quote_snippet: Seed1.8 achieves a score of 47.2 on the WorldTravel multimodal application task
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "best of five attempts",
"judge": null
}new-benchmark: worldtravel not yet in data/benchmarks.json. Prose value; page footnote: 'Scores related to WorldTravel are based on the best of five attempts.' Gemini-3-pro ties 47.2 per vision read.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Seed1.8's Generalized Agentic Capabilities Proven Performance in Diverse Real-World Tasks · row: WorldTravel (text) · figure: images/04.png (Economically Valuable Fields 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: worldtravel not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=52.1. All page score panels are images; no DOM table.Text-condition row appears only in chart; prose cites the multi-modal 47.2 value. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:WorldTravel text 行:GPT-5-high 56.4, Claude-Sonnet-4.5 53.3, Gemini-2.5-pro 44.5, Gemini-3-pro 53.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: AIME-25 · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-25 not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=94.3. All page score panels are images; no DOM table.Competitor cells carry public-technical-report asterisks (Gemini-3-pro 95.0* row best). 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 94.6*, Claude-Sonnet-4.5 87.0*, Gemini-2.5-pro 88.0*, Gemini-3-pro 95.0*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: HMMT25(Feb) · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hmmt-25 not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=89.7. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 88.3*, Claude-Sonnet-4.5 66.7, Gemini-2.5-pro 86.7, Gemini-3-pro 97.5*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: BeyondAIME · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: beyondaime not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=77.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 74.0, Claude-Sonnet-4.5 62.0, Gemini-2.5-pro 62.0, Gemini-3-pro 83.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: AMO-Bench · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: amo-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=60.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 50.0, Claude-Sonnet-4.5 32.0, Gemini-2.5-pro 38.7, Gemini-3-pro 64.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: IMO-AnswerBench · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: imo-answerbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=76.3. All page score panels are images; no DOM table.Chart marks the Seed cell 'w/code tools' - tool condition applies to Seed1.8's row. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 76.0*, Claude-Sonnet-4.5 68.3, Gemini-2.5-pro 57.5, Gemini-3-pro 83.3*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: GPQA-Diamond · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=83.8. All page score panels are images; no DOM table.Existing benchmark id gpqa is GPQA Diamond; chart label GPQA-Diamond. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 85.7*, Claude-Sonnet-4.5 83.4*, Gemini-2.5-pro 86.4*, Gemini-3-pro 91.9*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: PHYBench · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: phybench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=41.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 40.0, Claude-Sonnet-4.5 31.0, Gemini-2.5-pro 48.0, Gemini-3-pro 59.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: BIOBench · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: biobench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=42.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 48.0, Claude-Sonnet-4.5 44.6, Gemini-2.5-pro 41.5, Gemini-3-pro 51.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: KOR-Bench · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: kor-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=76.2. All page score panels are images; no DOM table.Distinct from korb (KernelBench); KOR-Bench is a rule-knowledge-operation-reasoning benchmark. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 77.4, Claude-Sonnet-4.5 74.5, Gemini-2.5-pro 74.2, Gemini-3-pro 75.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: ARC-AGI-1 · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=67.9. All page score panels are images; no DOM table.Maps to existing benchmark arc-agi, ARC-AGI-1 (original) variant; competitor cells carry public-report asterisks. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 65.7*, Claude-Sonnet-4.5 63.7*, Gemini-2.5-pro 37.0*, Gemini-3-pro 75.0*。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: MMLU · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=92.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 93.8, Claude-Sonnet-4.5 93.1, Gemini-2.5-pro 92.9, Gemini-3-pro 93.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: MMLU-pro · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8=84.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 87.2, Claude-Sonnet-4.5 88.8, Gemini-2.5-pro 86.9, Gemini-3-pro 90.2。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: SuperGPQA · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: supergpqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=64.8. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 66.8, Claude-Sonnet-4.5 66.1, Gemini-2.5-pro 64.9, Gemini-3-pro 75.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: LPFQA · figure: images/06.png (Math/STEM/Knowledge 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: lpfqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=49.1. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 54.4, Claude-Sonnet-4.5 49.5, Gemini-2.5-pro 47.7, Gemini-3-pro 51.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: Inverse IFEval · figure: images/07.png (Complex Instruction Following 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: inverse-ifeval not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=80.3. All page score panels are images; no DOM table.Distinct from existing ifeval (IFEval); inverse-constraint variant benchmark. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 78.9, Claude-Sonnet-4.5 70.2, Gemini-2.5-pro 75.3, Gemini-3-pro 80.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: MARS-Bench · figure: images/07.png (Complex Instruction Following 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mars-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=70.1. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 77.2, Claude-Sonnet-4.5 72.5, Gemini-2.5-pro 73.6, Gemini-3-pro 80.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: MultiChallenge · figure: images/07.png (Complex Instruction Following 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multichallenge not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=66.7. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 69.6, Claude-Sonnet-4.5 57.2, Gemini-2.5-pro 55.4, Gemini-3-pro 67.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: Collie-Hard · figure: images/07.png (Complex Instruction Following 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: collie not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=72.6. All page score panels are images; no DOM table.GPT-5-high cell 99.0* carries the public-report asterisk. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 99.0*, Claude-Sonnet-4.5 77.6, Gemini-2.5-pro 69.5, Gemini-3-pro 95.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for LLM Capabilities On Par with Leading Generalized Models · row: EIFBench · figure: images/07.png (Complex Instruction Following 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: eifbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=48.6. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:GPT-5-high 66.7, Claude-Sonnet-4.5 47.0, Gemini-2.5-pro 44.7, Gemini-3-pro 50.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: TVBench · figure: images/12.png (Seed1.8 Motion & Perception 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tvbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=71.5. All page score panels are images; no DOM table.Competitor columns limited to Gemini-2.5-pro 67.4, Gemini-3-pro 71.1, Seed1.5-VL 66.6. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:67.4 / 71.1 / 66.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: TempCompass · figure: images/12.png (Seed1.8 Motion & Perception 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tempcompass not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=86.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:83.9 / 88.0 / 83.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: TOMATO · figure: images/12.png (Seed1.8 Motion & Perception 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tomato not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=60.8. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:50.3 / 55.8 / 44.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: EgoTempo · figure: images/12.png (Seed1.8 Motion & Perception 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: egotempo not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=67.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:58.1 / 65.4 / 51.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MotionBench · figure: images/12.png (Seed1.8 Motion & Perception 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: motionbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=70.6. All page score panels are images; no DOM table.Gemini-2.5-pro 66.3* and Gemini-3-pro 70.3* carry the public-report asterisk. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:66.3* / 70.3* / 68.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: Countix · figure: images/12.png (Seed1.8 Motion & Perception 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: countix not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=31.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:18.6 / 18.7 / 26.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: VideoMME double-dagger · figure: Chart image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4og2ymjb26xj6.png (Seed1.8 vs Gemini-2.5-pro, Gemini-3-pro, Seed1.5-VL; Long Video group; VideoMME carries the double-dagger footnote 'subtitles are included for evaluation') · quote_snippet: achieving a high score of 87.8 on the VideoMME benchmark
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: video-mme not yet in data/benchmarks.json. Prose value; chart carries the double-dagger footnote 'subtitles are included for evaluation'. Prose also discloses the VideoCut video tool (slow-motion playback of selected segments) used for long-video reasoning - a tool-enabled condition; Gemini-2.5-pro 86.9* / Gemini-3-pro 88.4* carry the public-report asterisk.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: CGBench · figure: images/13.png (Seed1.8 Long Video 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cgbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=62.4. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:64.6 / 64.5 / 57.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: LongVideoBench · figure: images/13.png (Seed1.8 Long Video 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: longvideobench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=77.4. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:77.6 / 76.7 / 74.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: LVBench · figure: images/13.png (Seed1.8 Long Video 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: lvbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8=73.0. All page score panels are images; no DOM table.Gemini-3-pro cell is a dash. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:73.5 / '-' / 64.6。
Seed1.8 Thinking
Seed1.8 Thinking 是 Seed1.8 的思考模式呈现(发布图表中的 Thinking 列),多模态推理与视觉问答组评测以该列报告:MMMU 83.4、VideoMME(含字幕)87.8、MM-BrowseComp 46.3。本次发布未为其单独报告区别于主列的规格数值。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MMMU · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8-thinking=83.4. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:79.8 / 83.3 / 82.0* / 87.0 / 77.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MMMU-Pro · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8-thinking=73.2. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:68.0* / 76.0* / 68.0* / 81.0* / 67.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MathVista · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Vision-read value (unconfirmed): seed-1-8-thinking=87.7. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:80.4 / 80.6 / 82.7* / 89.8 / 85.6。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MathVision · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mathvision not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=81.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:73.6 / 77.2 / 73.3* / 86.1 / 68.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: DynaMath · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: dyna-math not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=61.5. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:52.7 / 61.5 / 56.3 / 63.3 / 57.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: LogicVista · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: logicvista not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=78.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:71.8 / 70.0 / 73.8 / 80.8 / 73.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: EMMA · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: emma not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=60.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:53.5 / 61.7 / 59.4 / 66.5 / 49.5。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: SFE · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: sfe not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=51.2. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:50.5 / 46.0 / 47.7 / 61.9 / 44.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: ZeroBench (main) · figure: Chart image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4og2ymjb22697.png (column 'Seed1.8 Thinking' vs Claude-Sonnet-4.5, GPT-5.1-high, Gemini-2.5-pro, Gemini-3-pro, Seed1.5-VL Thinking; MultiModal Reasoning group) · quote_snippet: Seed1.8 achieves a top score of 11.0 on the highly challenging visual reasoning benchmark ZeroBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: zerobench not yet in data/benchmarks.json. Prose value; chart row label 'ZeroBench (main)'.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: VPCT · figure: images/09.png (Seed1.8-Thinking Multimodal Reasoning 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vpct not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=61.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:41.0 / 56.0 / 52.0 / 90.0 / 35.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: VLMsAreBiased · figure: Chart image https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/user-upload/4og2ymjb24c2n.png (column 'Seed1.8 Thinking' vs same 5 columns; General Visual QA group, 10 rows) · quote_snippet: Seed1.8 achieves a score of 62.0 on the VLMsAreBiased benchmark, substantially outperforming other models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vlmsarebiased not yet in data/benchmarks.json. Prose value; row best per vision read (Gemini-3-pro 50.6* next).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: VLMsAreBlind · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vlmsareblind not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=93.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:80.9 / 84.2 / 84.3* / 97.5 / 92.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: SimpleVQA · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: simplevqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=65.4. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:48.1 / 56.1 / 62.0* / 69.7 / 63.4。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: HallusionBench · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hallusionbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=63.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:59.1 / 64.8 / 63.7* / 69.9 / 60.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MMStar · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmstar not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=79.9. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:74.1 / 77.8 / 77.5 / 83.1 / 77.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MMBench v1.1 EN · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=91.6. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:87.5 / 85.4 / 90.1 / 93.3 / 89.9。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MMBench v1.1 CN · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=90.6. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:86.2 / 84.9 / 89.7 / 91.3 / 89.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MME-CC · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mme-cc not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=43.4. All page score panels are images; no DOM table.GPT-5.1-high cell is a dash. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:27.5 / '-' / 42.7 / 56.9 / 33.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MUIRBench · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: muirbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=78.7. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:71.8 / 78.2 / 77.2 / 78.2 / 72.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MMVP · figure: images/10.png (Seed1.8-Thinking General VQA 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmvp not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=86.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:74.7 / 84.3 / 70.7 / 90.0 / 69.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: BLINK · figure: images/11.png (Seed1.8-Thinking 2D&3D Spatial 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: blink not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=74.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:63.4 / 69.6 / 70.6 / 77.1 / 72.1。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: MMSIBench (circular) · figure: images/11.png (Seed1.8-Thinking 2D&3D Spatial 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmsi-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=25.8. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:17.2 / 22.3 / 17.6 / 25.4 / 11.7。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: RefSpatialBench · figure: images/11.png (Seed1.8-Thinking 2D&3D Spatial 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: refspatialbench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=56.3. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:21.7 / 28.2* / 33.6* / 65.5* / 58.8。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: ERQA · figure: images/11.png (Seed1.8-Thinking 2D&3D Spatial 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: erqa not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=58.8. All page score panels are images; no DOM table.Same benchmark id as seed-2-1 erqa rows; different model tier. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:49.8 / 60.0* / 56.0* / 70.5* / 47.0。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: DA-2K · figure: images/11.png (Seed1.8-Thinking 2D&3D Spatial 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: da-2k not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=90.7. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:68.2 / 78.6 / 76.5 / 82.1 / 85.3。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation Results for VLM Multimodal Capabilities: Outstanding Performance with a Notable Leap in Scores · row: CV-Bench · figure: images/11.png (Seed1.8-Thinking 2D&3D Spatial 表)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: cv-bench not yet in data/benchmarks.json. Vision-read value (unconfirmed): seed-1-8-thinking=88.0. All page score panels are images; no DOM table. 视觉转写自归档图(2026-09-01 复核,与先前读数一致)。同表竞品列值:79.3 / 84.6 / 85.9 / 92.0 / 84.9。
Seed1.5-VL
Seed1.5-VL 是发布文中的前代对照列(Seed1.5-VL),仅在代理与长视频图表中作为比较基线出现(如 OSWorld 36.7、Realbench 46.0),本次发布未为其报告新评测数值。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。