← 模型目录

Muse Spark 1.1

Meta / Llama · 2026-07-09 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Muse Spark 1.1

Meta 将 Muse Spark 1.1 定位为可主动管理自身 1M token 上下文窗口的智能体基础模型,随发布同步上线 Meta Model API 公测,并已在 Meta AI 应用与 meta.ai 的 Thinking 模式可用。已收录评测覆盖智能体编码与工具调用、计算机操作、网络安全及生物/化学安全评估等领域,Terminal-Bench 2.1 80.0%、OSWorld Verified 80.8%。

输入模态
文本 / 图像
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

swebench 24 out of 42 unique tasks (resolved at least once) 模型 muse-spark-1-1 · 版本 Verified Hard · 指标 unique_tasks_resolved_at_least_once · 单位 tasks 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 39 · quote_snippet: resolved, at least once, 24 out of 42 unique tasks on SWE-Bench Verified Hard

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": "multiple runs implied ('at least once'); exact run count not stated here",
  "aggregation": "count of unique tasks resolved at least once across runs",
  "judge": null
}

Value 24 is a task count, not a percentage - do not read as 24%. Meta frames it against a capability checkpoint threshold ('resolving over half of unique tasks on SWE-Bench Verified Hard'). Retained evals reuse 'the same benchmark construction, scaffolds, and compute budgets' as the Muse Spark Safety & Preparedness Report per the report footnote. Not comparable to standard SWE-bench Verified resolved_rate rows (different variant and aggregation). [2026-09-01 audit: verbatim in the PDF text layer (pypdf default extraction): "Muse Spark 1.1 resolved, at least once, 24 out of 42 unique tasks on SWE-Bench Verified Hard" - value and framing confirmed.]

打开官方来源

terminalbench 官方未公布数值 模型 muse-spark-1-1 · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: For coding capabilities (e.g, Terminal-Bench 2.1, SWE-Bench Pro), Muse Spark 1.1 trails Claude 4.8 Opus and/or GPT 5.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Relative claim only; no Terminal-Bench 2.1 number printed in the report text.

打开官方来源

swebench-pro 官方未公布数值 模型 muse-spark-1-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: For coding capabilities (e.g, Terminal-Bench 2.1, SWE-Bench Pro), Muse Spark 1.1 trails Claude 4.8 Opus and/or GPT 5.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-pro not yet in data/benchmarks/ (SWE-Bench Pro, distinct from SWE-bench Verified). Relative claim only; no number printed.

打开官方来源

deepswe 官方未公布数值 模型 muse-spark-1-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: for long-horizon agentic tasks (e.g., DeepSWE and DeepSearchQA), significant improvements still lag behind or on par best performing competitor models

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark id deepswe already introduced by the kimi-k3 batch, still absent from data/benchmarks/. Relative claim only; no DeepSWE number printed. A separate blog demo shows the model running DeepSWE tasks in OpenCode, which is a product demo, not a reported score.

打开官方来源

deepsearchqa 官方未公布数值 模型 muse-spark-1-1 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: for long-horizon agentic tasks (e.g., DeepSWE and DeepSearchQA), significant improvements still lag behind or on par best performing competitor models

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepsearchqa not yet in data/benchmarks/. Relative claim only; no number printed.

打开官方来源

mbct 53.2% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: MBCT · page: PDF p.7 · quote_snippet: MBCT | 53.2 | 54.4 | 56.0 | - | 46.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: MBCT: Muse Spark 1.0 54.4 | GPT-5.5 56.0 | Claude - | Gemini 3.1 Pro 46.2. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

vct 52.0% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: VCT · page: PDF p.7 · quote_snippet: VCT | 52.0 | 49.7 | 52.1 | - | 44.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: VCT: Muse Spark 1.0 49.7 | GPT-5.5 52.1 | Claude - | Gemini 3.1 Pro 44.6. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

hpct 61.9% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: HPCT · page: PDF p.7 · quote_snippet: HPCT | 61.9 | 55.7 | 67.7 | - | 62.9

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: HPCT: Muse Spark 1.0 55.7 | GPT-5.5 67.7 | Claude - | Gemini 3.1 Pro 62.9. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

wmdp 89.0% 模型 muse-spark-1-1 · 版本 Bio · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: WMDP-Bio · page: PDF p.7 · quote_snippet: WMDP-Bio | 89.0 | 88.4 | 90.4 | - | 89.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: WMDP-Bio: Muse Spark 1.0 88.4 | GPT-5.5 90.4 | Claude - | Gemini 3.1 Pro 89.5. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

wmdp 87.0% 模型 muse-spark-1-1 · 版本 Chem · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: WMDP-Chem · page: PDF p.7 · quote_snippet: WMDP-Chem | 87.0 | 85.6 | 85.5 | - | 86.3(6)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: WMDP-Chem: Muse Spark 1.0 85.6 | GPT-5.5 85.5 | Claude - | Gemini 3.1 Pro extracts as 86.36 - likely 86.3 with a footnote-6 marker; competitor cell kept in notes for that reason. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

lab-bench 88.0% 模型 muse-spark-1-1 · 版本 ProtocolQA · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ProtocolQA · page: PDF p.7 · quote_snippet: ProtocolQA | 88.0 | 87.3 | 78.2 | - | 88.9

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: ProtocolQA: Muse Spark 1.0 87.3 | GPT-5.5 78.2 | Claude - | Gemini 3.1 Pro 88.9. Chemical & Biological domain; id reuses lab-bench minted by meta/muse-glimmer-30b. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

seqqa 98.2% 模型 muse-spark-1-1 · 版本 agentic · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: SeqQA (agentic) · page: PDF p.7 · quote_snippet: SeqQA (agentic) | 98.2 | 97.3 | 98.2 | - | 95.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: SeqQA: Muse Spark 1.0 97.3 | GPT-5.5 98.2 | Claude - | Gemini 3.1 Pro 95.4. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

abcbench 97.0% 模型 muse-spark-1-1 · 版本 Fragment Design · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ABCBench (FragmentDesign) · page: PDF p.7 · quote_snippet: ABCBench (FragmentDesign) | 97.0 | 96.8 | 92.2 | - | 95.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: FragmentDesign: Muse Spark 1.0 96.8 | GPT-5.5 92.2 | Claude - | Gemini 3.1 Pro 95.8. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

abcbench 93.7% 模型 muse-spark-1-1 · 版本 Liquid Handling · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ABCBench (LiquidHandling) · page: PDF p.7 · quote_snippet: ABCBench (LiquidHandling) | 93.7 | 92.3 | 93.2 | - | 99.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: LiquidHandling: Muse Spark 1.0 92.3 | GPT-5.5 93.2 | Claude - | Gemini 3.1 Pro 99.2. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

abcbench 63.2% 模型 muse-spark-1-1 · 版本 Screening Evasion · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ABCBench (ScreeningEvasion) · page: PDF p.7 · quote_snippet: ABCBench (ScreeningEvasion) | 63.2 | 54.1 | - | - | -

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: ScreeningEvasion: Muse Spark 1.0 54.1 | GPT-5.5/Claude/Gemini not reported. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

bio-design-tools 55.2% 模型 muse-spark-1-1 · 版本 avg · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: BioDesignTools (avg) · page: PDF p.7 · quote_snippet: BioDesignTools (avg) | 55.2 | 39.2 | 67.4 | - | 62.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: BioDesignTools: Muse Spark 1.0 39.2 | GPT-5.5 67.4 | Claude - | Gemini 3.1 Pro 62.4. Chemical & Biological domain; average across BioDesignTools tasks. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

cybench 92.9% 模型 muse-spark-1-1 · 版本 pass@1, overall (40 challenges) · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: Cybench (pass@1) · page: PDF p.7 · quote_snippet: Cybench (pass@1) | 92.9 | 65.4 | 100.0 | 95.0 | -

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: Cybench: Muse Spark 1.0 65.4 | GPT-5.5 100.0 | Claude Opus 4.8 95.0 | Gemini 3.1 Pro not reported. Cybersecurity domain. Corroborated by Table 3: Overall (40) 92.9 vs Muse Spark 1.0 65.4 (Easy 100.0 / Medium 94.2 / Hard 76.1). Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

curated-ctfs 89.9% 模型 muse-spark-1-1 · 版本 pass@1, overall (198 challenges) · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: Curated CTFs (pass@1) · page: PDF p.7 · quote_snippet: Curated CTFs (pass@1) | 89.9 | 72.0 | - | - | -

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: Curated CTFs: Muse Spark 1.0 72.0 | GPT-5.5/Claude/Gemini not reported. Cybersecurity domain. Corroborated by section prose (89.9% pass@1, 95.7% pass@5 vs 1.0 72.0/84.1) and Table 4 Overall (198) 89.9 vs 72.0. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

cybergym 59.0% 模型 muse-spark-1-1 · 版本 pass@1 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: CyberGym (pass@1) · page: PDF p.7 · quote_snippet: CyberGym (pass@1) | 59.0 | 43.5 | 81.8 | 78.8 | -

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: CyberGym: Muse Spark 1.0 43.5 | GPT-5.5 81.8 | Claude Opus 4.8 78.8 | Gemini 3.1 Pro not reported. Cybersecurity domain; framework proxy for the Cyber 2 (vulnerability discovery) outcome. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

exploitgym 0.8% 模型 muse-spark-1-1 · 版本 pass@1 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ExploitGym (pass@1) · page: PDF p.7 · quote_snippet: ExploitGym (pass@1) | 0.8 | - | 14.8 | - | 1.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: ExploitGym: Muse Spark 1.0 not reported | GPT-5.5 14.8 | Claude not reported | Gemini 3.1 Pro 1.4. Cybersecurity domain; near-floor score explicitly shown in the scorecard. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

cyscenariobench 0.5% 模型 muse-spark-1-1 · 版本 pass@1 · 指标 pass@1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: CyScenarioBench (pass@1) · page: PDF p.7 · quote_snippet: CyScenarioBench (pass@1) | 0.5 | 0.0 | 26.0 | 16.6 | -

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: CyScenarioBench: Muse Spark 1.0 0.0 | GPT-5.5 26.0 | Claude Opus 4.8 16.6 | Gemini 3.1 Pro not reported. Cybersecurity domain; framework proxy for the Cyber 1 (end-to-end network compromise) outcome. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

social-engineering 5.1% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: SocialEngineering · page: PDF p.7 · quote_snippet: SocialEngineering | 5.1 | - | 1.2 | 7.1 | 13.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: SocialEngineering: Muse Spark 1.0 not reported | GPT-5.5 1.2 | Claude Opus 4.8 7.1 | Gemini 3.1 Pro 13.7. Cybersecurity domain; metric direction/unit not labeled in the scorecard row (lower appears better from context, but the report section should confirm before comparing). Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

airs-bench 77.0% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: AIRS-Bench · page: PDF p.7 · quote_snippet: AIRS-Bench | 77.0 | - | 86.0 | 84.0 | 83.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: AIRS-Bench: Muse Spark 1.0 not reported | GPT-5.5 86.0 | Claude Opus 4.8 84.0 | Gemini 3.1 Pro 83.0. Cybersecurity domain. Table 6 reports valid submission rate and average normalized score across 20 research tasks with 95% CIs; the scorecard cell is the valid-submission-rate lane and Table 6 should be consulted for the normalized-score lane. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

shade-arena 6.8% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: SHADE-Arena · page: PDF p.7 · quote_snippet: SHADE-Arena | 6.8 | - | 0.5 | 7.6 | 1.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: SHADE-Arena: Muse Spark 1.0 not reported | GPT-5.5 0.5 | Claude Opus 4.8 7.6 | Gemini 3.1 Pro 1.6. Loss of Control domain; SHADE-Arena (GDM) covert-subtask-success style metric - lower is the safe direction, confirmed by the GDM Stealth table narrative. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

gdm-stealth 2 / 4 challenges 模型 muse-spark-1-1 · 版本 未说明 · 指标 未说明 · 单位 challenges 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: GDM-Stealth (of 4) · page: PDF p.7 · quote_snippet: GDM-Stealth (of 4) | 2/4 | 1/4 | 1/4 | 0/4 | 3/4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: GDM-Stealth: Muse Spark 1.0 1/4 | GPT-5.5 1/4 | Claude Opus 4.8 0/4 | Gemini 3.1 Pro 3/4. Loss of Control domain; fraction of 4 stealth challenges. Corroborated by Table 9 narrative. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

gdm-situational-awareness 55.1% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: GDM Situational Awareness · page: PDF p.7 · quote_snippet: GDM Situational Awareness | 55.1 | 29.3 | 60.0 | 53.1 | 54.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: GDM SA: Muse Spark 1.0 29.3 | GPT-5.5 60.0 | Claude Opus 4.8 53.1 | Gemini 3.1 Pro 54.5. Loss of Control domain. Table 10 narrative adds: Muse Spark 1.1 resolves 7 out of 11 tasks on GDM Situational Awareness. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.

打开官方来源

mcp-atlas 88.1% 模型 muse-spark-1-1 · 版本 Scaled tool use · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: MCP Atlas / Scaled tool use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: MCP Atlas / Scaled tool use = 88.1%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mcp-atlas (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: 82.2 | Gemini 78.2 | Opus 4.8 82.2 | GPT 5.5 75.3。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

jobbench 54.7% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: JobBench / Professional tool use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: JobBench / Professional tool use = 54.7%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: jobbench (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 17.0 | Gemini 15.9 | Opus 4.8 48.4 | GPT 5.5 38.3。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

toolathlon-verified 75.6% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: Toolathlon-Verified / Personal tool use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Toolathlon-Verified / Personal tool use = 75.6%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: toolathlon-verified (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 49.4 | Gemini 61.1 | Opus 4.8 76.2(lead) | GPT 5.5 73.5。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

osworld 80.8% 模型 muse-spark-1-1 · 版本 Verified · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: OSWorld-Verified / Agentic computer use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: OSWorld-Verified / Agentic computer use = 80.8%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: osworld (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 53.3 | Gemini 76.2 | Opus 4.8 83.4(lead) | GPT 5.5 78.7。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

hlehle 62.1% 模型 muse-spark-1-1 · 版本 w/ tools · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: Humanity's Last Exam / Multidisciplinary reasoning (w/ tools) · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Humanity's Last Exam / Multidisciplinary reasoning (w/ tools) = 62.1%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hlehle (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 50.4 | Gemini 51.4 | Opus 4.8 57.9 | GPT 5.5 52.2。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

finance-agent-v2 57.2% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: Finance Agent v2 / Agentic financial analysis(原图拼写 anaysis) · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Finance Agent v2 / Agentic financial analysis(原图拼写 anaysis) = 57.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: finance-agent-v2 (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 无数据(图中 -) | Gemini 43.0 | Opus 4.8 53.9 | GPT 5.5 51.8。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

terminalbench 80.0% 模型 muse-spark-1-1 · 版本 2.1 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: Terminal-Bench 2.1 / Agentic terminal coding · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Terminal-Bench 2.1 / Agentic terminal coding = 80.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: terminalbench (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 67.3 | Gemini 70.3 | Opus 4.8 82.7 | GPT 5.5 83.4(lead)。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。 跨 release 注意: muse-spark-1-2 发布图(2026-08-05)中同模型 Terminal-Bench 2.1 为 76.2(muse-code harness),与本表 80.0 差 3.8pp —— 两次 Meta 自报,发布时点/脚手架不同,比较前先对齐 harness。

打开官方来源

swebench-pro 61.5% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: SWE-Bench Pro / Diverse software engineering · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: SWE-Bench Pro / Diverse software engineering = 61.5%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-pro (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 55.0 | Gemini 54.2 | Opus 4.8 69.2(lead) | GPT 5.5 58.6。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

deepswe 53.3% 模型 muse-spark-1-1 · 版本 1.1 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: DeepSWE 1.1 / Long-horizon agentic coding · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: DeepSWE 1.1 / Long-horizon agentic coding = 53.3%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepswe (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 10.0 | Gemini 12.0 | Opus 4.8 59.0 | GPT 5.5 67.0(lead)。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

charxiv-reasoning 88.4% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: CharXiv Reasoning / Chart QA · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: CharXiv Reasoning / Chart QA = 88.4%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: charxiv-reasoning (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 88.9 | Gemini 81.6 | Opus 4.8 89.9(lead) | GPT 5.5 84.8。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

babyvision 76.3% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: BabyVision / Visual reasoning · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: BabyVision / Visual reasoning = 76.3%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: babyvision (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 39.9 | Gemini 51.5 | Opus 4.8 81.2 | GPT 5.5 83.6(lead)。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。

打开官方来源

deepsearchqa 84.9% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: DeepSearchQA (images/05.png) · figure: models/2026-07-09-muse-spark-1-1/images/05.png · quote_snippet: DeepSearchQA (images/05.png) = 84.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: deepsearchqa (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: GPT 5.5 87.8(lead) | Opus 4.8 84.3 | Muse Spark 76.8 | Gemini 71.3

打开官方来源

vibe-code-bench 72.2% 模型 muse-spark-1-1 · 版本 v1.1 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: Vibe Code Bench v1.1 (images/08.png) · figure: models/2026-07-09-muse-spark-1-1/images/08.png · quote_snippet: Vibe Code Bench v1.1 (images/08.png) = 72.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: vibe-code-bench (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: Muse Spark 19.7(仅两代对比图)

打开官方来源

swe-atlas-codebase-qna 42.0% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: SWE Atlas - Codebase QnA (images/08.png) · figure: models/2026-07-09-muse-spark-1-1/images/08.png · quote_snippet: SWE Atlas - Codebase QnA (images/08.png) = 42.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swe-atlas-codebase-qna (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: Muse Spark 24.2(仅两代对比图)

打开官方来源

meta-internal-coding-bench 68.3% 模型 muse-spark-1-1 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluations (launch-post benchmark table/charts) · row: Meta Internal Coding Bench (images/09.png) · figure: models/2026-07-09-muse-spark-1-1/images/09.png · quote_snippet: Meta Internal Coding Bench (images/09.png) = 68.3%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: meta-internal-coding-bench (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: Opus 4.8 69.0(lead) | GPT 5.5 67.1 | Gemini 59.2 | Muse Spark 58.8

打开官方来源