Muse Spark 1.1
Meta / Llama · 2026-07-09 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Muse Spark 1.1
Meta 将 Muse Spark 1.1 定位为可主动管理自身 1M token 上下文窗口的智能体基础模型,随发布同步上线 Meta Model API 公测,并已在 Meta AI 应用与 meta.ai 的 Thinking 模式可用。已收录评测覆盖智能体编码与工具调用、计算机操作、网络安全及生物/化学安全评估等领域,Terminal-Bench 2.1 80.0%、OSWorld Verified 80.8%。
- 输入模态
- 文本 / 图像
- 上下文
- 1M
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 39 · quote_snippet: resolved, at least once, 24 out of 42 unique tasks on SWE-Bench Verified Hard
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": "multiple runs implied ('at least once'); exact run count not stated here",
"aggregation": "count of unique tasks resolved at least once across runs",
"judge": null
}Value 24 is a task count, not a percentage - do not read as 24%. Meta frames it against a capability checkpoint threshold ('resolving over half of unique tasks on SWE-Bench Verified Hard'). Retained evals reuse 'the same benchmark construction, scaffolds, and compute budgets' as the Muse Spark Safety & Preparedness Report per the report footnote. Not comparable to standard SWE-bench Verified resolved_rate rows (different variant and aggregation). [2026-09-01 audit: verbatim in the PDF text layer (pypdf default extraction): "Muse Spark 1.1 resolved, at least once, 24 out of 42 unique tasks on SWE-Bench Verified Hard" - value and framing confirmed.]
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: For coding capabilities (e.g, Terminal-Bench 2.1, SWE-Bench Pro), Muse Spark 1.1 trails Claude 4.8 Opus and/or GPT 5.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Relative claim only; no Terminal-Bench 2.1 number printed in the report text.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: For coding capabilities (e.g, Terminal-Bench 2.1, SWE-Bench Pro), Muse Spark 1.1 trails Claude 4.8 Opus and/or GPT 5.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-pro not yet in data/benchmarks/ (SWE-Bench Pro, distinct from SWE-bench Verified). Relative claim only; no number printed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: for long-horizon agentic tasks (e.g., DeepSWE and DeepSearchQA), significant improvements still lag behind or on par best performing competitor models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark id deepswe already introduced by the kimi-k3 batch, still absent from data/benchmarks/. Relative claim only; no DeepSWE number printed. A separate blog demo shows the model running DeepSWE tasks in OpenCode, which is a product demo, not a reported score.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 2.3.1.1 Acceleration of AI Development (Results) · page: 40 · quote_snippet: for long-horizon agentic tasks (e.g., DeepSWE and DeepSearchQA), significant improvements still lag behind or on par best performing competitor models
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepsearchqa not yet in data/benchmarks/. Relative claim only; no number printed.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: MBCT · page: PDF p.7 · quote_snippet: MBCT | 53.2 | 54.4 | 56.0 | - | 46.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: MBCT: Muse Spark 1.0 54.4 | GPT-5.5 56.0 | Claude - | Gemini 3.1 Pro 46.2. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: VCT · page: PDF p.7 · quote_snippet: VCT | 52.0 | 49.7 | 52.1 | - | 44.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: VCT: Muse Spark 1.0 49.7 | GPT-5.5 52.1 | Claude - | Gemini 3.1 Pro 44.6. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: HPCT · page: PDF p.7 · quote_snippet: HPCT | 61.9 | 55.7 | 67.7 | - | 62.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: HPCT: Muse Spark 1.0 55.7 | GPT-5.5 67.7 | Claude - | Gemini 3.1 Pro 62.9. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: WMDP-Bio · page: PDF p.7 · quote_snippet: WMDP-Bio | 89.0 | 88.4 | 90.4 | - | 89.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: WMDP-Bio: Muse Spark 1.0 88.4 | GPT-5.5 90.4 | Claude - | Gemini 3.1 Pro 89.5. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: WMDP-Chem · page: PDF p.7 · quote_snippet: WMDP-Chem | 87.0 | 85.6 | 85.5 | - | 86.3(6)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: WMDP-Chem: Muse Spark 1.0 85.6 | GPT-5.5 85.5 | Claude - | Gemini 3.1 Pro extracts as 86.36 - likely 86.3 with a footnote-6 marker; competitor cell kept in notes for that reason. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ProtocolQA · page: PDF p.7 · quote_snippet: ProtocolQA | 88.0 | 87.3 | 78.2 | - | 88.9
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: ProtocolQA: Muse Spark 1.0 87.3 | GPT-5.5 78.2 | Claude - | Gemini 3.1 Pro 88.9. Chemical & Biological domain; id reuses lab-bench minted by meta/muse-glimmer-30b. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: SeqQA (agentic) · page: PDF p.7 · quote_snippet: SeqQA (agentic) | 98.2 | 97.3 | 98.2 | - | 95.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: SeqQA: Muse Spark 1.0 97.3 | GPT-5.5 98.2 | Claude - | Gemini 3.1 Pro 95.4. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ABCBench (FragmentDesign) · page: PDF p.7 · quote_snippet: ABCBench (FragmentDesign) | 97.0 | 96.8 | 92.2 | - | 95.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: FragmentDesign: Muse Spark 1.0 96.8 | GPT-5.5 92.2 | Claude - | Gemini 3.1 Pro 95.8. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ABCBench (LiquidHandling) · page: PDF p.7 · quote_snippet: ABCBench (LiquidHandling) | 93.7 | 92.3 | 93.2 | - | 99.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: LiquidHandling: Muse Spark 1.0 92.3 | GPT-5.5 93.2 | Claude - | Gemini 3.1 Pro 99.2. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ABCBench (ScreeningEvasion) · page: PDF p.7 · quote_snippet: ABCBench (ScreeningEvasion) | 63.2 | 54.1 | - | - | -
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: ScreeningEvasion: Muse Spark 1.0 54.1 | GPT-5.5/Claude/Gemini not reported. Chemical & Biological domain. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: BioDesignTools (avg) · page: PDF p.7 · quote_snippet: BioDesignTools (avg) | 55.2 | 39.2 | 67.4 | - | 62.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: BioDesignTools: Muse Spark 1.0 39.2 | GPT-5.5 67.4 | Claude - | Gemini 3.1 Pro 62.4. Chemical & Biological domain; average across BioDesignTools tasks. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: Cybench (pass@1) · page: PDF p.7 · quote_snippet: Cybench (pass@1) | 92.9 | 65.4 | 100.0 | 95.0 | -
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: Cybench: Muse Spark 1.0 65.4 | GPT-5.5 100.0 | Claude Opus 4.8 95.0 | Gemini 3.1 Pro not reported. Cybersecurity domain. Corroborated by Table 3: Overall (40) 92.9 vs Muse Spark 1.0 65.4 (Easy 100.0 / Medium 94.2 / Hard 76.1). Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: Curated CTFs (pass@1) · page: PDF p.7 · quote_snippet: Curated CTFs (pass@1) | 89.9 | 72.0 | - | - | -
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: Curated CTFs: Muse Spark 1.0 72.0 | GPT-5.5/Claude/Gemini not reported. Cybersecurity domain. Corroborated by section prose (89.9% pass@1, 95.7% pass@5 vs 1.0 72.0/84.1) and Table 4 Overall (198) 89.9 vs 72.0. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: CyberGym (pass@1) · page: PDF p.7 · quote_snippet: CyberGym (pass@1) | 59.0 | 43.5 | 81.8 | 78.8 | -
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: CyberGym: Muse Spark 1.0 43.5 | GPT-5.5 81.8 | Claude Opus 4.8 78.8 | Gemini 3.1 Pro not reported. Cybersecurity domain; framework proxy for the Cyber 2 (vulnerability discovery) outcome. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: ExploitGym (pass@1) · page: PDF p.7 · quote_snippet: ExploitGym (pass@1) | 0.8 | - | 14.8 | - | 1.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: ExploitGym: Muse Spark 1.0 not reported | GPT-5.5 14.8 | Claude not reported | Gemini 3.1 Pro 1.4. Cybersecurity domain; near-floor score explicitly shown in the scorecard. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: CyScenarioBench (pass@1) · page: PDF p.7 · quote_snippet: CyScenarioBench (pass@1) | 0.5 | 0.0 | 26.0 | 16.6 | -
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: CyScenarioBench: Muse Spark 1.0 0.0 | GPT-5.5 26.0 | Claude Opus 4.8 16.6 | Gemini 3.1 Pro not reported. Cybersecurity domain; framework proxy for the Cyber 1 (end-to-end network compromise) outcome. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: SocialEngineering · page: PDF p.7 · quote_snippet: SocialEngineering | 5.1 | - | 1.2 | 7.1 | 13.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: SocialEngineering: Muse Spark 1.0 not reported | GPT-5.5 1.2 | Claude Opus 4.8 7.1 | Gemini 3.1 Pro 13.7. Cybersecurity domain; metric direction/unit not labeled in the scorecard row (lower appears better from context, but the report section should confirm before comparing). Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: AIRS-Bench · page: PDF p.7 · quote_snippet: AIRS-Bench | 77.0 | - | 86.0 | 84.0 | 83.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: AIRS-Bench: Muse Spark 1.0 not reported | GPT-5.5 86.0 | Claude Opus 4.8 84.0 | Gemini 3.1 Pro 83.0. Cybersecurity domain. Table 6 reports valid submission rate and average normalized score across 20 research tasks with 95% CIs; the scorecard cell is the valid-submission-rate lane and Table 6 should be consulted for the normalized-score lane. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: SHADE-Arena · page: PDF p.7 · quote_snippet: SHADE-Arena | 6.8 | - | 0.5 | 7.6 | 1.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: SHADE-Arena: Muse Spark 1.0 not reported | GPT-5.5 0.5 | Claude Opus 4.8 7.6 | Gemini 3.1 Pro 1.6. Loss of Control domain; SHADE-Arena (GDM) covert-subtask-success style metric - lower is the safe direction, confirmed by the GDM Stealth table narrative. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: GDM-Stealth (of 4) · page: PDF p.7 · quote_snippet: GDM-Stealth (of 4) | 2/4 | 1/4 | 1/4 | 0/4 | 3/4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: GDM-Stealth: Muse Spark 1.0 1/4 | GPT-5.5 1/4 | Claude Opus 4.8 0/4 | Gemini 3.1 Pro 3/4. Loss of Control domain; fraction of 4 stealth challenges. Corroborated by Table 9 narrative. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Table 1 Muse Spark 1.1 capabilities scorecard · table: PDF Table 1 (pypdf layout-mode extraction, no OCR) · row: GDM Situational Awareness · page: PDF p.7 · quote_snippet: GDM Situational Awareness | 55.1 | 29.3 | 60.0 | 53.1 | 54.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark row added 2026-09-01 audit: the batch-5 premise that report tables are unreadable held only for default pypdf extraction - layout mode separates columns cleanly. Competitors: GDM SA: Muse Spark 1.0 29.3 | GPT-5.5 60.0 | Claude Opus 4.8 53.1 | Gemini 3.1 Pro 54.5. Loss of Control domain. Table 10 narrative adds: Muse Spark 1.1 resolves 7 out of 11 tasks on GDM Situational Awareness. Scorecard context: Table 1 is the report canonical capability snapshot (Claude column empty for chem/bio due to high refusal rates per table caption). Values extracted from the PDF text layer with pypdf layout mode - no OCR/vision involved.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: MCP Atlas / Scaled tool use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: MCP Atlas / Scaled tool use = 88.1%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mcp-atlas (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: 82.2 | Gemini 78.2 | Opus 4.8 82.2 | GPT 5.5 75.3。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: JobBench / Professional tool use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: JobBench / Professional tool use = 54.7%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: jobbench (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 17.0 | Gemini 15.9 | Opus 4.8 48.4 | GPT 5.5 38.3。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: Toolathlon-Verified / Personal tool use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Toolathlon-Verified / Personal tool use = 75.6%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: toolathlon-verified (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 49.4 | Gemini 61.1 | Opus 4.8 76.2(lead) | GPT 5.5 73.5。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: OSWorld-Verified / Agentic computer use · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: OSWorld-Verified / Agentic computer use = 80.8%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: osworld (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 53.3 | Gemini 76.2 | Opus 4.8 83.4(lead) | GPT 5.5 78.7。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: Humanity's Last Exam / Multidisciplinary reasoning (w/ tools) · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Humanity's Last Exam / Multidisciplinary reasoning (w/ tools) = 62.1%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hlehle (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 50.4 | Gemini 51.4 | Opus 4.8 57.9 | GPT 5.5 52.2。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: Finance Agent v2 / Agentic financial analysis(原图拼写 anaysis) · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Finance Agent v2 / Agentic financial analysis(原图拼写 anaysis) = 57.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: finance-agent-v2 (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 无数据(图中 -) | Gemini 43.0 | Opus 4.8 53.9 | GPT 5.5 51.8。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: Terminal-Bench 2.1 / Agentic terminal coding · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: Terminal-Bench 2.1 / Agentic terminal coding = 80.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: terminalbench (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 67.3 | Gemini 70.3 | Opus 4.8 82.7 | GPT 5.5 83.4(lead)。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。 跨 release 注意: muse-spark-1-2 发布图(2026-08-05)中同模型 Terminal-Bench 2.1 为 76.2(muse-code harness),与本表 80.0 差 3.8pp —— 两次 Meta 自报,发布时点/脚手架不同,比较前先对齐 harness。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: SWE-Bench Pro / Diverse software engineering · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: SWE-Bench Pro / Diverse software engineering = 61.5%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-pro (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 55.0 | Gemini 54.2 | Opus 4.8 69.2(lead) | GPT 5.5 58.6。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: DeepSWE 1.1 / Long-horizon agentic coding · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: DeepSWE 1.1 / Long-horizon agentic coding = 53.3%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepswe (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 10.0 | Gemini 12.0 | Opus 4.8 59.0 | GPT 5.5 67.0(lead)。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: CharXiv Reasoning / Chart QA · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: CharXiv Reasoning / Chart QA = 88.4%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: charxiv-reasoning (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 88.9 | Gemini 81.6 | Opus 4.8 89.9(lead) | GPT 5.5 84.8。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: BabyVision / Visual reasoning · figure: models/2026-07-09-muse-spark-1-1/images/02.png(主基准表,逐格目验) · quote_snippet: BabyVision / Visual reasoning = 76.3%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: babyvision (launch-post main table; 已注册候选或复用既有 id)。2026-09-01 二次审计补行: 该数值来自发布博客主基准表图,页面 DOM 的 img alt 误标为 "Inference-time compute scaling chart for Muse Image"(Meta 页面 alt 标注错误),归档图 models/2026-07-09-muse-spark-1-1/images/02.png 内容实为 Spark 1.1 主基准表,本次审计直接目验原图逐格转录。 对照值: Muse Spark 39.9 | Gemini 51.5 | Opus 4.8 81.2 | GPT 5.5 83.6(lead)。同表对照列: Gemini 3.1 Pro (high) / Google, Opus 4.8 (max) / Anthropic, GPT 5.5 (xhigh) / OpenAI, 及前代 Muse Spark / Meta。
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: DeepSearchQA (images/05.png) · figure: models/2026-07-09-muse-spark-1-1/images/05.png · quote_snippet: DeepSearchQA (images/05.png) = 84.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: deepsearchqa (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: GPT 5.5 87.8(lead) | Opus 4.8 84.3 | Muse Spark 76.8 | Gemini 71.3
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: Vibe Code Bench v1.1 (images/08.png) · figure: models/2026-07-09-muse-spark-1-1/images/08.png · quote_snippet: Vibe Code Bench v1.1 (images/08.png) = 72.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vibe-code-bench (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: Muse Spark 19.7(仅两代对比图)
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: SWE Atlas - Codebase QnA (images/08.png) · figure: models/2026-07-09-muse-spark-1-1/images/08.png · quote_snippet: SWE Atlas - Codebase QnA (images/08.png) = 42.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swe-atlas-codebase-qna (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: Muse Spark 24.2(仅两代对比图)
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluations (launch-post benchmark table/charts) · row: Meta Internal Coding Bench (images/09.png) · figure: models/2026-07-09-muse-spark-1-1/images/09.png · quote_snippet: Meta Internal Coding Bench (images/09.png) = 68.3%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: meta-internal-coding-bench (launch-post supplementary chart; 复用既有 id 或登记候选)。2026-09-01 二次审计补行,审计员直接目验归档图逐格转录(同页 alt 误标问题见主表行说明)。对照值: Opus 4.8 69.0(lead) | GPT 5.5 67.1 | Gemini 59.2 | Muse Spark 58.8