← 模型目录

Muse Spark 1.2 / Muse Code

Meta / Llama · 2026-08-05 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Muse Spark 1.2

Meta 将 Muse Spark 1.2 定位为 Muse Spark 1.1 的编码向升级版,与 Muse Code 终端代理协同训练,经 Muse Code、Meta Model API、OpenRouter 提供。已收录 5 项评测集中于终端与软件工程 Agent:亮点 Terminal-Bench 2.1 82.9、DeepSWE 1.1 59.3,另含 GDPval-AA v2 与 MCP Atlas 第三方条目(未录数值)。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

terminalbench 82.9 模型 muse-spark-1-2 · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation (benchmark chart) · row: Terminal-Bench 2.1 · figure: images/01.webp(Terminal-Bench 2.1 柱状图:Opus 5 86.7 | Muse Spark 1.2 82.9 | GPT 5.6 Terra 81.8 | Grok 4.5 81.6 | Gemini 3.6 Flash 78.9 | Muse Spark 1.1 76.2) · quote_snippet: Terminal-Bench 2.1

{
  "harness": "Muse Code (agent per model: Codex for GPT, Claude Code for Opus, Grok Build for Grok, Antigravity for Gemini, Kimi Code for Kimi, mini-swe-agent for Spark 1.1)",
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh (maximum available per model: high for Grok/Gemini, max for Opus/GPT/Kimi)",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "average task success rate (pass@1) across five attempts",
  "judge": "task's executable verifier on final container state"
}

视觉转写自归档图 images/01.webp(2026-09-01):Muse Spark 1.2 柱值如上,与同图对照模型逐格核对。原 verified 行补值。UPGRADED pending→verified in batch 6: benchmark named on the launch post and the full protocol is now machine-verified against the official Evaluation Methodology PDF (all 89 tasks of the official 2.1 release; each attempt in an isolated Daytona cloud sandbox; official datasets or faithful Harbor-format conversion; verifier grades final container state). The numeric value remains chart-image-only → score_status not_extracted per §12.5; widely-transcribed value 82.9% (second to Claude Opus 5 + Claude Code 86.7%) stays in this note pending a manual image read. Methodology also states Muse Spark 1.2 is served through the Meta Model API. [2026-09-01 audit: value re-verified against the live chart image.]

打开官方来源

deepswe 59.3 模型 muse-spark-1-2 · 版本 1.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation (benchmark chart) · row: DeepSWE 1.1 · figure: images/02.webp(DeepSWE 1.1 柱状图:Opus 5 65.0 | GPT 5.6 Terra 64.8 | Muse Spark 1.2 59.3 | Grok 4.5 56.6 | Muse Spark 1.1 53.0 | Gemini 3.6 Flash 40.0) · quote_snippet: DeepSWE 1.1

{
  "harness": "Muse Code per-model agent products (NOT the leaderboard's uniform mini-swe-agent); internal agent evaluation framework; Harbor-format dataset; follows Pier v0.3.0 official runner as closely as possible",
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 5,
  "aggregation": "average task success rate (pass@1) across five attempts",
  "judge": "handwritten functional verifier + regression checks in a fresh verifier container"
}

视觉转写自归档图 images/02.webp(2026-09-01):Muse Spark 1.2 柱值如上,与同图对照模型逐格核对。原 verified 行补值。UPGRADED pending→verified in batch 6 via the methodology PDF: 113 tasks / 91 repositories / five languages (TS, Go, Python, JS, Rust); agent's final patch applied to a pristine checkout; external internet BLOCKED during rollout and grading (only the model endpoint reachable); a task passes only when functional verifier AND regression checks succeed. Methodology explicitly notes this is NOT harness-identical to the official leaderboard (leaderboard uses mini-swe-agent for every model). Chart-image value 59.3% (behind Opus 5 65.0 / GPT-5.6 Terra 64.8) stays in notes. [2026-09-01 audit: value re-verified against the live chart image.]

打开官方来源

meta-internal-coding-bench 70.6 模型 muse-spark-1-2 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Evaluation (benchmark chart) · row: Meta Internal Coding Bench · figure: images/03.webp(Meta Internal Coding Bench 柱状图:Opus 5 79.4 | Muse Spark 1.2 70.6 | Muse Spark 1.1 68.3 | GPT 5.6 Terra 65.4 | Gemini 3.6 Flash 63.9) · quote_snippet: Meta Internal Coding Bench

{
  "harness": "separate internal agentic harness + dedicated grading containers",
  "tools": null,
  "shots": null,
  "reasoning_effort": "xhigh",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": 2,
  "aggregation": "per-task success rate averaged over attempts, then averaged over tasks (percent resolved, pass@1)",
  "judge": "unit tests in dedicated grading containers; internet disabled"
}

视觉转写自归档图 images/03.webp(2026-09-01):Muse Spark 1.2 柱值如上,与同图对照模型逐格核对。原 verified 行补值。UPGRADED pending→verified in batch 6 via the methodology PDF: 440 tasks from real internal pull requests (bug fixes, features, refactors, cleanup); two attempts per task. new-benchmark id meta-internal-coding-bench (batch 5). Chart-image value 70.6% (vs Opus 5 79.4, Spark 1.1 68.3, GPT-5.6 Terra 65.4) stays in notes; Meta cautions the cross-model comparison mixes each model with its own agent product. [2026-09-01 audit: value re-verified against the live chart image.]

打开官方来源

gdpval-aa 该来源尚无可读数值 模型 muse-spark-1-2 · 版本 v2 · 指标 knowledge_work_elo · 单位 elo 来源等级 A · third_party_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Professional tasks / GDPVal-AA v2 · row: GDPVal-AA v2 · quote_snippet: GDPVal-AA v2 results come from Artificial Analysis.

{
  "harness": "Artificial Analysis Stirrup agentic harness (results produced by Artificial Analysis, not Meta)",
  "tools": [
    "shell access",
    "web browsing"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "Elo from blind pairwise LLM-judged comparisons of anonymized deliverables; human baseline anchored at 1000",
  "judge": "LLM judge, head-to-head anonymized deliverable comparison"
}

NEW row added in batch 6 from the methodology PDF: 220 real-world professional tasks from OpenAI's GDPval dataset, 44 occupations, nine US industries; deliverables include documents/spreadsheets/decks/diagrams/reports. attribution_type = third_party_reported because the PDF states the results come FROM Artificial Analysis — NOT vendor-run, does not count toward the vendor_reported public total. Not a CLI-agent comparison per the PDF. Numeric value not printed in the PDF (chart-only) → not_extracted. [2026-09-01 audit: value re-verified against the live chart image.]

打开官方来源

mcp-atlas 该来源尚无可读数值 模型 muse-spark-1-2 · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · third_party_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: MCP tool use / MCP Atlas · row: MCP Atlas · quote_snippet: We use the result produced by Scale AI with the benchmark's own agent harness and scoring pipeline.

{
  "harness": "benchmark's own agent harness and scoring pipeline, result produced by Scale AI",
  "tools": [
    "containerized MCP servers with target and distractor tools"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "pass rate = percentage of tasks with claim-coverage >= 0.75 (LLM judge scores each ground-truth claim 1 / 0.5 / 0)",
  "judge": "LLM judge per ground-truth claim"
}

NEW row added in batch 6 from the methodology PDF: 1,000 human-authored tasks / 36 real MCP servers / 220 tools; public and private splits of 500 each. attribution_type = third_party_reported (result produced by Scale AI, benchmark owner side — not Meta-run). Numeric value not in the PDF → not_extracted. Split (public/private) for the Spark 1.2 figure not stated — Muse Glimmer's 75.5 was the Public split. [2026-09-01 audit: value re-verified against the live chart image.]

打开官方来源

Muse Code

Muse Code 是 macOS/Linux 平台的终端编码代理(Beta),其 harness 与模型协同训练。已收录评测均以 Muse Spark 1.2 为被测模型,Muse Code 本身无独立评测数值。

输入模态
文本
上下文
官方资料未说明
参数
官方资料未说明
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

尚无对应评测记录。缺少证据不代表能力为零。