Muse Spark 1.2 / Muse Code
Meta / Llama · 2026-08-05 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Muse Spark 1.2
Meta 将 Muse Spark 1.2 定位为 Muse Spark 1.1 的编码向升级版,与 Muse Code 终端代理协同训练,经 Muse Code、Meta Model API、OpenRouter 提供。已收录 5 项评测集中于终端与软件工程 Agent:亮点 Terminal-Bench 2.1 82.9、DeepSWE 1.1 59.3,另含 GDPval-AA v2 与 MCP Atlas 第三方条目(未录数值)。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation (benchmark chart) · row: Terminal-Bench 2.1 · figure: images/01.webp(Terminal-Bench 2.1 柱状图:Opus 5 86.7 | Muse Spark 1.2 82.9 | GPT 5.6 Terra 81.8 | Grok 4.5 81.6 | Gemini 3.6 Flash 78.9 | Muse Spark 1.1 76.2) · quote_snippet: Terminal-Bench 2.1
{
"harness": "Muse Code (agent per model: Codex for GPT, Claude Code for Opus, Grok Build for Grok, Antigravity for Gemini, Kimi Code for Kimi, mini-swe-agent for Spark 1.1)",
"tools": null,
"shots": null,
"reasoning_effort": "xhigh (maximum available per model: high for Grok/Gemini, max for Opus/GPT/Kimi)",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "average task success rate (pass@1) across five attempts",
"judge": "task's executable verifier on final container state"
}视觉转写自归档图 images/01.webp(2026-09-01):Muse Spark 1.2 柱值如上,与同图对照模型逐格核对。原 verified 行补值。UPGRADED pending→verified in batch 6: benchmark named on the launch post and the full protocol is now machine-verified against the official Evaluation Methodology PDF (all 89 tasks of the official 2.1 release; each attempt in an isolated Daytona cloud sandbox; official datasets or faithful Harbor-format conversion; verifier grades final container state). The numeric value remains chart-image-only → score_status not_extracted per §12.5; widely-transcribed value 82.9% (second to Claude Opus 5 + Claude Code 86.7%) stays in this note pending a manual image read. Methodology also states Muse Spark 1.2 is served through the Meta Model API. [2026-09-01 audit: value re-verified against the live chart image.]
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation (benchmark chart) · row: DeepSWE 1.1 · figure: images/02.webp(DeepSWE 1.1 柱状图:Opus 5 65.0 | GPT 5.6 Terra 64.8 | Muse Spark 1.2 59.3 | Grok 4.5 56.6 | Muse Spark 1.1 53.0 | Gemini 3.6 Flash 40.0) · quote_snippet: DeepSWE 1.1
{
"harness": "Muse Code per-model agent products (NOT the leaderboard's uniform mini-swe-agent); internal agent evaluation framework; Harbor-format dataset; follows Pier v0.3.0 official runner as closely as possible",
"tools": null,
"shots": null,
"reasoning_effort": "xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 5,
"aggregation": "average task success rate (pass@1) across five attempts",
"judge": "handwritten functional verifier + regression checks in a fresh verifier container"
}视觉转写自归档图 images/02.webp(2026-09-01):Muse Spark 1.2 柱值如上,与同图对照模型逐格核对。原 verified 行补值。UPGRADED pending→verified in batch 6 via the methodology PDF: 113 tasks / 91 repositories / five languages (TS, Go, Python, JS, Rust); agent's final patch applied to a pristine checkout; external internet BLOCKED during rollout and grading (only the model endpoint reachable); a task passes only when functional verifier AND regression checks succeed. Methodology explicitly notes this is NOT harness-identical to the official leaderboard (leaderboard uses mini-swe-agent for every model). Chart-image value 59.3% (behind Opus 5 65.0 / GPT-5.6 Terra 64.8) stays in notes. [2026-09-01 audit: value re-verified against the live chart image.]
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Evaluation (benchmark chart) · row: Meta Internal Coding Bench · figure: images/03.webp(Meta Internal Coding Bench 柱状图:Opus 5 79.4 | Muse Spark 1.2 70.6 | Muse Spark 1.1 68.3 | GPT 5.6 Terra 65.4 | Gemini 3.6 Flash 63.9) · quote_snippet: Meta Internal Coding Bench
{
"harness": "separate internal agentic harness + dedicated grading containers",
"tools": null,
"shots": null,
"reasoning_effort": "xhigh",
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": 2,
"aggregation": "per-task success rate averaged over attempts, then averaged over tasks (percent resolved, pass@1)",
"judge": "unit tests in dedicated grading containers; internet disabled"
}视觉转写自归档图 images/03.webp(2026-09-01):Muse Spark 1.2 柱值如上,与同图对照模型逐格核对。原 verified 行补值。UPGRADED pending→verified in batch 6 via the methodology PDF: 440 tasks from real internal pull requests (bug fixes, features, refactors, cleanup); two attempts per task. new-benchmark id meta-internal-coding-bench (batch 5). Chart-image value 70.6% (vs Opus 5 79.4, Spark 1.1 68.3, GPT-5.6 Terra 65.4) stays in notes; Meta cautions the cross-model comparison mixes each model with its own agent product. [2026-09-01 audit: value re-verified against the live chart image.]
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Professional tasks / GDPVal-AA v2 · row: GDPVal-AA v2 · quote_snippet: GDPVal-AA v2 results come from Artificial Analysis.
{
"harness": "Artificial Analysis Stirrup agentic harness (results produced by Artificial Analysis, not Meta)",
"tools": [
"shell access",
"web browsing"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "Elo from blind pairwise LLM-judged comparisons of anonymized deliverables; human baseline anchored at 1000",
"judge": "LLM judge, head-to-head anonymized deliverable comparison"
}NEW row added in batch 6 from the methodology PDF: 220 real-world professional tasks from OpenAI's GDPval dataset, 44 occupations, nine US industries; deliverables include documents/spreadsheets/decks/diagrams/reports. attribution_type = third_party_reported because the PDF states the results come FROM Artificial Analysis — NOT vendor-run, does not count toward the vendor_reported public total. Not a CLI-agent comparison per the PDF. Numeric value not printed in the PDF (chart-only) → not_extracted. [2026-09-01 audit: value re-verified against the live chart image.]
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: MCP tool use / MCP Atlas · row: MCP Atlas · quote_snippet: We use the result produced by Scale AI with the benchmark's own agent harness and scoring pipeline.
{
"harness": "benchmark's own agent harness and scoring pipeline, result produced by Scale AI",
"tools": [
"containerized MCP servers with target and distractor tools"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "pass rate = percentage of tasks with claim-coverage >= 0.75 (LLM judge scores each ground-truth claim 1 / 0.5 / 0)",
"judge": "LLM judge per ground-truth claim"
}NEW row added in batch 6 from the methodology PDF: 1,000 human-authored tasks / 36 real MCP servers / 220 tools; public and private splits of 500 each. attribution_type = third_party_reported (result produced by Scale AI, benchmark owner side — not Meta-run). Numeric value not in the PDF → not_extracted. Split (public/private) for the Spark 1.2 figure not stated — Muse Glimmer's 75.5 was the Public split. [2026-09-01 audit: value re-verified against the live chart image.]
Muse Code
Muse Code 是 macOS/Linux 平台的终端编码代理(Beta),其 harness 与模型协同训练。已收录评测均以 Muse Spark 1.2 为被测模型,Muse Code 本身无独立评测数值。
- 输入模态
- 文本
- 上下文
- 官方资料未说明
- 参数
- 官方资料未说明
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
尚无对应评测记录。缺少证据不代表能力为零。