← 模型目录

Muse Spark 1.3

Meta / Llama · 2026-09-02 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Muse Spark 1.3

Meta 专为长程软件工程与智能体工具循环优化的前沿多模态大模型,相较 1.2 版本减少 20% 工具调用与 25% 输出 token;在 DeepSWE v1.1 达到 75.4%、Terminal-Bench 2.1 88.8%、SWE-Atlas Codebase QnA 59.4% 以及 1M 上下文 MRCR 98.1%,支持 max 与 xhigh 双推理档位。

输入模态
文本 / 代码 / 多模态
上下文
1M
参数
官方资料未说明
价格(每百万 tokens)
USD 输入 1.25 / 输出 4.25 · Meta Model API / OpenRouter 定价:输入 $1.25 / 输出 $4.25 每百万 tokens

本变体的评测证据

deepswe 75.4% 模型 muse-spark-1-3 · 版本 1.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance · table: Software Engineering Benchmarks · row: DeepSWE v1.1

{
  "harness": "Muse Code agentic harness",
  "tools": [
    "bash",
    "editor"
  ],
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Meta 官方博文发布,DeepSWE v1.1 报告 75.4%,使用 Muse Code 自研 agentic harness,在 max reasoning 设置下运行。

打开官方来源

terminalbench 88.8% 模型 muse-spark-1-3 · 版本 2.1 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance · table: Software Engineering Benchmarks · row: Terminal-Bench 2.1

{
  "harness": "Terminus-2",
  "tools": [
    "bash"
  ],
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Meta 官方博文发布,Terminal-Bench 2.1 报告 88.8%(Terminus-2 harness,max reasoning)。

打开官方来源

swe-atlas-codebase-qna 59.4% 模型 muse-spark-1-3 · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance · table: Software Engineering Benchmarks · row: SWE-Atlas Codebase QnA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Meta 官方博文发布,SWE-Atlas Codebase QnA 报告 59.4%,检验大型代码仓深度语义理解。

打开官方来源

mrcr 98.1% 模型 muse-spark-1-3 · 版本 v2 1M · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-11 · 距快照 23 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmark Performance · table: Software Engineering Benchmarks · row: MRCR v2 (512K-1M)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "max",
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Meta 官方博文披露,在 512K-1M 超长上下文检索召回率(MRCR v2)保持 98.1% 准确率。

打开官方来源