← 模型目录

Muse Glimmer 30B

Meta / Llama · 2026-08-10 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Muse Glimmer 30B

Muse Glimmer 30B 被 Meta 定位为可在消费级设备上运行的开源代理模型(Apache 2.0,由 Muse Spark 蒸馏,24-32GB 显存可跑)。评测覆盖通用代理、代理编码、多模态与安全:MCP Atlas 75.5%、SWE-bench Verified 76.0%、AIME 2026 94.7%;另提供 K-Quant 量化档与 DFlash 投机解码加速。

输入模态
文本 / 图像
上下文
官方资料未说明
参数
29.6B(dense)+ 1.8B ViT
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mcp-atlas 75.5% 模型 muse-glimmer-30b · 版本 Public · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / MCP Atlas (Public) · row: MCP Atlas (Public) · quote_snippet: MCP Atlas (Public) | 75.5 | 54.2 | 62.5

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id mcp-atlas; Public-subset variant recorded (Google 3.5 Flash's 83.6% does not state the subset — variant alignment needed before comparing). Sampling params are the card's recommended settings (temp 1.0 / top_p 0.95 / top_k 64), applied to all rows of this release. Competitors: Gemma4-31B thinking 54.2, Qwen3.6-27B thinking 62.5.

打开官方来源

deepsearchqa 74.6% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / DeepSearch QA · row: DeepSearch QA · quote_snippet: DeepSearch QA | 74.6 | 61.7 | 71.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id deepsearchqa (registered batch 2). Competitors: Gemma4 61.7, Qwen3.6 71.1.

打开官方来源

tau3-bench 23.5% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / τ3-Banking · row: τ3-Banking · quote_snippet: 𝛕3-Banking | 23.5 | 15.1 | 16.7

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: tau3-banking not yet in data/benchmarks/ (τ³-Bench Banking domain; card text names τ3-Bench). Distinct from existing tau-bench (retail/airline) — τ³ is a different generation, keep ids separate.

打开官方来源

wildclawbench 47.6% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / WildClawBench · row: WildClawBench · quote_snippet: WildClawBench | 47.6 | 37.6 | 43.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id wildclawbench (registered batch 3).

打开官方来源

gdpval-aa 953 Elo 模型 muse-glimmer-30b · 版本 v2 · 指标 knowledge_work_elo · 单位 elo 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / GDPVal-AA v2 · row: GDPVal-AA v2 · quote_snippet: GDPVal-AA v2 | 953 | 811 | 1141

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": "leaderboard-style Elo",
  "judge": null
}

Reuses candidate id gdpval-aa variant v2. Glimmer loses this row to Qwen3.6-27B (1141) — card bolds the competitor; useful bracket against Google 3.5 Flash-Lite's 1140.

打开官方来源

gaia2 43.3% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / Gaia2 · row: Gaia2 · quote_snippet: Gaia2 | 43.3 | 36.4 | 40.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: gaia2 not yet in data/benchmarks/ — successor generation to existing gaia; keep ids distinct.

打开官方来源

skillsbench 44.3% 模型 muse-glimmer-30b · 版本 with skills · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / SkillsBench (with skills) · row: SkillsBench (with skills) · quote_snippet: SkillsBench (with skills) | 44.3 | 32.4 | 46.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: skillsbench not yet in data/benchmarks/. Variant qualifier (with skills) is part of the row label.

打开官方来源

osworld 65.9% 模型 muse-glimmer-30b · 版本 Verified · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / OSWorld-Verified · row: OSWorld-Verified · quote_snippet: OSWorld-Verified | 65.9 | 58.5 | 75.6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark osworld variant Verified. Qwen3.6-27B wins (75.6).

打开官方来源

swebench-pro 51.2% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / SWE-Bench Pro · row: SWE-Bench Pro · quote_snippet: SWE-Bench Pro | 51.2 | 36.9 | 50.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id swebench-pro (registered batch 2).

打开官方来源

swebench 76.0% 模型 muse-glimmer-30b · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / SWE-Bench Verified · row: SWE-Bench Verified · quote_snippet: SWE-Bench Verified | 76.0 | 66.6 | 77.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark swebench (SWE-bench Verified). 30B open-weight model within 1.2pp of Qwen3.6-27B (77.2).

打开官方来源

terminalbench 51.7% 模型 muse-glimmer-30b · 版本 2.1, with terminus2 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / TerminalBench 2.1 (with terminus2) · row: TerminalBench 2.1 (with terminus2) · quote_snippet: TerminalBench 2.1 (with terminus2) | 51.7 | 43.4 | 60.7

{
  "harness": "terminus2",
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark terminalbench variant 2.1; harness terminus2 recorded (Terminus-2 harness also named on Google's 3.6 model page — harness-matched cross-vendor comparison possible here, unlike unnamed-harness rows).

打开官方来源

scicode 43.6% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / SciCode · row: SciCode · quote_snippet: SciCode | 43.6 | 43.4 | 39.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id scicode (registered batch 3).

打开官方来源

charxiv-reasoning 78.8% 模型 muse-glimmer-30b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / Charxiv Reasoning · row: Charxiv Reasoning · quote_snippet: Charxiv Reasoning | 78.8 | 77.7 | 78.4

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id charxiv-reasoning (registered batch 2). No-tools/with-tools split not present on this card (single figure).

打开官方来源

screenspot-pro 75.4% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / ScreenSpot Pro · row: ScreenSpot Pro · quote_snippet: ScreenSpot Pro | 75.4 | 75.9 | 76.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: screenspot-pro not yet in data/benchmarks/ (GUI grounding).

打开官方来源

omnidocbench 75.8% 模型 muse-glimmer-30b · 版本 v1.5 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / OmniDocBench v1.5 · row: OmniDocBench v1.5 · quote_snippet: OmniDocBench v1.5 | 75.8 | 72.5 | 77.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Reuses candidate id omnidocbench (registered batch 3), variant v1.5.

打开官方来源

mmmu 74% 模型 muse-glimmer-30b · 版本 Pro · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / MMMU Pro · row: MMMU Pro · quote_snippet: MMMU Pro | 74 | 73 | 75

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark mmmu variant Pro — comparable lane to Gemini 3 Flash's 81.2% MMMU Pro row.

打开官方来源

ifbench 77.0% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / IFBench · row: IFBench · quote_snippet: IFBench | 77.0 | 76.0 | 70.8

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ifbench not yet in data/benchmarks/ (instruction-following bench, IFBench; distinct from existing ifeval).

打开官方来源

aime-26 94.7% 模型 muse-glimmer-30b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / AIME 2026 · row: AIME 2026 · quote_snippet: AIME 2026 | 94.7 | 89.2 | 94.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aime-26 (AIME 2026 edition; id pattern aligned with aime-25). Sampling/aggregation not stated — do not assume pass@1.

打开官方来源

gpqa 83.5% 模型 muse-glimmer-30b · 版本 AA · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / GPQA Diamond (AA) · row: GPQA Diamond (AA) · quote_snippet: GPQA Diamond (AA) | 83.5 | 85.7 | 84.2

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark gpqa; variant 'AA' marks the Artificial-Analysis-run condition (card header 'GPQA Diamond (AA)') — a third-party-run lane, recorded as vendor_reported because the card publishes it, but noted as AA-run.

打开官方来源

hlehle 22.0% 模型 muse-glimmer-30b · 版本 Text (AA) · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / HLE Text (AA) · row: HLE Text (AA) · quote_snippet: HLE Text (AA) | 22.0 | 23.6 | 23.1

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Maps to existing benchmark hlehle, variant Text (AA) — text-only subset, third-party-run lane. Comparable only to other text-only HLE rows (e.g. Grok 4 Heavy 50.7% text-only, though different runner).

打开官方来源

aa-lcr 80.0 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / AA-LCR · row: AA-LCR · quote_snippet: AA-LCR | 80.0 | 68.3 | 73.3

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: aa-lcr not yet in data/benchmarks/ (Artificial Analysis Live Context Requests).

打开官方来源

beam128k 65.1% 模型 muse-glimmer-30b · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / Beam128K · row: Beam128K · quote_snippet: Beam128K | 65.1 | 58.2 | 63.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: beam128k not yet in data/benchmarks/ (long-context benchmark at 128k).

打开官方来源

mbct 41.5% 模型 muse-glimmer-30b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: MBCT · quote_snippet: MBCT | 41.5% | 50.6% | 45.9% | 58.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mbct - internal Meta chem/bio preparedness benchmark (no expansion printed on the card). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 50.6 | Qwen3.6-27B 45.9 | Kimi K3 58.9. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).

打开官方来源

hpct 52.3% 模型 muse-glimmer-30b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: HPCT · quote_snippet: HPCT | 52.3% | 54.0% | 48.7% | 59.6%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hpct - internal Meta chem/bio preparedness benchmark (no expansion printed on the card). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 54.0 | Qwen3.6-27B 48.7 | Kimi K3 59.6. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).

打开官方来源

vct 37.0% 模型 muse-glimmer-30b · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: VCT · quote_snippet: VCT | 37.0% | 43.5% | 33.7% | 48.0%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: vct - internal Meta chem/bio preparedness benchmark (no expansion printed on the card). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 43.5 | Qwen3.6-27B 33.7 | Kimi K3 48.0. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).

打开官方来源

wmdp 86.5% 模型 muse-glimmer-30b · 版本 Bio · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: WMDP (Bio) · quote_snippet: WMDP (Bio) | 86.5% | 85.9% | 84.8% | 89.1%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: wmdp - public Weapons-of-Mass-Destruction-Proxy benchmark, Bio subset. Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 85.9 | Qwen3.6-27B 84.8 | Kimi K3 89.1. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).

打开官方来源

wmdp 75.2% 模型 muse-glimmer-30b · 版本 Chem · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: WMDP (Chem) · quote_snippet: WMDP (Chem) | 75.2% | 80.5% | 74.8% | 84.2%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: wmdp - public Weapons-of-Mass-Destruction-Proxy benchmark, Chem subset (same id as the Bio row, variant splits the subset). Competitors: Gemma4-31B 80.5 | Qwen3.6-27B 74.8 | Kimi K3 84.2. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).

打开官方来源

lab-bench 80.2% 模型 muse-glimmer-30b · 版本 ProtocolQA · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: Lab Bench (ProtocolQA) · quote_snippet: Lab Bench (ProtocolQA) | 80.2% | 75.8% | 69.1% | 81.9%

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
  "temperature": 1,
  "top_p": 0.95,
  "top_k": 64,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: lab-bench - public LAION LabBench, ProtocolQA subset (wet-lab protocol debugging). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 75.8 | Qwen3.6-27B 69.1 | Kimi K3 81.9. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).

打开官方来源