Muse Glimmer 30B
Meta / Llama · 2026-08-10 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Muse Glimmer 30B
Muse Glimmer 30B 被 Meta 定位为可在消费级设备上运行的开源代理模型(Apache 2.0,由 Muse Spark 蒸馏,24-32GB 显存可跑)。评测覆盖通用代理、代理编码、多模态与安全:MCP Atlas 75.5%、SWE-bench Verified 76.0%、AIME 2026 94.7%;另提供 K-Quant 量化档与 DFlash 投机解码加速。
- 输入模态
- 文本 / 图像
- 上下文
- 官方资料未说明
- 参数
- 29.6B(dense)+ 1.8B ViT
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / MCP Atlas (Public) · row: MCP Atlas (Public) · quote_snippet: MCP Atlas (Public) | 75.5 | 54.2 | 62.5
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id mcp-atlas; Public-subset variant recorded (Google 3.5 Flash's 83.6% does not state the subset — variant alignment needed before comparing). Sampling params are the card's recommended settings (temp 1.0 / top_p 0.95 / top_k 64), applied to all rows of this release. Competitors: Gemma4-31B thinking 54.2, Qwen3.6-27B thinking 62.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / DeepSearch QA · row: DeepSearch QA · quote_snippet: DeepSearch QA | 74.6 | 61.7 | 71.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id deepsearchqa (registered batch 2). Competitors: Gemma4 61.7, Qwen3.6 71.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / τ3-Banking · row: τ3-Banking · quote_snippet: 𝛕3-Banking | 23.5 | 15.1 | 16.7
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: tau3-banking not yet in data/benchmarks/ (τ³-Bench Banking domain; card text names τ3-Bench). Distinct from existing tau-bench (retail/airline) — τ³ is a different generation, keep ids separate.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / WildClawBench · row: WildClawBench · quote_snippet: WildClawBench | 47.6 | 37.6 | 43.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id wildclawbench (registered batch 3).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / GDPVal-AA v2 · row: GDPVal-AA v2 · quote_snippet: GDPVal-AA v2 | 953 | 811 | 1141
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": "leaderboard-style Elo",
"judge": null
}Reuses candidate id gdpval-aa variant v2. Glimmer loses this row to Qwen3.6-27B (1141) — card bolds the competitor; useful bracket against Google 3.5 Flash-Lite's 1140.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / Gaia2 · row: Gaia2 · quote_snippet: Gaia2 | 43.3 | 36.4 | 40.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: gaia2 not yet in data/benchmarks/ — successor generation to existing gaia; keep ids distinct.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / SkillsBench (with skills) · row: SkillsBench (with skills) · quote_snippet: SkillsBench (with skills) | 44.3 | 32.4 | 46.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: skillsbench not yet in data/benchmarks/. Variant qualifier (with skills) is part of the row label.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Agentic / OSWorld-Verified · row: OSWorld-Verified · quote_snippet: OSWorld-Verified | 65.9 | 58.5 | 75.6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark osworld variant Verified. Qwen3.6-27B wins (75.6).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / SWE-Bench Pro · row: SWE-Bench Pro · quote_snippet: SWE-Bench Pro | 51.2 | 36.9 | 50.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id swebench-pro (registered batch 2).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / SWE-Bench Verified · row: SWE-Bench Verified · quote_snippet: SWE-Bench Verified | 76.0 | 66.6 | 77.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark swebench (SWE-bench Verified). 30B open-weight model within 1.2pp of Qwen3.6-27B (77.2).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / TerminalBench 2.1 (with terminus2) · row: TerminalBench 2.1 (with terminus2) · quote_snippet: TerminalBench 2.1 (with terminus2) | 51.7 | 43.4 | 60.7
{
"harness": "terminus2",
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark terminalbench variant 2.1; harness terminus2 recorded (Terminus-2 harness also named on Google's 3.6 model page — harness-matched cross-vendor comparison possible here, unlike unnamed-harness rows).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Agentic Coding / SciCode · row: SciCode · quote_snippet: SciCode | 43.6 | 43.4 | 39.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id scicode (registered batch 3).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / Charxiv Reasoning · row: Charxiv Reasoning · quote_snippet: Charxiv Reasoning | 78.8 | 77.7 | 78.4
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id charxiv-reasoning (registered batch 2). No-tools/with-tools split not present on this card (single figure).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / ScreenSpot Pro · row: ScreenSpot Pro · quote_snippet: ScreenSpot Pro | 75.4 | 75.9 | 76.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: screenspot-pro not yet in data/benchmarks/ (GUI grounding).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / OmniDocBench v1.5 · row: OmniDocBench v1.5 · quote_snippet: OmniDocBench v1.5 | 75.8 | 72.5 | 77.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Reuses candidate id omnidocbench (registered batch 3), variant v1.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), Multimodal / MMMU Pro · row: MMMU Pro · quote_snippet: MMMU Pro | 74 | 73 | 75
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark mmmu variant Pro — comparable lane to Gemini 3 Flash's 81.2% MMMU Pro row.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / IFBench · row: IFBench · quote_snippet: IFBench | 77.0 | 76.0 | 70.8
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ifbench not yet in data/benchmarks/ (instruction-following bench, IFBench; distinct from existing ifeval).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / AIME 2026 · row: AIME 2026 · quote_snippet: AIME 2026 | 94.7 | 89.2 | 94.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aime-26 (AIME 2026 edition; id pattern aligned with aime-25). Sampling/aggregation not stated — do not assume pass@1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / GPQA Diamond (AA) · row: GPQA Diamond (AA) · quote_snippet: GPQA Diamond (AA) | 83.5 | 85.7 | 84.2
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark gpqa; variant 'AA' marks the Artificial-Analysis-run condition (card header 'GPQA Diamond (AA)') — a third-party-run lane, recorded as vendor_reported because the card publishes it, but noted as AA-run.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / HLE Text (AA) · row: HLE Text (AA) · quote_snippet: HLE Text (AA) | 22.0 | 23.6 | 23.1
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Maps to existing benchmark hlehle, variant Text (AA) — text-only subset, third-party-run lane. Comparable only to other text-only HLE rows (e.g. Grok 4 Heavy 50.7% text-only, though different runner).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / AA-LCR · row: AA-LCR · quote_snippet: AA-LCR | 80.0 | 68.3 | 73.3
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: aa-lcr not yet in data/benchmarks/ (Artificial Analysis Live Context Requests).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Model card benchmark table (DOM), General Capabilities and Reasoning / Beam128K · row: Beam128K · quote_snippet: Beam128K | 65.1 | 58.2 | 63.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: beam128k not yet in data/benchmarks/ (long-context benchmark at 128k).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: MBCT · quote_snippet: MBCT | 41.5% | 50.6% | 45.9% | 58.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mbct - internal Meta chem/bio preparedness benchmark (no expansion printed on the card). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 50.6 | Qwen3.6-27B 45.9 | Kimi K3 58.9. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: HPCT · quote_snippet: HPCT | 52.3% | 54.0% | 48.7% | 59.6%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hpct - internal Meta chem/bio preparedness benchmark (no expansion printed on the card). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 54.0 | Qwen3.6-27B 48.7 | Kimi K3 59.6. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: VCT · quote_snippet: VCT | 37.0% | 43.5% | 33.7% | 48.0%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vct - internal Meta chem/bio preparedness benchmark (no expansion printed on the card). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 43.5 | Qwen3.6-27B 33.7 | Kimi K3 48.0. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: WMDP (Bio) · quote_snippet: WMDP (Bio) | 86.5% | 85.9% | 84.8% | 89.1%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: wmdp - public Weapons-of-Mass-Destruction-Proxy benchmark, Bio subset. Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 85.9 | Qwen3.6-27B 84.8 | Kimi K3 89.1. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: WMDP (Chem) · quote_snippet: WMDP (Chem) | 75.2% | 80.5% | 74.8% | 84.2%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: wmdp - public Weapons-of-Mass-Destruction-Proxy benchmark, Chem subset (same id as the Bio row, variant splits the subset). Competitors: Gemma4-31B 80.5 | Qwen3.6-27B 74.8 | Kimi K3 84.2. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Trust and Safety (chem/bio preparedness) · table: Model card chem/bio preparedness table (DOM) · row: Lab Bench (ProtocolQA) · quote_snippet: Lab Bench (ProtocolQA) | 80.2% | 75.8% | 69.1% | 81.9%
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": "high (column header: Muse Glimmer-30B High Reasoning)",
"temperature": 1,
"top_p": 0.95,
"top_k": 64,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: lab-bench - public LAION LabBench, ProtocolQA subset (wet-lab protocol debugging). Converted from release notes to its own row in the 2026-09-01 audit. Competitors: Gemma4-31B 75.8 | Qwen3.6-27B 69.1 | Kimi K3 81.9. Table context: card section "In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging" feeding the Chem/Bio: Moderate-or-lower risk designation. Kimi K3 included by Meta "for context" (larger open-weight model, outside the size class).