Qwen3.8-Flash-Next / Qwen3.8-Flash-Next-Base
Alibaba / Qwen · 2026-08-26 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Qwen3.8-Flash-Next
Qwen3.8-Flash-Next 被发布文定位为迈向极致性价比的全新架构预览(GDN+QSA 混合注意力、N-gram Embedding、Muon 优化器),为 Qwen4 架构探路。评测覆盖代理编码/办公、多模态代理与通用推理:SWE-bench Pro 62.5、CoWorkBench 73.9、Toolathlon Verified 73.5、AndroidWorld 84.5。
- 输入模态
- 文本 / 图像 / 视频
- 上下文
- 256K(YaRN 可扩至 1M)
- 参数
- 125B-A6B(另含 51B N-gram embedding)
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: DeepSWE 1.1 (agentic coding)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B 42.2, Qwen3.7-Plus 16.5, DS-V4-Flash-0731 54.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SWE-bench Pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B 61.7, Qwen3.7-Plus 55.8, DS-V4-Flash-0731 56.0, Opus-4.6 53.4.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SWE-bench Multilingual
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-multilingual. Competitor cells: Qwen3.8-27B 73.8, Qwen3.7-Plus 75.8, Opus-4.6 77.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: NL2Repo-Bench (repo-level codegen)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B 42.3, Qwen3.7-Plus 41.1, DS-V4-Flash-0731 54.2, Opus-4.6 47.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: CoWorkBench (long-horizon office)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: coworkbench. Competitor cells: Qwen3.8-27B 70.7, Qwen3.7-Plus 65.1, DS-V4-Flash-0731 45.1, Opus-4.6 68.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: JobBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: jobbench. Competitor cells: Qwen3.8-27B 33.4, Qwen3.7-Plus 27.6, DS-V4-Flash-0731 41.3, Opus-4.6 36.6.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: Agents' Last Exam
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Dual metric cell: Pass@1 24.3 AND Score 51.2. Competitor cells: Qwen3.8-27B 20.4/42.9, Qwen3.7-Plus 13.2/33.6, DS-V4-Flash-0731 25.2/-.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: Toolathlon Verified (Pass@1)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B 67.1, Qwen3.7-Plus 50.6, DS-V4-Flash-0731 70.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: IFBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: ifbench. Competitor cells: Qwen3.8-27B 79.5, Qwen3.7-Plus 79.1, DS-V4-Flash-0731 79.2, Opus-4.6 62.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: GPQA Diamond
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B 89.2, Qwen3.7-Plus 90.3, DS-V4-Flash-0731 90.8, Opus-4.6 91.3.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: HLE
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B 30.8, Qwen3.7-Plus 34.7, DS-V4-Flash-0731 33.8, Opus-4.6 40.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: LiveCodeBench v6
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B 90.3, Qwen3.7-Plus 89.6, DS-V4-Flash-0731 90.6, Opus-4.6 88.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: ClawEval-MM (multimodal tool use)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: claweval-mm (Pass@3 64.4 AND Average 60.4). Competitor cells: Qwen3.8-27B 57.4/56.9, Qwen3.7-Plus 57.4/60.1, Opus-4.6 52.5/54.7.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: RecreationBench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: recreationbench already introduced by prior batches. Competitor cells: Qwen3.8-27B 47.1, Qwen3.7-Plus 30.2.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: AndroidWorld
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}androidworld id exists in data/benchmarks/. Competitor cells: Qwen3.8-27B 81.9, Qwen3.7-Plus 81.0, Opus-4.6 62.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: OSWorld 2.0
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Dual metric: Binary 19.4 AND Partial 52.3. Competitor cells: Qwen3.8-27B 19.4/48.0, Qwen3.7-Plus 2.8/21.5.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: Vision2Web
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: vision2web not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B 62.9, Qwen3.7-Plus 42.1.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: ERQA (embodied)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: erqa already introduced by prior batches. Competitor cells: Qwen3.8-27B 65.5, Qwen3.7-Plus 69.8, Opus-4.6 40.8.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: LVBench (long video)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: lvbench already introduced by prior batches. Competitor cells: Qwen3.8-27B 72.4, Qwen3.7-Plus 76.2, Opus-4.6 63.0.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: RealWorldQA
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: realworldqa already introduced by prior batches. Competitor cells: Qwen3.8-27B 85.9, Qwen3.7-Plus 86.9, Opus-4.6 73.9.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MathVision
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Dual condition: Without code-interpreter 90.6 AND With CI 95.7. Competitor cells: Qwen3.8-27B 90.0/94.6, Qwen3.7-Plus 90.3/88.7, Opus-4.6 65.5/-.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: CharXiv (RQ)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}charxiv-reasoning id already introduced by prior batches. Competitor cells: Qwen3.8-27B 83.7/90.2, Qwen3.7-Plus 85.8/85.9, Opus-4.6 66.0/-.
Qwen3.8-Flash-Next-Base
Qwen3.8-Flash-Next-Base 是同架构的预训练基座版本(未做后训练),发布文表 3 单独报告其通用、数学 STEM、编码与多语言基线:MMLU 90.36、BBH 90.87、GSM8K 93.29。
- 输入模态
- 官方资料未说明
- 上下文
- 256K(YaRN 可扩至 1M)
- 参数
- 125B-A6B(另含 51B N-gram embedding)
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMLU (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Base-model table (pre-post-training). Competitor cells: Qwen3.8-27B-Base 87.51, Qwen3.7-Plus-Base 90.43.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMLU-Redux (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B-Base 87.26, Qwen3.7-Plus-Base 91.47.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMLU-Pro (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B-Base 68.60, Qwen3.7-Plus-Base 70.90.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SuperGPQA (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: supergpqa. Competitor cells: Qwen3.8-27B-Base 44.86, Qwen3.7-Plus-Base 48.42.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: BBH (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: bbh (BIG-Bench Hard) not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 89.56, Qwen3.7-Plus-Base 89.41.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: GPQA (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Base GPQA (main set, not Diamond). Competitor cells: Qwen3.8-27B-Base 45.01, Qwen3.7-Plus-Base 51.52.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: GSM8K (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Competitor cells: Qwen3.8-27B-Base 93.18, Qwen3.7-Plus-Base 92.95.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MATH (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: math (also introduced in qwen2-5.json this batch). Competitor cells: Qwen3.8-27B-Base 60.54, Qwen3.7-Plus-Base 74.38.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: EvalPlus (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: evalplus (also in qwen2-5.json). Competitor cells: Qwen3.8-27B-Base 76.05, Qwen3.7-Plus-Base 78.06.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MultiPL-E (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: multipl-e. Competitor cells: Qwen3.8-27B-Base 74.50, Qwen3.7-Plus-Base 81.68.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SWEBench-Pretrain (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swebench-pretrain not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 41.66, Qwen3.7-Plus-Base 49.24.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MGSM (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}mgsm id exists in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 86.37, Qwen3.7-Plus-Base 85.42.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMMLU (Base, multilingual)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: mmmlu (multilingual MMLU) not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 79.74, Qwen3.7-Plus-Base 84.53.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: INCLUDE (Base)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: include (multilingual) not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 74.37, Qwen3.7-Plus-Base 78.90.