← 模型目录

Qwen3.8-Flash-Next / Qwen3.8-Flash-Next-Base

Alibaba / Qwen · 2026-08-26 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next 被发布文定位为迈向极致性价比的全新架构预览(GDN+QSA 混合注意力、N-gram Embedding、Muon 优化器),为 Qwen4 架构探路。评测覆盖代理编码/办公、多模态代理与通用推理:SWE-bench Pro 62.5、CoWorkBench 73.9、Toolathlon Verified 73.5、AndroidWorld 84.5。

输入模态
文本 / 图像 / 视频
上下文
256K(YaRN 可扩至 1M)
参数
125B-A6B(另含 51B N-gram embedding)
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

deepswe 58.7 模型 qwen3-8-flash-next · 版本 1.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: DeepSWE 1.1 (agentic coding)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B 42.2, Qwen3.7-Plus 16.5, DS-V4-Flash-0731 54.4.

打开官方来源

swebench-pro 62.5 模型 qwen3-8-flash-next · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SWE-bench Pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B 61.7, Qwen3.7-Plus 55.8, DS-V4-Flash-0731 56.0, Opus-4.6 53.4.

打开官方来源

swebench-multilingual 81 模型 qwen3-8-flash-next · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SWE-bench Multilingual

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-multilingual. Competitor cells: Qwen3.8-27B 73.8, Qwen3.7-Plus 75.8, Opus-4.6 77.5.

打开官方来源

nl2repo 48.1 模型 qwen3-8-flash-next · 版本 未说明 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: NL2Repo-Bench (repo-level codegen)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B 42.3, Qwen3.7-Plus 41.1, DS-V4-Flash-0731 54.2, Opus-4.6 47.6.

打开官方来源

coworkbench 73.9 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: CoWorkBench (long-horizon office)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: coworkbench. Competitor cells: Qwen3.8-27B 70.7, Qwen3.7-Plus 65.1, DS-V4-Flash-0731 45.1, Opus-4.6 68.2.

打开官方来源

jobbench 55.7 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: JobBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: jobbench. Competitor cells: Qwen3.8-27B 33.4, Qwen3.7-Plus 27.6, DS-V4-Flash-0731 41.3, Opus-4.6 36.6.

打开官方来源

agents-last-exam 24.3 / 51.2 模型 qwen3-8-flash-next · 版本 Pass@1 / Score dual · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: Agents' Last Exam

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Dual metric cell: Pass@1 24.3 AND Score 51.2. Competitor cells: Qwen3.8-27B 20.4/42.9, Qwen3.7-Plus 13.2/33.6, DS-V4-Flash-0731 25.2/-.

打开官方来源

toolathlon 73.5 模型 qwen3-8-flash-next · 版本 Verified, Pass@1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: Toolathlon Verified (Pass@1)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B 67.1, Qwen3.7-Plus 50.6, DS-V4-Flash-0731 70.3.

打开官方来源

ifbench 81.3 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: IFBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: ifbench. Competitor cells: Qwen3.8-27B 79.5, Qwen3.7-Plus 79.1, DS-V4-Flash-0731 79.2, Opus-4.6 62.5.

打开官方来源

gpqa 91.7 模型 qwen3-8-flash-next · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: GPQA Diamond

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B 89.2, Qwen3.7-Plus 90.3, DS-V4-Flash-0731 90.8, Opus-4.6 91.3.

打开官方来源

hlehle 35.9 模型 qwen3-8-flash-next · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: HLE

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B 30.8, Qwen3.7-Plus 34.7, DS-V4-Flash-0731 33.8, Opus-4.6 40.0.

打开官方来源

lcb 91.9 模型 qwen3-8-flash-next · 版本 v6 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: LiveCodeBench v6

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B 90.3, Qwen3.7-Plus 89.6, DS-V4-Flash-0731 90.6, Opus-4.6 88.8.

打开官方来源

claw-eval 64.4 / 60.4 模型 qwen3-8-flash-next · 版本 Pass@3 / Average dual · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: ClawEval-MM (multimodal tool use)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: claweval-mm (Pass@3 64.4 AND Average 60.4). Competitor cells: Qwen3.8-27B 57.4/56.9, Qwen3.7-Plus 57.4/60.1, Opus-4.6 52.5/54.7.

打开官方来源

recreationbench 49.9 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: RecreationBench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: recreationbench already introduced by prior batches. Competitor cells: Qwen3.8-27B 47.1, Qwen3.7-Plus 30.2.

打开官方来源

androidworld 84.5 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: AndroidWorld

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

androidworld id exists in data/benchmarks/. Competitor cells: Qwen3.8-27B 81.9, Qwen3.7-Plus 81.0, Opus-4.6 62.0.

打开官方来源

osworld 19.4 / 52.3 模型 qwen3-8-flash-next · 版本 2.0, Binary/Partial dual · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: OSWorld 2.0

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Dual metric: Binary 19.4 AND Partial 52.3. Competitor cells: Qwen3.8-27B 19.4/48.0, Qwen3.7-Plus 2.8/21.5.

打开官方来源

vision2web 64 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: Vision2Web

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: vision2web not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B 62.9, Qwen3.7-Plus 42.1.

打开官方来源

erqa 72.3 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: ERQA (embodied)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: erqa already introduced by prior batches. Competitor cells: Qwen3.8-27B 65.5, Qwen3.7-Plus 69.8, Opus-4.6 40.8.

打开官方来源

lvbench 76.6 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: LVBench (long video)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: lvbench already introduced by prior batches. Competitor cells: Qwen3.8-27B 72.4, Qwen3.7-Plus 76.2, Opus-4.6 63.0.

打开官方来源

realworldqa 88.5 模型 qwen3-8-flash-next · 版本 未说明 · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: RealWorldQA

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: realworldqa already introduced by prior batches. Competitor cells: Qwen3.8-27B 85.9, Qwen3.7-Plus 86.9, Opus-4.6 73.9.

打开官方来源

mathvision 90.6 / 95.7 模型 qwen3-8-flash-next · 版本 Without CI / With CI dual · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MathVision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Dual condition: Without code-interpreter 90.6 AND With CI 95.7. Competitor cells: Qwen3.8-27B 90.0/94.6, Qwen3.7-Plus 90.3/88.7, Opus-4.6 65.5/-.

打开官方来源

charxiv-reasoning 84.6 / 90.6 模型 qwen3-8-flash-next · 版本 RQ, Without CI / With CI dual · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: CharXiv (RQ)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

charxiv-reasoning id already introduced by prior batches. Competitor cells: Qwen3.8-27B 83.7/90.2, Qwen3.7-Plus 85.8/85.9, Opus-4.6 66.0/-.

打开官方来源

Qwen3.8-Flash-Next-Base

Qwen3.8-Flash-Next-Base 是同架构的预训练基座版本(未做后训练),发布文表 3 单独报告其通用、数学 STEM、编码与多语言基线:MMLU 90.36、BBH 90.87、GSM8K 93.29。

输入模态
官方资料未说明
上下文
256K(YaRN 可扩至 1M)
参数
125B-A6B(另含 51B N-gram embedding)
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmlu 90.36 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMLU (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Base-model table (pre-post-training). Competitor cells: Qwen3.8-27B-Base 87.51, Qwen3.7-Plus-Base 90.43.

打开官方来源

mmlu-redux 90.68 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMLU-Redux (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B-Base 87.26, Qwen3.7-Plus-Base 91.47.

打开官方来源

mmlu-pro 73.23 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMLU-Pro (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B-Base 68.60, Qwen3.7-Plus-Base 70.90.

打开官方来源

supergpqa 51.36 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SuperGPQA (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: supergpqa. Competitor cells: Qwen3.8-27B-Base 44.86, Qwen3.7-Plus-Base 48.42.

打开官方来源

bbh 90.87 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: BBH (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: bbh (BIG-Bench Hard) not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 89.56, Qwen3.7-Plus-Base 89.41.

打开官方来源

gpqa 51.42 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: GPQA (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Base GPQA (main set, not Diamond). Competitor cells: Qwen3.8-27B-Base 45.01, Qwen3.7-Plus-Base 51.52.

打开官方来源

gsm8k 93.29 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: GSM8K (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Competitor cells: Qwen3.8-27B-Base 93.18, Qwen3.7-Plus-Base 92.95.

打开官方来源

math 72.78 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MATH (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: math (also introduced in qwen2-5.json this batch). Competitor cells: Qwen3.8-27B-Base 60.54, Qwen3.7-Plus-Base 74.38.

打开官方来源

evalplus 78.76 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: EvalPlus (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: evalplus (also in qwen2-5.json). Competitor cells: Qwen3.8-27B-Base 76.05, Qwen3.7-Plus-Base 78.06.

打开官方来源

multipl-e 79.09 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MultiPL-E (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: multipl-e. Competitor cells: Qwen3.8-27B-Base 74.50, Qwen3.7-Plus-Base 81.68.

打开官方来源

swebench-pretrain 50.99 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: SWEBench-Pretrain (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swebench-pretrain not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 41.66, Qwen3.7-Plus-Base 49.24.

打开官方来源

mgsm 89.33 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MGSM (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

mgsm id exists in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 86.37, Qwen3.7-Plus-Base 85.42.

打开官方来源

mmmlu 84.86 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: MMMLU (Base, multilingual)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: mmmlu (multilingual MMLU) not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 79.74, Qwen3.7-Plus-Base 84.53.

打开官方来源

include 78.4 模型 qwen3-8-flash-next-base · 版本 Base model · 指标 未说明 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-08-31 · 距快照 34 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: 模型表现 tables (DOM x3) · table: DOM tables: instruct (12 benchmark rows), multimodal (10 benchmark rows), base model (14 benchmark rows) - all machine-readable · row: INCLUDE (Base)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: include (multilingual) not yet in data/benchmarks/. Competitor cells: Qwen3.8-27B-Base 74.37, Qwen3.7-Plus-Base 78.90.

打开官方来源