← 模型目录

Step 3.7 Flash

StepFun / 阶跃星辰 · 2026-05-29 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Step 3.7 Flash

Step 3.7 Flash 被阶跃星辰定位为面向真实场景代理的高效率 Flash 模型,支持三档推理级别、最高 400 tok/s。评测覆盖通用代理、编码与长上下文/视觉:BrowseComp 75.8%、HLE w. tool 47.2%、SWE-bench Verified 76.5%、V* 95.29%。

输入模态
文本 / 图像
上下文
256K
参数
196B-A11B(另含 1.8B ViT)
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

hlehle 47.2 模型 step-3-7-flash · 版本 w. tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: hlehle

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hlehle not yet in data/benchmarks/.

打开官方来源

browsecomp 75.8 模型 step-3-7-flash · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: browsecomp

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

deepsearchqa 92.8 模型 step-3-7-flash · 版本 F1 · 指标 f1 · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: deepsearchqa

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

deepsearchqa 81.7 模型 step-3-7-flash · 版本 acc · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: deepsearchqa

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

researchrubrics 71.7 模型 step-3-7-flash · 版本 未说明 · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: researchrubrics

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: researchrubrics not yet in data/benchmarks/.

打开官方来源

toolathlon 49.5 模型 step-3-7-flash · 版本 未说明 · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: toolathlon

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

claw-eval 67.1 模型 step-3-7-flash · 版本 v1.1 pass^3 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: claw-eval

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

swebench-multilingual 72.4 模型 step-3-7-flash · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: swe-mtlg

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: swe-mtlg not yet in data/benchmarks/.

打开官方来源

swebench-pro 56.3 模型 step-3-7-flash · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: swebench-pro

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

swebench 76.5 模型 step-3-7-flash · 版本 Verified · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: swebench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

terminalbench 59.6 模型 step-3-7-flash · 版本 2.1 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: terminalbench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

aa-lcr 63.9 模型 step-3-7-flash · 版本 avg@16 · 指标 score · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: aa-lcr

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

打开官方来源

gdpval-aa 1415.8 模型 step-3-7-flash · 版本 Stirrup · 指标 elo_rating · 单位 points 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: GDPval-Stirrup (Elo)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Page note: on GDPval, the Step 3.7 Flash score is obtained through internal pairwise evaluation, while comparison models are sourced from the official Artificial Analysis Leaderboard. internal pairwise evaluation vs official AA leaderboard competitors

打开官方来源

gdpval-aa 45.8 模型 step-3-7-flash · 版本 internal pairwise (ii) · 指标 win_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: GDPval internal pairwise (ii)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Page note: on GDPval, the Step 3.7 Flash score is obtained through internal pairwise evaluation, while comparison models are sourced from the official Artificial Analysis Leaderboard. score obtained through internal pairwise evaluation per page note

打开官方来源

vstar 95.29 模型 step-3-7-flash · 版本 with Python tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: vstar

{
  "harness": null,
  "tools": [
    "python_tool (crop/zoom/draw)"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).

打开官方来源

hr-bench 89.13 模型 step-3-7-flash · 版本 4K, with Python tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: hr-bench

{
  "harness": null,
  "tools": [
    "python_tool (crop/zoom/draw)"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hr-bench not yet in data/benchmarks/. Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).

打开官方来源

hr-bench 86.34 模型 step-3-7-flash · 版本 8K, with Python tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: hr-bench

{
  "harness": null,
  "tools": [
    "python_tool (crop/zoom/draw)"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: hr-bench not yet in data/benchmarks/. Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).

打开官方来源

visualprobe 65.05 模型 step-3-7-flash · 版本 with Python tool · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: visualprobe

{
  "harness": null,
  "tools": [
    "python_tool (crop/zoom/draw)"
  ],
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: visualprobe not yet in data/benchmarks/. Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).

打开官方来源

step-swe-bench 67.08 模型 step-3-7-flash · 版本 未说明 · 指标 resolved_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Agentic Coding → Step-SWE-Bench · table: Step-SWE-Bench per-harness table · row: Step 3.7 Flash (avg)

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

new-benchmark: step-swe-bench not yet in data/benchmarks/. Vendor in-house SWE benchmark; avg across six harnesses (Hermes Agent 67.5 / OpenClaw 67.0 / Claude Code 71.5 / KiloCode 67.5 / OpenCode 64.5 / RooCode 64.5). Step 3.5 Flash same-table avg 56.50% (cross-release competitor row not transcribed separately).

打开官方来源

tau2-bench 官方未公布数值 模型 step-3-7-flash · 版本 Telecom · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Sharpened for Enterprise Tasks · row: Tau2-bench Telecom · quote_snippet: passes at over 98% across different reasoning difficulty tiers on Tau2-bench Telecom

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

Prose gives only a lower-bound phrasing ('over 98%') with no exact number; value kept null per data rules, quote preserved for locating.

打开官方来源