Step 3.7 Flash
StepFun / 阶跃星辰 · 2026-05-29 · 类别未确认
官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录
Step 3.7 Flash
Step 3.7 Flash 被阶跃星辰定位为面向真实场景代理的高效率 Flash 模型,支持三档推理级别、最高 400 tok/s。评测覆盖通用代理、编码与长上下文/视觉:BrowseComp 75.8%、HLE w. tool 47.2%、SWE-bench Verified 76.5%、V* 95.29%。
- 输入模态
- 文本 / 图像
- 上下文
- 256K
- 参数
- 196B-A11B(另含 1.8B ViT)
- 价格(每百万 tokens)
- 无此口径报价;不等于免费
本变体的评测证据
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: hlehle
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hlehle not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: browsecomp
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: deepsearchqa
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: deepsearchqa
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: researchrubrics
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: researchrubrics not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: toolathlon
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: claw-eval
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: swe-mtlg
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: swe-mtlg not yet in data/benchmarks/.
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: swebench-pro
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: swebench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: terminalbench
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: aa-lcr
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: GDPval-Stirrup (Elo)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page note: on GDPval, the Step 3.7 Flash score is obtained through internal pairwise evaluation, while comparison models are sourced from the official Artificial Analysis Leaderboard. internal pairwise evaluation vs official AA leaderboard competitors
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Benchmarks · table: Benchmarks (Flash Level / Pro Level) · row: GDPval internal pairwise (ii)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Page note: on GDPval, the Step 3.7 Flash score is obtained through internal pairwise evaluation, while comparison models are sourced from the official Artificial Analysis Leaderboard. score obtained through internal pairwise evaluation per page note
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Recognition with Visual Search · table: Visual Recognition with Visual Search · row: simplevqa
{
"harness": null,
"tools": [
"visual_search"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Recognition with Visual Search · table: Visual Recognition with Visual Search · row: worldvqa
{
"harness": null,
"tools": [
"visual_search"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Recognition with Visual Search · table: Visual Recognition with Visual Search · row: bc-vl
{
"harness": null,
"tools": [
"visual_search"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: bc-vl not yet in data/benchmarks/. Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: vstar
{
"harness": null,
"tools": [
"python_tool (crop/zoom/draw)"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: hr-bench
{
"harness": null,
"tools": [
"python_tool (crop/zoom/draw)"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hr-bench not yet in data/benchmarks/. Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: hr-bench
{
"harness": null,
"tools": [
"python_tool (crop/zoom/draw)"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: hr-bench not yet in data/benchmarks/. Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Visual Perception with Python Tool · table: Visual Perception with Python Tool · row: visualprobe
{
"harness": null,
"tools": [
"python_tool (crop/zoom/draw)"
],
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: visualprobe not yet in data/benchmarks/. Table footnote: * denotes a self-tested score (competitor cells); GLM results aligned with official GLM personnel using crop + search tools. BC-VL likely denotes BrowseComp-VL (page prints only the abbreviation).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Agentic Coding → Step-SWE-Bench · table: Step-SWE-Bench per-harness table · row: Step 3.7 Flash (avg)
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}new-benchmark: step-swe-bench not yet in data/benchmarks/. Vendor in-house SWE benchmark; avg across six harnesses (Hermes Agent 67.5 / OpenClaw 67.0 / Claude Code 71.5 / KiloCode 67.5 / OpenCode 64.5 / RooCode 64.5). Step 3.5 Flash same-table avg 56.50% (cross-release competitor row not transcribed separately).
归一化读数:不计算;样本及方差不完整,无法计算置信区间。
原文位置与完整协议
heading: Sharpened for Enterprise Tasks · row: Tau2-bench Telecom · quote_snippet: passes at over 98% across different reasoning difficulty tiers on Tau2-bench Telecom
{
"harness": null,
"tools": null,
"shots": null,
"reasoning_effort": null,
"temperature": null,
"top_p": null,
"token_budget": null,
"turn_limit": null,
"time_limit": null,
"run_count": null,
"aggregation": null,
"judge": null
}Prose gives only a lower-bound phrasing ('over 98%') with no exact number; value kept null per data rules, quote preserved for locating.