← 模型目录

Step 3

StepFun / 阶跃星辰 · 2025-07-31 · 类别未确认

官方发布来源 · 按需求选型 · 查看覆盖缺口 · 归档与转录记录

Step 3

StepFun 将 Step 3 定位为高性价比的多模态智能模型,以 38B 激活的 MoE 架构与 MFA 注意力 + AFD 并行设计平衡性能与推理成本,并以 Apache 2.0 开源。已收录评测覆盖多模态理解、数学与科学推理、代码等领域,AIME 2025 82.9、MMMU 74.2。

输入模态
文本 / 图像
上下文
64K(65536)
参数
321B-A38B MoE(VLM 321B / LLM 316B)
价格(每百万 tokens)
无此口径报价;不等于免费

本变体的评测证据

mmmu 74.2 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: mmmu

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

mathvision 64.8 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: mathvision

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

simplevqa 62.2 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: simplevqa

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

hallusionbench 64.2 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: hallusionbench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

zerobench 23 模型 step-3 · 版本 sub · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: zerobench

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

dyna-math 50.1 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: dyna-math

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

aime-25 82.9 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: aime-25

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

gpqa 73 模型 step-3 · 版本 Diamond · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: gpqa

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

lcb 67.1 模型 step-3 · 版本 202408-202505 · 指标 pass_rate · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: lcb

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

hmmt25 70 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: hmmt25

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源

cnmo-2024 83.7 模型 step-3 · 版本 未说明 · 指标 accuracy · 单位 percent 来源等级 A · vendor_reported · 核对日期 2026-09-01 · 距快照 33 天

归一化读数:不计算;样本及方差不完整,无法计算置信区间。

原文位置与完整协议

heading: Overall Performance · table: Overall Performance (VLM + LLM benchmarks) · row: cnmo-2024

{
  "harness": null,
  "tools": null,
  "shots": null,
  "reasoning_effort": null,
  "temperature": null,
  "top_p": null,
  "token_budget": null,
  "turn_limit": null,
  "time_limit": null,
  "run_count": null,
  "aggregation": null,
  "judge": null
}

2026-09-01 audit: value cross-confirmed against official GitHub figure figures/step3_bmk.jpeg.

打开官方来源