Step 3.7 Flash
Stepfun🇨🇳
stepfun/step-3.7-flashStep 3.7 Flash is StepFun's multimodal vision-language model, built on Step 3.5 Flash with an added vision encoder for native image understanding. With a 256k context window, selectable reasoning levels, and high throughput, it targets agentic, coding, and search workflows that mix text and visual input.
费用比率
0.2x
上下文
262K
发布日期
2026年5月28日
输入
TextImageVideo
输出
Text
支持
推理工具调用结构化输出