Step 3.7 Flash
Stepfun🇨🇳
stepfun/step-3.7-flashStep 3.7 Flash is StepFun's multimodal vision-language model, built on Step 3.5 Flash with an added vision encoder for native image understanding. With a 256k context window, selectable reasoning levels, and high throughput, it targets agentic, coding, and search workflows that mix text and visual input.
Kostenfaktor
0.2x
Kontext
262K
Veröffentlicht
28. Mai 2026
Eingabe
TextImageVideo
Ausgabe
Text
Unterstützung
SchlussfolgernTool-AufrufeStrukturierte Ausgaben