Step 3.7 Flash Model to read images, video and code fast
StepFun Chat and reasoning models Upgrade to use stepfun/step-3.7-flash
Step 3.7 Flash is StepFun’s fast multimodal chat model that combines a large sparse language core with a built-in vision encoder for native image and video understanding. It is for developers and teams who need low-cost visual question answering, agentic chains, or coding help from screenshots and short recordings.
What it can do
- Reads images and screenshots
- Searches the live web and cites sources
- Thinks step by step on hard questions
- Understands video input
What people use it for
Visual QA from screenshots and photos
Upload product photos, screenshots or receipts and get concise answers, labelled fields and verified measurements. The model reads images natively and summarizes visual content for help desks and catalog tagging.
Agentic workflows that combine search, tools and vision
Run conversational agents that inspect an image or short video, call a search or a tool, then return stepwise recommendations. Step 3.7 Flash is built for agentic pipelines where perception and reasoning must run together.
Code reading, debugging and code-aware screenshots
Paste code or upload a screen recording of a bug and get explanations, diffs and suggested fixes. The model’s language backbone and vision encoder let it comment on snippets and image-based stack traces.
Why it is worth it
- Multimodal answers that combine text, images and short video inputs.
- Sparse MoE architecture for higher efficiency at inference time.
- Works well in agent chains that require tool use and stepwise reasoning.
- Runs faster and cheaper than larger dense multimodal models on many prompts.
Questions people ask
What types of images and video can Step 3.7 Flash handle?
It handles common image formats and short video clips for scene description, OCR-style reading and visual question answering. For best results use clear frames and short clips under a few seconds.
How is billing handled on Katteb for this model?
Katteb bills Step 3.7 Flash per message on actual usage, charged in credits. This model costs 3 credits per message in Katteb, billed on send.
When should I choose Step 3.7 Flash over a dense language model?
Choose Step 3.7 Flash when you need native vision or video input plus tool-oriented reasoning at lower inference cost. For pure long-form writing without images a dense model may be preferable.