Katteb

Step 3.7 Flash Model to read images, video and code fast

StepFun Chat and reasoning models Upgrade to use stepfun/step-3.7-flash

Step 3.7 Flash is StepFun’s fast multimodal chat model that combines a large sparse language core with a built-in vision encoder for native image and video understanding. It is for developers and teams who need low-cost visual question answering, agentic chains, or coding help from screenshots and short recordings.

Use Step 3.7 Flash now All models

What it can do

  • Reads images and screenshots
  • Searches the live web and cites sources
  • Thinks step by step on hard questions
  • Understands video input
Takestext
Returnstext

What people use it for

Visual QA from screenshots and photos

Upload product photos, screenshots or receipts and get concise answers, labelled fields and verified measurements. The model reads images natively and summarizes visual content for help desks and catalog tagging.

Agentic workflows that combine search, tools and vision

Run conversational agents that inspect an image or short video, call a search or a tool, then return stepwise recommendations. Step 3.7 Flash is built for agentic pipelines where perception and reasoning must run together.

Code reading, debugging and code-aware screenshots

Paste code or upload a screen recording of a bug and get explanations, diffs and suggested fixes. The model’s language backbone and vision encoder let it comment on snippets and image-based stack traces.

Why it is worth it

Questions people ask

What types of images and video can Step 3.7 Flash handle?

It handles common image formats and short video clips for scene description, OCR-style reading and visual question answering. For best results use clear frames and short clips under a few seconds.

How is billing handled on Katteb for this model?

Katteb bills Step 3.7 Flash per message on actual usage, charged in credits. This model costs 3 credits per message in Katteb, billed on send.

When should I choose Step 3.7 Flash over a dense language model?

Choose Step 3.7 Flash when you need native vision or video input plus tool-oriented reasoning at lower inference cost. For pure long-form writing without images a dense model may be preferable.

Related models