Coding agents just jumped 5 points on SWE-Bench Pro

Coding agents that can’t see screenshots or understand visual documentation miss half the context they need to fix real bugs.

Most AI coding tools still work in text-only mode, forcing developers to describe UI bugs, error dialogs, and visual layouts in words instead of simply showing the model what’s broken.

StepFun Releases Step 3.7 Flash combines a 196B-parameter language model with native vision processing to read screenshots, understand visual interfaces, and execute coding tasks across both text and images. You upload code repositories along with screenshots or UI mockups, and it generates working solutions that account for both the codebase and visual requirements.

Three types of teams will feel this immediately

  • DevOps engineers who debug deployment dashboards and need agents that can read error screens
  • Frontend developers who build from design mockups and want AI that sees layout requirements
  • QA automation specialists who write tests based on visual bug reports from non-technical users

StepFun’s model scored 56.26% on SWE-Bench Pro compared to the previous version’s 51.3% — a jump that puts it ahead of several commercial alternatives. The vision capability means coding agents can finally work with the full context that human developers use daily.

What you can automate with visual coding agents

  • Generate frontend components directly from design screenshots or wireframes
  • Debug UI issues by analyzing error screenshots alongside code
  • Write automated tests that validate visual layouts and component behavior
  • Convert visual documentation into working code implementations

Pricing not listed — check our directory.

The 198B parameter count means you’ll need substantial compute resources for self-hosting.

AnthropClaude and OpenAI’s GPT-4V offer similar multimodal coding but through API-only access. Step 3.7 Flash runs Apache 2.0 licensed for complete control over your codebase.

Vision-first coding tools are becoming table stakes

The gap between text-only and multimodal coding agents is widening faster than most teams realize. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.