MiniMax M3 Fits 1M Tokens and Sees Your Screen

Every model that forces you to chunk your codebase, summarize your own documents, or switch tools mid-task is costing you compounding hours you never get back.

Long-context coding has been a duct-tape problem until now

Engineers working on large repositories or analysts processing massive document sets have had to pre-filter, truncate, or batch inputs just to fit model limits. That preprocessing step introduces errors and buries the context that would have changed the answer.

One model reads your million-token codebase and watches your desktop do it

MiniMax Releases MiniMax M3 with MSA Architecture Supporting 1M accepts text, image, and video input natively, processes up to one million tokens in a single context window using its new MSA (MiniMax Sparse Attention) architecture, and can operate a desktop computer directly for agentic coding tasks. You call it through the MiniMax API, MiniMax Code, or the Token Plan, and the open weights drop within ten days of the June 1 launch. The MSA architecture processes that context more than four times faster than comparable open-source sparse attention implementations by reading each KV cache block only once in contiguous memory.

Three roles will feel this difference by end of week

  • Staff engineers maintaining legacy monorepos who need a model that ingests the entire codebase without summarization loss
  • AI researchers benchmarking open-weight frontier models who want weights they can inspect, fine-tune, and run on their own infrastructure
  • Automation engineers building agentic pipelines who need a single model that can read a screen, write code, and act on a desktop without chaining separate tools

The combination of all three capabilities in one open-weight release is what separates this from incremental context bumps other labs have shipped.

The open-weight frontier just moved faster than most teams expected

Closed models from Anthropic and OpenAI have held the multimodal agentic lead for most of 2025, but MiniMax M3 is the first open-weight model to pair a verified 1M-token window with native vision and computer-use in a single architecture. If the weights perform at API parity when they ship, enterprise teams with on-premise requirements get a credible alternative they did not have last quarter.

What you can actually do with it today

  • Send a full million-token codebase and ask for a refactor plan without chunking
  • Pass video input directly into a coding or analysis prompt as a native modality
  • Run agentic desktop tasks end-to-end without a separate computer-use model
  • Pull open weights within ten days and deploy on private infrastructure

MiniMax M3 is available now via the MiniMax API with pricing tied to the Token Plan. Pricing not listed for enterprise tiers — check our directory.

The one gap worth watching before you commit

The weights are not public yet, so independent benchmark verification against the API performance claims is still ten days out at minimum.

For teams that need open weights today, Qwen2.5 covers long-context reasoning at a lower token ceiling. For closed-model multimodal agents, Gemini 1.5 Pro holds a longer public track record on million-token tasks.

Open-weight agentic models are closing the gap faster than the labs want to admit

This is exactly the kind of structural shift we track week over week — who is closing the capability gap, who is losing it, and what it means for your stack. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.