GPT-5.6 Sol beats Anthropic Fable on every coding benchmark

If your team chose Anthropic’s Fable last quarter based on benchmarks, those numbers just got older.

Enterprise coding reviews were eating hours that nobody was billing

Security-conscious engineering teams have been stuck choosing between model quality and token cost when running code review, threat modeling, and blue team simulations at scale. That tradeoff is what this release directly attacks.

OpenAI just shipped its most cost-efficient model family yet

OpenAI launches its new family of models with GPT comes in three variants: Sol (the heavy-duty workhorse), Terra (mid-range), and Luna (budget). You access them through the existing API or the new ChatGPT Work desktop and web app, where you can paste code, upload documents, or run threat modeling prompts and get structured outputs including patched code, reviewed spreadsheets, or blue team reports.

Security engineers and enterprise developers feel this first

  • Security engineers running blue team exercises who need frontier-level threat modeling without burning through token budgets on every session
  • Enterprise developers who benchmark coding agents against Anthropic Fable and need a current, documented comparison before committing to a vendor
  • IT procurement leads who approve AI tooling costs and need the token efficiency numbers to justify switching or expanding API usage

The timing is not incidental. OpenAI cites the Artificial Analysis Coding Agent Index to claim Sol outperforms Anthropic’s models at every data point, and the release lands the same week that SpaceXAI and Meta both pushed new model families into the market.

Anthropic built its enterprise lead slowly — OpenAI is trying to erase it in one week

Sol is 54% more token-efficient on coding tasks than prior OpenAI versions, a number that materially changes monthly API invoices for teams running high-volume agents. If that efficiency gap holds up under independent testing, the enterprise cost argument that has quietly favored Anthropic gets a lot harder to make.

What you can actually do with it today

  • Run blue team attack simulations against your own codebase to surface vulnerabilities
  • Submit code for automated review and receive patched output directly
  • Draft enterprise documents, spreadsheets, and presentations inside ChatGPT Work
  • Benchmark Sol against Fable using the Artificial Analysis Coding Agent Index for your own stack

Pricing follows OpenAI’s existing API tiers, with Luna positioned as the low-cost entry point — check the OpenAI pricing page for current per-token rates across all three variants.

The cybersecurity capabilities drew enough regulatory concern from the Trump administration to nearly delay the rollout, which means real-world misuse potential is not hypothetical and responsible deployment guardrails will matter in practice.

Anthropic Claude remains the default choice for teams that built workflows around its API, and switching costs are real regardless of benchmark results.

The coding model wars are now a benchmark arms race with real budget consequences

The winner of this cycle will not be decided by launch announcements but by which model holds its efficiency claims under production workloads over the next 90 days. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.