Opus 5 Hit 96% on SWE-bench Verified. Three Labs Are Now 1.1 Points Apart.
Claude Opus 5 launched July 24 and essentially saturated SWE-bench Verified at 96.0%. On the harder Pro variant, the top three flagship models from three different labs land within 1.1 points of each other. Model choice stopped being your coding differentiator.
NeuroX AI · July 27, 2026

Anthropic shipped Claude Opus 5 on July 24 at $5/$25 per million tokens — the same price as Opus 4.8. It scores 96.0% on SWE-bench Verified, which is the polite way of saying that benchmark is finished.
The number worth staring at is the other one. On SWE-bench Pro — the messier multi-file variant — the leaderboard reads Mythos 5 at 80.3%, Fable 5 at 80.0%, Opus 5 at 79.2%. Three flagship models, three different labs, 1.1 points end to end. Whatever you were arguing about in your model-selection doc is now inside the error bars.
That convergence is specific to coding. On ARC-AGI-3, which tests novel reasoning rather than patching known repos, Opus 5 scores 30.2% against 7.8% for the next publicly listed model — a 4x spread. And Frontier-Bench, which measures agentic coding end to end, still tops out at 43.3%. So: writing the diff is solved, judgment isn't, and running the whole task unsupervised isn't close.
If you're shipping agents, that reprices your decisions. Swapping models to chase a point of SWE-bench Pro buys you nothing measurable. The 57 points missing from Frontier-Bench are not a model problem — they're scoped tasks, verification the agent can't fake, and a definition of done a human can check in under a minute.
Stop shopping models. Build the harness.