Claude Opus 5 Is Anthropic’s Fourth Frontier Model in Two Months. Here’s What Actually Changed

A wall display showing the Claude logo beside a hardware dial labeled low, medium and high, with a developer silhouetted in front of it

Anthropic released Claude Opus 5 on Friday at $5 per million input tokens and $25 per million output tokens, the exact price of the model it replaces, while claiming it roughly doubles that model’s performance on the benchmarks the company chose to publish.

The price tag is what everyone will quote, but the more consequential change is sitting in the safety section of Anthropic’s own announcement, where the company discloses that Opus 5’s cybersecurity filters step in about 85% less often than its flagship’s do.

The Price Did Not Move. The Math Did.

Holding a list price flat is not a price cut, and Anthropic is careful never to call it one. But if the per-token cost stays at $5 and $25 while the model finishes more work per token, the effective cost of a completed task falls, and that is the number enterprise buyers actually track.

The company’s Opus 5 announcement puts real figures behind that. On CursorBench 3.2, a coding evaluation, Opus 5 lands within half a percentage point of Fable 5, Anthropic’s own top-end model, at half the cost per task. On OSWorld 2.0, which tests a model driving an actual computer, it beats Fable 5 at roughly a third of the cost. On ARC-AGI 3, a benchmark deliberately built out of problems no model has seen before, Anthropic says Opus 5 scores three times higher than the next best system on the board.

Treat vendor benchmarks with the skepticism they have earned. Every lab picks the evals it wins. What is harder to wave away is the shape of the claim: Anthropic is not arguing that Opus 5 is the smartest thing it has ever built. It is arguing that Opus 5 is the cheapest way to get most of the way there, and it is willing to point at its own flagship as the thing being undercut.

The Effort Dial Is the Real Product

The feature Anthropic is leading with is a setting, not a capability. Developers can now tell Opus 5 how hard to think, choosing low, medium, or high effort on a per-request basis, and pay accordingly.

That sounds mundane. It is the most honest thing about the release. For two years the industry sold reasoning as an unalloyed good, and customers discovered that a model which thinks for ninety seconds about a formatting question generates a bill that thinks for ninety seconds about a formatting question. Anthropic is now conceding, in product form, that most work does not need frontier intelligence and that the customer is a better judge of which work does.

The early numbers from customers with a reason to measure are specific. The legal AI company Harvey reported that Opus 5 matched the previous model’s maximum-reasoning output while generating 26% fewer tokens. Box measured an 8% overall improvement, rising to 17% on due-diligence workflows. Those are efficiency gains, not intelligence gains, and the distinction matters because efficiency is the thing enterprise finance departments have been shouting about since the first eye-watering invoice landed.

Anthropic Is Undercutting Anthropic

This is the fourth frontier model Anthropic has shipped in under two months, after Mythos 5, Fable 5, and Sonnet 5 in June. Fortune noted in its coverage of the launch that Anthropic openly concedes Fable 5 is still the better choice for long-running autonomous projects, positioning Opus 5 as the affordable alternative for everything else.

Read that as strategy rather than modesty. A company that files a release cadence like this is not chasing a leaderboard, it is defending an install base. Roughly four fifths of Anthropic’s revenue comes from enterprises and developers rather than consumers, which makes API spend the whole ballgame, and API spend is exactly what a rival undercuts first. We looked at how Anthropic built its lead in the enterprise Claude market earlier this year, and the defensive logic here follows directly from it.

OpenAI made the same bet two weeks earlier, pitching GPT-5.6 on token economy when it launched on July 9. When two labs that disagree about nearly everything both decide that the pitch is thrift, the competitive frontier has moved. It is no longer about who has the smartest model. It is about who can deliver an acceptable answer for the least money, which is a commodity fight, and commodity fights compress margins for everyone in them.

The Guardrails Got Looser, Not Tighter

Here is the part that deserves more attention than it will get.

Anthropic describes Opus 5 as its most aligned Opus model, with the lowest misaligned-behavior score of any recent release. In the same document, it discloses that the model’s cybersecurity classifiers intervene about 85% less often than Fable 5’s, and that the API now automatically reroutes to a different model when Opus 5 declines a request on safety grounds.

Both things can be true. A model that refuses fewer legitimate security-research requests is a better tool for the defenders who make up most of that user base, and over-refusal is a genuine failure mode that has frustrated professionals for years. But “we tuned our safety system to trigger far less often” is a decision with a direction, and it arrived in the same release that Anthropic marketed on cost and convenience. Loosening a filter has a measurable upside on the product side and a diffuse, deferred, largely external downside. That asymmetry is precisely why this class of decision should not be graded by the company making it.

The automatic fallback is the stranger piece. When a request trips a safety limit, the system does not simply stop, it quietly hands the work to another model. Engadget’s write-up of the launch treats that as a smoother developer experience, and for the developer it is. As a safety architecture, a refusal that reroutes rather than a refusal that holds is a softer boundary than the word “refusal” implies.

What to Watch

The number worth tracking is not a benchmark. It is whether anyone outside Anthropic can reproduce the cost claims on their own workloads over the next quarter, because vendor efficiency figures are measured on tasks vendors picked and real workloads are messier than evaluations.

The second thing to watch is what happens to Fable 5. A company that just shipped a cheaper model matching most of its flagship’s output has created an obvious question about who is still paying double, and the answer will show up in Anthropic’s pricing sheet long before it shows up in a press release. When a lab starts competing against its own premium tier, the honest read is that it expects someone else to do it first.