Anthropic has released Claude Opus 5, a new frontier model that approaches the intelligence of its higher-end Claude Fable 5 offering at half the price. It replaces its predecessor, Opus 4.8, delivering significantly better performance without increasing the cost per task.
For the AI market, the headline is efficiency. Opus 5 is now the default model on Claude Max and the strongest model on Claude Pro. It allows users to adjust an "effort setting" to dial up intelligence or conserve tokens for cheaper, faster results. Across coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, it has set new state-of-the-art benchmarks, though Anthropic notes it still trails Mythos 5 on cybersecurity tasks.
Multi-Step Reliability is the New Baseline
What makes Opus 5 notable isn't just its raw scoring—it's how it handles complex, multi-step execution without failing halfway through.
On Zapier’s AutomationBench, which tests a model's ability to complete business tasks from start to finish, Opus 5 achieves a pass rate roughly 1.5 times higher than the next-best model at the same cost. Zapier CEO Wade Foster noted that Opus 5 topped their leaderboard by taking a raw account-health workbook and successfully running a full churn-prevention sequence end-to-end. It flagged at-risk accounts, alerted the appropriate owner, and summarized the data for retention operations. Where previous models failed to pass, Opus 5 hit a 100% success rate.
This matters for small and mid-sized businesses. When building operations automations, the primary bottleneck isn't usually writing emails; it's reliably passing structured data between systems without requiring a human to babysit the process. Models that fail at step four of a five-step Zapier workflow are essentially useless for business administration. Opus 5 proves the technology is moving past that hurdle.
Verification and Agency in the Wild
Anthropic’s release heavily emphasizes Opus 5's capability to verify its own work and iterate carefully until a task succeeds.
The examples provided highlight a distinct shift toward autonomous problem-solving. In one instance, when tasked with rebuilding a machine part as a 3D FreeCAD model but intentionally blocked from directly viewing the drawing, Opus 5 wrote its own computer vision pipeline to extract the geometry from raw pixels. In another real-world test, an engineer at a trading firm used the model to build a market data feed for a new exchange. Because there was no live feed to validate against, Opus 5 built its own test harness to ensure its code parsed the data correctly.
Enterprise software leaders are already seeing the downstream effects of this reasoning capability. Box CTO Ben Kus reported that Opus 5 outperforms Opus 4.8 by 8% overall, delivering an 11% improvement in data analysis and a 17% jump in due diligence workflows—the exact type of dense, unstructured data processing that professional services and healthcare organizations rely on daily.
What Opus 5 Means for SMB Automations
For businesses looking to automate their back-office operations, the release of Opus 5 translates directly to higher ROI on automation investments.
First, cost predictability. By delivering Fable 5-level performance at Opus 4.8 prices, businesses can run sophisticated, high-token analyses—like processing large volumes of enterprise content or financial modeling—without blowing up their API budgets. On OSWorld 2.0, a computer-use benchmark, Opus 5 surpassed Fable 5’s best result at just over a third of the cost.
Second, consistency. As Lovable co-founder Fabian Hedin noted in the release, Opus 5 has "far less variance run to run." For SMBs, consistency is the entire ballgame. A fractional AI partner can build a workflow, but if the underlying model responds differently to the same invoice format on Tuesday than it did on Monday, the automation becomes a liability.
Opus 5 shows that frontier AI is maturing into steady, industrial-grade software. The technology is now fully capable of handling dense administrative tasks, root-cause debugging, and deep data analysis with a high degree of autonomy. For businesses, the challenge is no longer waiting for the AI to get smart enough—it's identifying which operational bottlenecks to point it at first.