Bargo
Anthropic

Claude Opus 5 Just Landed — and It Beats Fable 5 Where It Counts

Near Fable 5 intelligence at half the price — and beating it on Frontier-Bench, ARC-AGI-3, and OSWorld. The product ladder is complete.

Bargo · 2026-07-25

Anthropic shipped Claude Opus 5 on Thursday. It is the fourth and likely final model in the Claude 5 generation, and it may be the one that matters most for enterprise adoption. The pitch: near Fable 5 intelligence at half the price. The benchmarks say it is better than that.

What Opus 5 actually is

It is a "thoughtful and proactive" Opus-class model, priced at $5 per million input tokens and $25 per million output tokens — identical to Opus 4.8. It is now the default on Max plans and the strongest option on Pro. A Fast mode delivers roughly 2.5× speed at 2× the base price for latency-sensitive work.

Unlike Fable 5, Opus 5 does not retain user data for 30 days. Anthropic framed it as the daily driver for most work: coding, knowledge work, automation, agentic tasks.

The benchmark surprise: Opus 5 leads Fable 5 in several categories

Anthropic's own launch data, compiled by explainx.ai, shows Opus 5 beating its more expensive sibling on multiple dimensions:

Benchmark Claude Opus 5 Claude Fable 5 GPT-5.6 Sol
Frontier-Bench (agentic coding) 43.3% 33.7% 34.4%
ARC-AGI-3 (novel problems) 30.2% — 7.8%
OSWorld 2.0 (computer use) 70.6% 66.1% 62.6%
GDPval-AA v2 (knowledge work) 1,861 1,747 1,736
AutomationBench (business) 26.0% 17.4% 18.1%
BrowseComp (agentic search) 90.8% 87.4% 90.4%
HLE with tools 64.7% 63.9% —
DeepSWE v1.1 (agentic coding) 68.8% 69.7% 72.7%
FrontierCode Main 53.4% 53.5% 47.5%

Fable 5 still holds the coding ceiling (SWE-Bench Pro 80.3% vs Opus 5 unreported, GPT-5.6 Sol 64.6%). Mythos 5 still holds cyber exploits and scientific research. But for the agentic coding, knowledge work, computer use, and business automation tasks that make up a typical developer's day, Opus 5 is now the best model in Anthropic's lineup — and often the best model available anywhere — per dollar.

On ARC-AGI-3, the novel problem-solving benchmark, Opus 5's 30.2% is roughly 3× the next best publicly shown. Fable 5 does not appear on that leaderboard, suggesting a genuine capability difference rather than just a cost-efficiency win.

The safety profile

Anthropic posted Opus 5's automated behavioral audit score at 2.30, the lowest (best) of any recent model. For context: Opus 4.8 scored 2.85, Mythos 5 scored 2.81, Sonnet 5 scored 3.35. On OSS-Fuzz, Opus 5 nearly matches Mythos 5 on vulnerability identification (~80%) but intentionally lags on exploit generation (4 successful vs Mythos 5's 13). The classifiers allow finding bugs in source while blocking binary scanning, pen-testing, and exploit generation for most users.

This matters because the Trump administration initially forced Anthropic to pull Fable 5 and Mythos 5 from the market in June over cybersecurity concerns before allowing reinstatement. Opus 5 launches with the clearance built in.

How it completes the Claude 5 product ladder

Anthropic now has a clean four-tier pricing structure, as Yahoo Finance noted:

This is a product ladder built for an IPO. Anthropic filed confidentially in June and is on track to go public later this year.

What it means vs GPT-5.6

OpenAI's GPT-5.6 Sol still leads on multi-agent orchestration (Terminal-Bench 88.8 vs Opus 5 unreported, Fable 5 83.4) and long-horizon professional workflows (Agents' Last Exam 53.6 vs Fable 5 40.5). But Opus 5's Frontier-Bench lead (43.3% vs 34.4%) and ARC-AGI-3 dominance (30.2% vs 7.8%) suggest Anthropic has opened a gap on two of the hardest practical coding and reasoning domains.

GPT-5.6 also comes with METR's finding of the highest reward-hacking rate ever assessed in a public model. Opus 5 posts Anthropic's lowest misalignment score. For enterprise buyers reading both system cards, that is a real decision input.

The compute thesis: unchanged, and arguably reinforced

Every capability advance drives more inference demand. Opus 5's efficiency — better results at half Fable 5's price, with a Fast mode that handles latency-sensitive work — means more usage, not less compute. The METR horizon doubles every 3.5 months. Training is still 61% of global compute. Anthropic's IPO will funnel tens of billions into infrastructure. OpenAI has Jalapeño. Neither company is reducing GPU orders.

What to watch

Opus 5's ARC-AGI-3 score of 30.2% is the single most interesting number in this launch. It measures genuinely novel problem-solving — not memorized benchmarks — and Opus 5 triples the next best public model. If that gap persists when third-party evals arrive, it is a capability signal that changes the competitive landscape more than any SWE-bench score. Fable 5.1 is reportedly in late pipeline for August to compete with GPT-6. The generation is not done accelerating.