Claude Opus 5 Just Landed — and It Beats Fable 5 Where It Counts
Near Fable 5 intelligence at half the price — and beating it on Frontier-Bench, ARC-AGI-3, and OSWorld. The product ladder is complete.
Anthropic shipped Claude Opus 5 on Thursday. It is the fourth and likely final model in the Claude 5 generation, and it may be the one that matters most for enterprise adoption. The pitch: near Fable 5 intelligence at half the price. The benchmarks say it is better than that.
What Opus 5 actually is
It is a "thoughtful and proactive" Opus-class model, priced at $5 per million input tokens and $25 per million output tokens — identical to Opus 4.8. It is now the default on Max plans and the strongest option on Pro. A Fast mode delivers roughly 2.5× speed at 2× the base price for latency-sensitive work.
Unlike Fable 5, Opus 5 does not retain user data for 30 days. Anthropic framed it as the daily driver for most work: coding, knowledge work, automation, agentic tasks.
The benchmark surprise: Opus 5 leads Fable 5 in several categories
Anthropic's own launch data, compiled by explainx.ai, shows Opus 5 beating its more expensive sibling on multiple dimensions:
| Benchmark | Claude Opus 5 | Claude Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| Frontier-Bench (agentic coding) | 43.3% | 33.7% | 34.4% |
| ARC-AGI-3 (novel problems) | 30.2% | — | 7.8% |
| OSWorld 2.0 (computer use) | 70.6% | 66.1% | 62.6% |
| GDPval-AA v2 (knowledge work) | 1,861 | 1,747 | 1,736 |
| AutomationBench (business) | 26.0% | 17.4% | 18.1% |
| BrowseComp (agentic search) | 90.8% | 87.4% | 90.4% |
| HLE with tools | 64.7% | 63.9% | — |
| DeepSWE v1.1 (agentic coding) | 68.8% | 69.7% | 72.7% |
| FrontierCode Main | 53.4% | 53.5% | 47.5% |
Fable 5 still holds the coding ceiling (SWE-Bench Pro 80.3% vs Opus 5 unreported, GPT-5.6 Sol 64.6%). Mythos 5 still holds cyber exploits and scientific research. But for the agentic coding, knowledge work, computer use, and business automation tasks that make up a typical developer's day, Opus 5 is now the best model in Anthropic's lineup — and often the best model available anywhere — per dollar.
On ARC-AGI-3, the novel problem-solving benchmark, Opus 5's 30.2% is roughly 3× the next best publicly shown. Fable 5 does not appear on that leaderboard, suggesting a genuine capability difference rather than just a cost-efficiency win.
The safety profile
Anthropic posted Opus 5's automated behavioral audit score at 2.30, the lowest (best) of any recent model. For context: Opus 4.8 scored 2.85, Mythos 5 scored 2.81, Sonnet 5 scored 3.35. On OSS-Fuzz, Opus 5 nearly matches Mythos 5 on vulnerability identification (~80%) but intentionally lags on exploit generation (4 successful vs Mythos 5's 13). The classifiers allow finding bugs in source while blocking binary scanning, pen-testing, and exploit generation for most users.
This matters because the Trump administration initially forced Anthropic to pull Fable 5 and Mythos 5 from the market in June over cybersecurity concerns before allowing reinstatement. Opus 5 launches with the clearance built in.
How it completes the Claude 5 product ladder
Anthropic now has a clean four-tier pricing structure, as Yahoo Finance noted:
- Mythos 5: the real frontier, gated behind Project Glasswing (government/cyber)
- Fable 5: the public frontier, days-long autonomous projects ($10/$50)
- Opus 5: the daily workhorse, near-frontier at half the price ($5/$25)
- Sonnet 5: mass-market agentic, Free/Pro default ($3/$15)
This is a product ladder built for an IPO. Anthropic filed confidentially in June and is on track to go public later this year.
What it means vs GPT-5.6
OpenAI's GPT-5.6 Sol still leads on multi-agent orchestration (Terminal-Bench 88.8 vs Opus 5 unreported, Fable 5 83.4) and long-horizon professional workflows (Agents' Last Exam 53.6 vs Fable 5 40.5). But Opus 5's Frontier-Bench lead (43.3% vs 34.4%) and ARC-AGI-3 dominance (30.2% vs 7.8%) suggest Anthropic has opened a gap on two of the hardest practical coding and reasoning domains.
GPT-5.6 also comes with METR's finding of the highest reward-hacking rate ever assessed in a public model. Opus 5 posts Anthropic's lowest misalignment score. For enterprise buyers reading both system cards, that is a real decision input.
The compute thesis: unchanged, and arguably reinforced
Every capability advance drives more inference demand. Opus 5's efficiency — better results at half Fable 5's price, with a Fast mode that handles latency-sensitive work — means more usage, not less compute. The METR horizon doubles every 3.5 months. Training is still 61% of global compute. Anthropic's IPO will funnel tens of billions into infrastructure. OpenAI has Jalapeño. Neither company is reducing GPU orders.
What to watch
Opus 5's ARC-AGI-3 score of 30.2% is the single most interesting number in this launch. It measures genuinely novel problem-solving — not memorized benchmarks — and Opus 5 triples the next best public model. If that gap persists when third-party evals arrive, it is a capability signal that changes the competitive landscape more than any SWE-bench score. Fable 5.1 is reportedly in late pipeline for August to compete with GPT-6. The generation is not done accelerating.