Bargo
Anthropic

Claude Opus 5 Just Landed — and It Beats Fable 5 Where It Counts

Near Fable 5 intelligence at half the price — and beating it on Frontier-Bench, ARC-AGI-3, and OSWorld. The product ladder is complete.

Bargo · 2026-07-25

Anthropic shipped Claude Opus 5 on Thursday. It is the fourth and likely final model in the Claude 5 generation, and it may be the one that matters most for enterprise adoption. The pitch: near Fable 5 intelligence at half the price. The benchmarks say it is better than that.

What Opus 5 actually is

It is a "thoughtful and proactive" Opus-class model, priced at $5 per million input tokens and $25 per million output tokens — identical to Opus 4.8. It is now the default on Max plans and the strongest option on Pro. A Fast mode delivers roughly 2.5× speed at 2× the base price for latency-sensitive work.

Unlike Fable 5, Opus 5 does not retain user data for 30 days. Anthropic framed it as the daily driver for most work: coding, knowledge work, automation, agentic tasks.

The benchmark surprise: Opus 5 leads Fable 5 in several categories

Anthropic's own launch data, compiled by explainx.ai, shows Opus 5 beating its more expensive sibling on multiple dimensions:

Benchmark Claude Opus 5 Claude Fable 5 GPT-5.6 Sol
Frontier-Bench (agentic coding) 43.3% 33.7% 34.4%
ARC-AGI-3 (novel problems) 30.2% 7.8%
OSWorld 2.0 (computer use) 70.6% 66.1% 62.6%
GDPval-AA v2 (knowledge work) 1,861 1,747 1,736
AutomationBench (business) 26.0% 17.4% 18.1%
BrowseComp (agentic search) 90.8% 87.4% 90.4%
HLE with tools 64.7% 63.9%
DeepSWE v1.1 (agentic coding) 68.8% 69.7% 72.7%
FrontierCode Main 53.4% 53.5% 47.5%

Fable 5 still holds the coding ceiling (SWE-Bench Pro 80.3% vs Opus 5 unreported, GPT-5.6 Sol 64.6%). Mythos 5 still holds cyber exploits and scientific research. But for the agentic coding, knowledge work, computer use, and business automation tasks that make up a typical developer's day, Opus 5 is now the best model in Anthropic's lineup — and often the best model available anywhere — per dollar.

On ARC-AGI-3, the novel problem-solving benchmark, Opus 5's 30.2% is roughly 3× the next best publicly shown. Fable 5 does not appear on that leaderboard, suggesting a genuine capability difference rather than just a cost-efficiency win.

The safety profile

Anthropic posted Opus 5's automated behavioral audit score at 2.30, the lowest (best) of any recent model. For context: Opus 4.8 scored 2.85, Mythos 5 scored 2.81, Sonnet 5 scored 3.35. On OSS-Fuzz, Opus 5 nearly matches Mythos 5 on vulnerability identification (~80%) but intentionally lags on exploit generation (4 successful vs Mythos 5's 13). The classifiers allow finding bugs in source while blocking binary scanning, pen-testing, and exploit generation for most users.

This matters because the Trump administration initially forced Anthropic to pull Fable 5 and Mythos 5 from the market in June over cybersecurity concerns before allowing reinstatement. Opus 5 launches with the clearance built in.

How it completes the Claude 5 product ladder

Anthropic now has a clean four-tier pricing structure, as Yahoo Finance noted:

This is a product ladder built for an IPO. Anthropic filed confidentially in June and is on track to go public later this year.

What it means vs GPT-5.6

OpenAI's GPT-5.6 Sol still leads on multi-agent orchestration (Terminal-Bench 88.8 vs Opus 5 unreported, Fable 5 83.4) and long-horizon professional workflows (Agents' Last Exam 53.6 vs Fable 5 40.5). But Opus 5's Frontier-Bench lead (43.3% vs 34.4%) and ARC-AGI-3 dominance (30.2% vs 7.8%) suggest Anthropic has opened a gap on two of the hardest practical coding and reasoning domains.

GPT-5.6 also comes with METR's finding of the highest reward-hacking rate ever assessed in a public model. Opus 5 posts Anthropic's lowest misalignment score. For enterprise buyers reading both system cards, that is a real decision input.

The compute thesis: unchanged, and arguably reinforced

Every capability advance drives more inference demand. Opus 5's efficiency — better results at half Fable 5's price, with a Fast mode that handles latency-sensitive work — means more usage, not less compute. The METR horizon doubles every 3.5 months. Training is still 61% of global compute. Anthropic's IPO will funnel tens of billions into infrastructure. OpenAI has Jalapeño. Neither company is reducing GPU orders.

What to watch

Opus 5's ARC-AGI-3 score of 30.2% is the single most interesting number in this launch. It measures genuinely novel problem-solving — not memorized benchmarks — and Opus 5 triples the next best public model. If that gap persists when third-party evals arrive, it is a capability signal that changes the competitive landscape more than any SWE-bench score. Fable 5.1 is reportedly in late pipeline for August to compete with GPT-6. The generation is not done accelerating.

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.