Grok 4.6 at $2/$6: Frontier Intelligence at Workhorse Pricing
It matches GPT-5.6 Sol on the Intelligence Index and leads on GDPVal-AA v2, but the real disruption is cost per task, not absolute SoTA
Grok 4.6 is not the smartest model. It may be the most disruptive on price.
xAI released Grok 4.6 on Aug 12, and the headline is not a new high score. It is the same score for a lot less money. Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and sitting one point behind Claude Fable 5 at 62 and two behind Claude Opus 5 at 63, while holding pricing flat at $2 per million input tokens and $6 per million output tokens.
That pricing is 80% cheaper on input and 88% cheaper on output than Fable 5 at $10/$50, and 60% and 76% cheaper than Opus 5 at $5/$25. For agentic workloads where output tokens dominate cost, the gap is what matters. Artificial Analysis measured Grok 4.6 at $0.84 per task on its Intelligence Index suite, the same as Kimi K3 and on the Pareto frontier for cost versus intelligence.
What Gavin Baker got right, and what he rounded
Investor Gavin Baker framed Grok 4.6 as matching Fable 5 Max at an 85% discount on input and 88% on output. The output math is exact. The input math is 80%, not 85%. Blended at a typical 75% input / 25% output agent mix, the discount is 85% versus Fable 5 and 70% versus Opus 5.
The point stands. At $3.00 blended versus $20.00 for Fable 5 and $10.00 for Opus 5, Grok 4.6 delivers 98.4% of Fable 5's Intelligence Index score for 15% of the price. Cost per Index point is $0.049 for Grok versus $0.323 for Fable 5, about 6.6 times cheaper.
Pricing is unchanged from Grok 4.5, which is unusual. Frontier gains usually come with price hikes. Grok 4.6 adds 5 points on the Intelligence Index in just over a month, and 23 points since Grok 4.3, without raising the sticker price. Cache hits did tick up to $0.50 per million from $0.30 on Grok 4.5, per Artificial Analysis.
Sources: xAI Grok 4.6 launch post for $2/$6 and 500K context; Claude Platform pricing for Fable 5 $10/$50 and Opus 5 $5/$25; Artificial Analysis for $0.84 per task and Pareto frontier.
Benchmarks: SoTA where it counts for agents, second place where it counts for hard code
xAI's launch post says Grok 4.6 "matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks." Artificial Analysis confirms the 61 score and adds context: Grok 4.6 is back on the frontier alongside OpenAI, behind only Anthropic.
The wins are concentrated in agentic knowledge work:
| Evaluation | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 (Elo) | 1753 | 1526 | 1728 | 1741 |
| AA-Briefcase (Elo) | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| FrontierCode v1.1 Extended | 61.3% | 56.6% | 60.6% | 64.9% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
GDPVal-AA v2 at 1753 is the standout. Artificial Analysis calls it "behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max." On AA-Briefcase, its private benchmark for long horizon knowledge work, Grok 4.6 sits at Fable 5 tier with an Elo of 1577, behind the Opus 5 family.
The losses matter too. On DeepSWE v1.1 Grok trails Sol by 7.1 points, and on Terminal-Bench v3.0 it trails by 8.6 points. Those are hard coding and terminal use tasks where Fable 5 and Sol still lead. Synthwavedd summarized it accurately: near SoTA on the Intelligence Index, SoTA on GDPVal-AA v2, AA-Briefcase and Harvey LAB, and second only to Fable 5 on CursorBench, FrontierCode and APEX — see Synthwavedd table.
Are benchmarks cherry picked? Partly. xAI headlines the four wins and does not headline the two clear lags. The composite Intelligence Index hides that dispersion. If your workload is long running research, analysis across a codebase, or turning an idea into a polished artifact, the SoTA claims are relevant. If it is competitive coding or terminal heavy SWE, Fable 5 and Sol remain ahead.
*Elo scale for GDPVal and Briefcase; % for others. Source: xAI and Artificial Analysis.
CursorBench and latency: where price becomes product
For Cursor users, CursorBench v3.2 is the real world test. Grok 4.6 High scores 69.9% versus Fable 5 Max at 70.5%, a 0.6 point gap. The cost gap is larger. Cursor's leaderboard lists Grok 4.6 Extra High at 70.8% for $2.81 versus $17.32 for Fable 5 Max and $8.23 for Opus 5 Max. Grok 4.6 Extra High actually edges Fable 5 at 70.8% for about one sixth the cost.
One caveat carries over from Grok 4.5. Cursor disclosed that an older snapshot of the Cursor codebase was accidentally in Grok 4.5's training data, giving it an advantage on CursorBench that Cursor said it could not fully measure. xAI says that data was removed for future runs, but treat 4.6's CursorBench score as directionally strong rather than definitive until independent retests land.
Latency reinforces the value case. Artificial Analysis measures Grok 4.6 High at 85.8 tokens per second on SpaceXAI's own endpoint, versus Grok 4.5 at 80 tokens per second. Frontier peers typically test at 30 to 50 tokens per second in the same harness. For agentic loops that make dozens of tool calls, speed is cost. Grok 4.6 also shows turn efficiency on AA-Briefcase: about 53 turns and 0.5 billion input tokens on average versus about 103 turns and 2.0 billion for Opus 5 Max to reach a comparable answer. Fewer turns means less accumulated context to pay for.
Context window stays at 500K tokens, unchanged from Grok 4.5.
Why this pricing matters beyond one model
This is the same pattern we have seen when a frontier capable model reprices the market. DeepSeek's $0.14 per million pricing forced a conversation about sustainable inference economics, and Nvidia's own commentary has noted that closed model inference still drives the spend, but price per token sets who can afford to run agents at scale.
Grok 4.6 does not win by being the smartest. It wins by being smart enough at a price where you can afford to let it run for 50 turns. That is a genuine disruption to the pricing model for agentic coding and knowledge work, even if it is not a disruption to the capability frontier itself.
What to watch
- Independent Artificial Analysis retests and CursorBench reruns without any training data overlap, to confirm the 69.9% to 70.8% range holds.
- Whether xAI holds $2/$6 under load. Rate limits are listed at 150 requests per second and 50 million tokens per minute, but sustained agentic use will test that.
- Grok 4.7, which Baker notes will be a much larger model with Cursor and SpaceX data in pretraining. If xAI can keep pricing flat again, the Pareto frontier moves further.
- How Anthropic and OpenAI respond on output pricing. Output tokens dominate reasoning heavy workloads, and Grok's $6 versus $25 to $50 is where the 70% to 88% gap bites hardest.
More research at bargo.ai/research.
Sources
- xAI Grok 4.6 launch post — pricing, context, training details
- Artificial Analysis Grok 4.6 benchmarks and analysis — Intelligence Index, GDPVal-AA v2, cost per task, turn efficiency
- Claude Platform pricing — Fable 5 $10/$50, Opus 5 $5/$25
- Synthwavedd benchmark table
- CursorBench leaderboard