Claude Opus 5 Becomes Top Model In The World On Artificial Analysis Intelligence Index, Beats Fable 5

Anthropic has just bested its previous best publicly-available model — and is offering it at half the price.

Claude Opus 5 has debuted at the top of the Artificial Analysis Intelligence Index with a score of 61, edging past Claude Fable 5’s 60 and pushing GPT-5.6 Sol into third place at 59. It’s a narrow gap on paper, one point separating first and second, but the context around that point is what makes this result notable.

Claude Opus 5 Artificial Analysis Intelligence Index

Opus 5 is scoring 61 as a standalone, fully available model that anyone can call through the API today. Anthropic beating its own currently-available flagship while that flagship is hobbled by regulation is one thing. Doing it at a price Anthropic says is half of what Fable costs per task is the part that actually matters for anyone deciding which model to build on.

Where the rest of the field sits

The index has gotten considerably more crowded over the past month. GPT-5.6 Sol lands at 59, just behind Fable’s fallback score, and Moonshot AI’s Kimi K3 follows close after at 57 — a genuinely strong showing for an open-weights model going up against the two most talked-about closed releases of the year. Grok 4.5 sits at 54, GLM-5.2 and Meta’s Muse Spark 1.1 are tied at 51, and Gemini 3.6 Flash rounds out the models scoring above 50 at an even 50.

Below that band, DeepSeek V4 Pro and MiniMax-M3 are tied at 44, Nvidia’s Nemotron 3 Ultra sits at 38, and gpt-oss-120b brings up the rear at 24. What stands out looking at the full board is how many labs have now cleared the 50-point mark. A few months ago that threshold separated the frontier from everyone else. Now six different labs have a model above it, and the gap between first place and fifth place has compressed to single digits.

The agentic knowledge work numbers stand out more than the headline score

The more decisive result in Artificial Analysis’s evaluation isn’t the Intelligence Index at all, it’s GDPval-AA v2, which measures how well a model performs real professional knowledge work using Artificial Analysis’s open-source reference agent harness, Stirrup. Opus 5 (max) scores 1861 Elo there, over 100 points clear of both Fable 5 and GPT-5.6 Sol, and more than 114 points ahead of Fable specifically. On AA-Briefcase, Artificial Analysis’s own proprietary benchmark for agentic knowledge work, Opus 5 scores 1720 Elo, a 146-point lead over Fable 5. Both results make Opus 5 the new leader in agentic professional work by a wide margin, which is a different kind of claim than edging out a rival by a single point on a composite index.

Coding tells a similarly strong story. Running inside Claude Code at xhigh effort, Opus 5 takes joint first place on the Artificial Analysis Coding Agent Index and posts the highest score of any model on SWE-Atlas-QnA. On Terminal-Bench v2.1, a test of agentic terminal use, Opus 5 hits 89% at max effort, putting it roughly in line with the current leader, GPT-5.6 Sol at xhigh effort.

Where Opus 5 doesn’t lead

Anthropic’s model isn’t ahead everywhere, and Artificial Analysis’s breakdown is fairly direct about where it falls short. On Humanity’s Last Exam, Opus 5 scores 53%, in line with Fable 5 rather than ahead of it. On CritPt, a frontier physics evaluation built by researchers at Argonne National Laboratory and the University of Illinois Urbana-Champaign, it also matches Fable 5 but trails GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra.

Factual knowledge is the more notable gap. On AA-Omniscience, Opus 5 improves 7 points over Opus 4.8 on accuracy, but still lags behind Fable 5 — a difference Artificial Analysis attributes to the models’ relative size classes. The trade-off gets more pointed once you look at how Opus 5 handles uncertainty: it answers more often instead of declining when unsure, and its hallucination rate on the same benchmark climbs 14 points to 50%. Artificial Analysis also flags that Opus 5’s efficiency gains only really show up at the high end of the intelligence-versus-cost curve — at lower effort settings, it sits just behind the GPT-5.6 family on cost per task for a given level of intelligence.

Effort settings, context, and pricing

Opus 5 ships with five effort settings — low, medium, high, xhigh, and max — and Artificial Analysis found that gap between them unusually wide. On GDPval-AA v2 alone, moving from low to max effort spans 407 Elo points, with output token usage climbing roughly 8x across that same range. That mirrors how GPT-5.6 Sol behaves across its own effort settings, and it means Opus 5 can be run either far cheaper or far more capable than models from other labs, depending entirely on how a developer configures it.

The model carries a 1 million token context window, matching Opus 4.8, and keeps the same $5 per million input / $25 per million output pricing as its predecessor. Cache writes carry the usual 25% premium at $6.25 per million tokens with a five-minute time to live, while cache hits get a 90% discount down to $0.50 per million tokens. Server-side fallback, the same mechanism keeping Fable 5 online through export-control restrictions, is supported on Opus 5 as well.

Why the price angle matters more than the point gap

A one-point lead at the top of a benchmark index rarely tells the whole story, and Anthropic knows this better than most labs on the chart. What’s changed with Opus 5 is the shape of the trade-off. Fable 5 was Anthropic’s answer to the question of raw capability, priced and gated accordingly. Opus 5 is being pitched as the model that gets you nearly all of that capability, running independently rather than through a fallback, at a fraction of the cost per task, while also taking outright leads on the agentic knowledge work and coding benchmarks that matter most for enterprise deployment. Given how tightly bunched the top of the Intelligence Index has become, cost per point of intelligence is quickly turning into the metric labs actually compete on, not the top-line score itself.

That’s a shift worth watching. When the field was spread out, being three or four points ahead meant something concrete. Now that six labs are clustered within roughly ten points of each other, the model that delivers a score of 61 at 26% lower cost than the model sitting at 60 is arguably making the more compelling case, regardless of which one technically tops the chart this week.

Posted in AI