Claude Sonnet 5.5 Scores 56 On Artificial Analysis Intelligence Index, 3 Points Ahead Of GPT-6 Astra

Anthropic has opened a commanding lead between itself and the rest of the AI field.

Anthropic’s newly launched Claude Sonnet 5.5 has scored 56 points on the Artificial Analysis Intelligence Index at max effort, landing just two points behind Claude Opus 5.5 at the top of the leaderboard and three points ahead of OpenAI’s GPT-6 Astra, which sits tied with Claude Fable 5.1 at 53. That puts Sonnet 5.5 in second place overall, an 18-point jump over Sonnet 5’s score of 38, though Artificial Analysis notes the model gets there by generating the highest volume of output tokens per task it has measured on any model to date.

The index rolls up ten evaluations, including AA-Briefcase, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, and Humanity’s Last Exam, into a single composite score. On several of the individual components, Sonnet 5.5 essentially matches Opus 5.5 outright: it posts a GDPval-AA Elo of 1844 against Opus 5.5’s 1846, an AA-Briefcase Elo of 1811 against 1822, and actually edges ahead on AutomationBench-AA with a 71% headline score versus Opus 5.5’s 70%. On Terminal-Bench 4.0, a test of agentic command-line work, Sonnet 5.5 scores 63.6%, ahead of both Opus 5.5 (59.6%) and GPT-6 Astra (59.1%) — a roughly 50-point jump over Sonnet 5’s own Terminal-Bench score of 14.1%. On the newer Terminal-Bench-Science benchmark, which tests agentic research workflows and isn’t yet part of the Intelligence Index, Sonnet 5.5 scores 53.3%, trailing only GPT-6 Astra (63.3%) and Opus 5.5 (59.0%).

The catch is how much compute it takes to get there. At max effort, Artificial Analysis measured Sonnet 5.5 using roughly 193,000 output tokens per Intelligence Index task — about 60% more than either Opus 5.5 or Sonnet 5 at their own max settings, and close to seven times what GPT-6 Astra uses at max. That verbosity shows up directly in cost: Anthropic has kept Sonnet 5.5’s list price identical to Sonnet 5’s at $0.20 per million tokens for cache reads, $2 per million input tokens, and $10 per million output tokens, matching GPT-6 Sol’s pricing. But because it runs so many more tokens per task, Artificial Analysis calculates an effective cost of $7.60 per Intelligence Index task at max effort, about 50% higher than Sonnet 5.

That combination leaves Sonnet 5.5 sitting off the Pareto frontier that plots intelligence against cost per task. At high effort, it lands just behind GPT-6 Sol for roughly the same price, which Artificial Analysis flags as its most cost-competitive setting. At max effort it trails Opus 5.5 on the same frontier, and at lower effort levels, GPT-6 Astra and GPT-6 Sol configurations deliver comparable Intelligence Index scores for less money. Five effort levels are available in total — low, medium, high, xhigh, and max — and Artificial Analysis ran its evaluations across all of them with Anthropic’s default fallback behavior enabled. Sonnet 5.5 fell back to Sonnet 5 in about 0.1% of tasks across the index, almost entirely on Terminal-Bench 4.0 runs.

Sonnet 5.5’s gaps against Opus 5.5 show up more clearly outside agentic and knowledge-work tasks. On AA-Omniscience, a factual knowledge benchmark, it scores 54% on accuracy against Opus 5.5’s 66%, though it also hallucinates less often, at a 47% rate versus Opus 5.5’s 59%. It also runs about six points behind Opus 5.5 on both Humanity’s Last Exam and SciCode, evaluations that lean more heavily on scientific reasoning — a gap that lines up with how Anthropic has continued to position Sonnet as the faster, lower-cost option built for everyday and well-scoped work rather than the most demanding research-grade tasks that it reserves for the Opus tier.

Artificial Analysis notes one caveat on its numbers: the evaluations ran on a pre-release deployment of Sonnet 5.5 that Anthropic later found had a bug degrading responses to requests using structured outputs. Anthropic says that’s fixed in the public release, and Artificial Analysis expects the impact on scores to be minimal or, if anything, to have understated Sonnet 5.5’s actual performance — it plans to re-run the affected evaluations to confirm.

Outside of the benchmark results, Sonnet 5.5 keeps the same 1 million token context window as Sonnet 5, with support for both image and text input, and the same $2.50 per million token cache-write price. The launch follows Opus 5.5’s own run at the top of the index, where it had opened up a lead over GPT-6 Astra shortly after release, and suggests Anthropic’s second-tier model is now closing that gap from below rather than simply trailing it.

Posted in AI