Anthropic Releases Claude Sonnet 5.5, Beats GPT-6 Sol On Some Benchmarks

Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family following Claude Opus 5.5 earlier this month. The company says Sonnet 5.5 is a clear upgrade over Sonnet 5, running more than 30% faster and costing up to 30% less for most workloads, while slotting in as a cheaper, quicker complement to Opus 5.5 for well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets. A Haiku 5.5 model is expected to round out the family in the coming weeks.

Claude Sonnet 5.5 benchmarks

On benchmarks, Anthropic says the jump from Sonnet 5 is dramatic. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, a test of agentic command-line work, up from Sonnet 5’s 10.3% and actually ahead of Opus 5.5’s reported 66.4%. It also comes in just two points behind Opus 5.5 on GDPval-AA, Anthropic’s benchmark for real-world work across dozens of occupations, and the company notes it is the first Sonnet model to beat Pokémon Red using only screenshots as input.

Claude Sonnet 5.5 benchmarks

Against OpenAI’s GPT-6 Sol, Anthropic’s own charts put Sonnet 5.5 ahead on some measures, including GDPval-AA and AA-Briefcase, while Sol edges out Sonnet 5.5 on FrontierCode’s coding benchmark. The comparison echoes the framing Anthropic used when it said Opus 5.5 had built a lead over GPT-6 Astra on the Artificial Analysis Intelligence Index.

Claude Sonnet 5.5 pricing

Pricing for Sonnet 5.5 stays the same as its predecessor at $2 per million input tokens and $10 per million output tokens, but Anthropic says the model typically needs far fewer tokens to complete the same task, cutting real-world costs by up to 30%. The company is also pointing to design as a strength: early testers said Sonnet 5.5 adds polish to interfaces and can follow a slide template closely enough that decks need minimal editing, with one internal test producing a 10-slide earnings review that two reviewers judged ready to send without changes.

Enterprise customers cited in Anthropic’s announcement echoed that efficiency angle. Slack’s Principal Engineer Curtis Allen said the model improved results on the company’s offline evaluations for its Slackbot product using around 14% fewer output tokens without any prompt changes, while Epic Games credited the model with holding up on system design and data-flow reviews across large gameplay codebases. It’s a similar customer-facing rollout to how Anthropic introduced Opus 5, which it said beat Fable 5 on many benchmarks at roughly half the price when it launched in July.

On safety, Anthropic says Sonnet 5.5 improves on or matches Sonnet 5 across most measures in its automated behavioral audit, a roughly 1,850-scenario test of alignment and honesty, and comes close to Opus 5.5 on newer containment evaluations that check how often models try to escape their sandboxes. Because its cybersecurity capabilities are now comparable to Sonnet 5’s, Anthropic is deploying Sonnet 5.5 with the same kind of cyber safeguards and fallbacks it built for its most capable models — meaning some higher-risk cybersecurity requests will be routed to Sonnet 5 instead, while Anthropic says routine software development is unaffected. Biology safeguards remain unchanged from Sonnet 5. The model also ships with new classifiers meant to block distillation attacks, where large numbers of fake accounts are used to extract a model’s capabilities into an unsafeguarded copy.

Claude Sonnet 5.5 is available now across the Claude apps, the Claude Platform under the model name claude-sonnet-5-5, and on AWS, Google Cloud, and Microsoft Azure.

Posted in AI