Anthropic Launches Claude 5.5 Haiku, Beats GPT-6 Luna Across Most Benchmarks While Priced at 1/20th Of Sonnet 5.5

Anthropic has launched Claude Haiku 5.5, the small-model entry in its 5.5 family, and is pitching it as the cheapest, fastest and most capable small model it has ever released. The company says Haiku 5.5 costs around 75% less to run than Claude Haiku 4.5 on average, while posting large gains over its predecessor on coding, computer use and knowledge work benchmarks.

Haiku 5.5 is aimed at high-volume, cost-sensitive work: summaries, compaction, database queries and classification. Anthropic also says it works well as a subagent alongside Claude Opus 5.5 and Sonnet 5.5 on coding tasks, and that it’s fast enough for speed-sensitive uses like live customer support and browser use. It’s the last piece of the 5.5 lineup, following Sonnet 5.5, which Anthropic launched at the end of September and which beat OpenAI’s GPT-6 Sol on some benchmarks, and Opus 5.5 before that.

Haiku 5.5 is also the first Haiku-class model to come with an adjustable effort setting, which lets developers choose whether to optimize a given task for cost or for intelligence, as they can with Anthropic’s larger models.

Claude Haiku 5.5 benchmarks

Anthropic compared Haiku 5.5 against Haiku 4.5, OpenAI’s GPT-6 Luna, and, for reference, its own Sonnet 5.5. Here’s how the numbers stack up:

All figures are as reported by Anthropic. Knowledge work scores are Elo-style ratings; the rest are percentages.

A huge jump over Haiku 4.5

The generational leap is the headline. On the two knowledge work benchmarks, Haiku 5.5 more than doubles its predecessor: GDPval-AA goes from 735 to 1620, and AA-Briefcase from 614 to 1578. On OSWorld 2.1, which measures how well an agent can operate a real computer to complete long, multi-step tasks, the offline subset score rises from 15.7% to 72.4%.

The reasoning and vision numbers show similar jumps. On Humanity’s Last Exam, Haiku 5.5 scores 45.9% without tools and 57.4% with tools, against 10.2% and 18.7% for Haiku 4.5. On Chartography, a visual reasoning benchmark run without tools, it scores 46.4% against 6.4% for the older model, a roughly seven-fold improvement.

The starkest gap is in agentic coding. Haiku 4.5 scored 0.0% on Terminal-Bench 4.0, which tests how well a model can complete complex, multi-step professional tasks in a command-line interface. Haiku 5.5 scores 39.2%.

Ahead of GPT-6 Luna

Against OpenAI’s GPT-6 Luna, the one competitor model in Anthropic’s table, Haiku 5.5 comes out ahead on every benchmark where Luna has a reported score. It leads on GDPval-AA (1620 vs 1437), AA-Briefcase (1578 vs 1336), OSWorld 2.1 (72.4% vs 48.9%), Terminal-Bench 4.0 (39.2% vs 16.4%), FrontierCode 1.1 (46.4% vs 42.4%) and Chartography (46.4% vs 29.1%). Luna has no reported scores on the two Humanity’s Last Exam variants.

The FrontierCode margin is the narrowest of the set, at four points. The computer use and Terminal-Bench gaps are far wider.

How close does it get to Sonnet 5.5?

Sonnet 5.5 is included in the table purely as a reference point, and it remains ahead of Haiku 5.5 on every benchmark shown. The size of the gap depends heavily on the task type.

On knowledge work, Haiku 5.5 is relatively close: 1620 vs 1840 on GDPval-AA and 1578 vs 1824 on AA-Briefcase. On computer use, it reaches 72.4% against Sonnet’s 83.9%. On FrontierCode, the gap is under six points (46.4% vs 52.1%, with Sonnet’s score at xhigh effort).

Where the gap is widest is Terminal-Bench 4.0, where Sonnet 5.5 scores 70.6% to Haiku’s 39.2%. Anthropic is upfront about this. The company says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding, while Haiku 5.5 is best suited to more narrowly scoped tasks that might previously have been too expensive to run with earlier Claude models, such as compaction, summarization and subagent work.

Performance by effort level

Anthropic also published charts showing how Haiku 5.5 performs at each effort setting (low, medium, high, xhigh and max) on three benchmarks: OSWorld 2.1 (offline subset), GDPval-AA and Humanity’s Last Exam. The charts plot accuracy against cost per attempt on a log scale, with Haiku 4.5, Sonnet 5.5 and GPT-6 Luna alongside, so developers can see what they give up in accuracy for each step down in cost.

The company has made a similar point about Sonnet 5.5. On Terminal-Bench 4.0, Anthropic plotted that model’s performance relative to cost both before and after its cache read price cut (more on that below).

Claude Haiku 5.5 Pricing

Haiku 5.5 is priced per million tokens at $0.10 for input and $0.50 for output on prompts up to 100,000 tokens. For prompts over 100,000 tokens, the rates are $0.50 for input and $2.50 for output. Anthropic says the model is especially good value for prompts under 100,000 tokens, which make up around 90% of requests to the previous Haiku model.

Price per 1M tokensHaiku 5.5 (prompts up to 100k / over 100k)Haiku 4.5Sonnet 5.5
Cache reads$0.01 / $0.05$0.10$0.10
Cache writes$0.125 / $0.625$1.25$2.50
Input tokens$0.10 / $0.50$1.00$2.00
Output tokens$0.50 / $2.50$5.00$10.00

For prompts up to 100,000 tokens, Haiku 5.5 costs a tenth of what Haiku 4.5 did per token, and a twentieth of what Sonnet 5.5 costs for input and output. The 75% average saving Anthropic cites is lower than the per-token drop, since it reflects typical workloads rather than list prices alone.

Cheaper Sonnet 5.5 cache reads, and API credits for Max and Team

Alongside the Haiku launch, Anthropic is cutting the price of cache reads on Sonnet 5.5 by 50%, from $0.20 to $0.10 per million tokens. Because cache reads account for a large share of token consumption in agentic work, the company says this lowers the cost of running Sonnet 5.5 on most agentic tasks by around 20%.

Anthropic is also introducing a monthly API credit for Claude Max and Team subscribers, to be rolled out this week, for building agents and applications on the Claude Platform. Max 5x subscribers get $100 per month, Max 20x subscribers get $200, and Team subscribers get up to $500, pooled across their users. The credits can be used on any Anthropic model.

For developers, Anthropic is also updating its Python and TypeScript SDKs to add beta support for computer use and browser use, two areas where it says Haiku 5.5’s mix of speed, capability and price makes it a particularly good fit.

Claude Haiku 5.5 Availability

Claude Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure. Developers can access it on the Claude Platform using the model name claude-haiku-5-5.

The pricing and speed pitch matters more for Haiku than for most models. Small models tend to be the ones running in the background at enormous scale, summarizing, routing and classifying, and sometimes acting as the cheap workers beneath a more expensive orchestrating model. Anthropic is betting that a Haiku which now handles computer use and command-line work, at a fraction of the old price, will pull more of that volume onto Claude. How that plays out against the cost of competing small models, and against Anthropic’s own pricing on Sonnet 5.5, which also scored well on independent evaluations like the Artificial Analysis Intelligence Index, will be worth watching.

Posted in AI