Artificial Analysis has put Anthropic’s newly released Claude Haiku 5.5 through its Intelligence Index, and the small model lands at 43 at max effort. That’s a 26-point jump over Claude 4.5 Haiku, which scored 17 on the same index a year earlier, and it puts Haiku 5.5 ahead of several of the best-known small models, including Google’s Gemini 3.8 Flash (41) and OpenAI’s GPT-6 Luna (38). It sits one point behind Moonshot AI’s 2.8-trillion-parameter Kimi K3, which scores 44, and 13 points behind Claude Sonnet 5.5 (max), which scores 56.
The catch is token usage. At max effort, Haiku 5.5 uses roughly 162,000 output tokens per Intelligence Index task, about three times as many as GPT-6 Luna at max effort, and Artificial Analysis says it needs more tokens than Luna to reach similar scores across effort settings.

Where Haiku 5.5 sits on the Intelligence Index
Haiku 5.5 is the first Haiku model with Anthropic’s effort settings and adaptive thinking. Artificial Analysis tested it across all five effort levels, and its scores climb steadily with effort:
| Effort setting | Haiku 5.5 Intelligence Index | GPT-6 Luna Intelligence Index |
|---|---|---|
| Max | 43 | 38 |
| Xhigh | 41 | 35 |
| High | 38 | 33 |
| Medium | 34 | 30 |
| Low | 29 | 22 |
Scores from Artificial Analysis Intelligence Index v4.3.2, which incorporates 10 evaluations. Haiku 5.5 entries are listed as “with fallback” in Artificial Analysis’s charts.
At max effort, Haiku 5.5 is slightly ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41), DeepSeek V4.1 Flash (39) and GPT-6 Luna (38), and its score is comparable to Kimi K3’s 44. Haiku 5.5 at xhigh effort (41) still matches Gemini 3.8 Flash, and Haiku 5.5 at high effort (38) matches GPT-6 Luna at max. The model scores 43 at the top setting, but even at medium effort, its 34 is ahead of Luna’s high-effort 33.
For a sense of the gap to Anthropic’s larger models, Sonnet 5.5 at max effort scored 56 on the same index.
Heavy on tokens
Artificial Analysis’s main caveat concerns how many tokens Haiku 5.5 burns to get those scores. Here’s the weighted average of output tokens used per Intelligence Index task, by effort setting:
| Effort setting | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Max | ~162k | ~50k |
| Xhigh | ~89k | ~27k |
| High | ~55k | ~20k |
| Medium | ~33k | ~11k |
| Low | ~17k | ~2k |
For comparison, Claude 4.5 Haiku used about 18k tokens per task, though it scored only 17.
Moving from xhigh to max adds two points to Haiku 5.5’s score at a cost of roughly 1.8x the tokens, with most of the max-effort total (about 129k of the 162k) going to reasoning rather than the final answer. Artificial Analysis also notes that Haiku 5.5 at max uses more output tokens than Opus 5.5 at max.
The comparison with GPT-6 Luna is where the token appetite shows up most clearly. Haiku 5.5 (high) scores 38 using about 55k tokens per task, while GPT-6 Luna (max) scores the same 38 with about 50k. The gap widens at lower effort settings. Since output tokens are billed, this matters for the “cheapest model” pitch, though the effort setting gives developers a way to dial it back.
Pricing: same as Luna, with a catch above 100k
Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That’s the same as GPT-6 Luna and 10% of the price of the previous Haiku model. Above 100,000 tokens, pricing rises 5x to $0.50 and $2.50.

Artificial Analysis says its site doesn’t yet reflect the tiered pricing, so its provisional cost figures for Haiku 5.5 don’t include the step up. The firm says it’s working on support and will follow up with cost-per-task coverage soon, which will be the real test of how Haiku 5.5’s extra token use plays out against its low per-token price. Cache reads cost $0.01 per million tokens ($0.05 above 100k), and five-minute cache writes cost $0.125 ($0.625 above 100k).
Other model details Artificial Analysis lists: a 1-million-token context window, up from 200,000 for Claude 4.5 Haiku, and text and image input with text output.
Claude 5.5 Haiku Benchmark highlights
Agentic knowledge work. On AA-Briefcase, Artificial Analysis’s private evaluation of realistic knowledge work tasks, Haiku 5.5 (max) reaches 1578 Elo. That’s ahead of models including Kimi K3 and GLM-5.3, comparable to Muse Spark 1.3 (max), and within the confidence intervals of GPT-6 Astra (max) and Claude Fable 5.1 (high), two far larger and costlier models.
Terminal use. On Terminal-Bench 4.0, Haiku 5.5 scores 33% in Artificial Analysis’s testing, up from 0% for Haiku 4.5. That’s level with GLM-5.3 Flash and ahead of Gemini 3.8 Flash (20%) and GPT-6 Luna (13%). (Anthropic’s own launch figure for the benchmark is 39.2%.)
Factual knowledge and hallucinations. As expected for a smaller model, Haiku 5.5 trails on raw factual knowledge. Its accuracy on AA-Omniscience is 36%, against 55% for Gemini 3.8 Flash and 44% for GPT-6 Luna. Part of that gap comes from the model being more willing to admit when it doesn’t know. Its hallucination rate on the test is 40%, against 55% and 77% for those two models.
AutomationBench-AA. Haiku 5.5 scores just 35%, well below the 53–60% range for GPT-6 Luna, Gemini 3.8 Flash and GLM-5.3 Flash. Artificial Analysis says the result is likely understated: during pre-release testing, a safety refusal issue caused the model to over-refuse. Anthropic is working on a fix, and Artificial Analysis plans to re-run the evaluation once it lands and expects the score to rise.
The takeaway
Anthropic launched Haiku 5.5 as the cheapest and fastest model it has released, and the independent numbers back up the “most capable small model” part of that pitch: it tops the small-model field on Artificial Analysis’s index at max effort, handles agentic knowledge work at a level that rivals much larger models, and has made a huge leap on terminal tasks. The open questions are all about efficiency. Between its heavy token use and a pricing step-up above 100k tokens, the real cost of running Haiku 5.5 against rivals like GPT-6 Luna will only become clear once Artificial Analysis publishes its cost-per-task figures.