Products & Models · Oct 9, 2026
Claude Haiku 5.5 confirmed from the primary source: 75% cheaper per token, but it uses 3 times as many tokens as Luna
It now sits next to GPT-6 Luna on the price list at the same rates, but Artificial Analysis found the two use very different numbers of tokens per task. The only way to tell which model is cheaper is to measure it on real work, not on the price list
Rie Suzuki · Technology Editor

Key points
- We confirmed Haiku 5.5 on Anthropic's announcement page. For prompts under 100,000 tokens, it costs $0.10 per million input tokens and $0.50 per million output tokens, the same as GPT-6 Luna
- Artificial Analysis found that Haiku 5.5 uses about 3 times as many tokens per task as Luna. Since the per-token prices are the same, that works out to roughly 3 times the cost per task
- Above 100,000 tokens, per-token prices rise fivefold. A model that uses more tokens is more likely to build up long contexts and cross into that higher tier
Anthropic has officially released Claude Haiku 5.5 on its own announcement page (anthropic.com/claude-haiku-5-5, primary source). Our earlier coverage relied only on a VentureBeat report, and checks of the anthropic.com/news listing gave different answers depending on when we looked. Now that we have checked the announcement page itself, we treat the release as fact rather than as a report.
Its per-token price is about 75% lower than Haiku 4.5, which puts it at the same price as OpenAI's GPT-6 Luna. But matching on the price list is not the same as costing less in practice. Independent evaluator Artificial Analysis found (artificialanalysis.ai) that Haiku 5.5 uses about 3 times as many tokens as Luna to finish a task. Comparing per-token prices alone does not tell you which one is cheaper.
Per-token price: matched to Luna
For prompts under 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens, those rates rise to $0.50 for input and $2.50 for output. The lower-tier rates are exactly the same as GPT-6 Luna's $0.10 input and $0.50 output; Luna was released on September 22. Anthropic has matched OpenAI's prices again. On September 28, Sonnet 5.5 launched at $2 input and $10 output, and GPT-6.1 Sol came out the next day at exactly the same rates. In both the mid-size and the small tier, the two companies now charge the same and compete on what the models can do.
The size of the price cut needs care. Anthropic's own comparison puts Haiku 5.5 about 75% below Haiku 4.5. VentureBeat, in the report we cited earlier, described "a roughly 90% price cut." Both figures depend on which price tiers are compared and how they are combined, for example whether you look only at the under-100,000-token tier or include the higher tier too. Arguing over which calculation is correct does not help readers decide. What matters, as the rest of this article shows, is that the size of the price cut does not by itself set the cost per task.
Anthropic has also reportedly cut Sonnet 5.5's cache-read price to $0.10 per million tokens and made the model available on AWS, Google Cloud and Azure.
Token usage: about 3 times Luna for the same work
Artificial Analysis gives each model the same set of tasks and measures its score, the tokens it used and what it cost. In that testing, Haiku 5.5 used about 3 times as many tokens per task as GPT-6 Luna.
Because the two models charge the same per-token prices, the math is simple. Most of the cost comes from output tokens, so 3 times the tokens means roughly 3 times the cost per task. The models sit side by side on the price list, but give them the same job and Haiku 5.5's bill can come close to 3 times Luna's.
This kind of gap is not unique to Haiku 5.5. Gemini 4 Argon, which we covered earlier, scored 53 on the Intelligence Index, the same as GPT-6 Astra. Yet it produced about 62,000 output tokens per task, more than twice Astra's roughly 27,000. In September, Grok 4.7 scored 46, tying Xiaomi's MiMo-V2.6-Pro, but it cost Artificial Analysis about 24 times as much to run the index. With each new release, there are more examples of models with the same score whose token use differs severalfold. The price per million tokens now tells only half of the cost story.
The long-context tier: more tokens, more likely to hit the higher rate
Haiku 5.5's pricing has another easy-to-miss feature. Once a prompt goes above 100,000 tokens, both input and output prices rise fivefold.
This matters for agents running long-horizon tasks. Tool-call results and the output of earlier steps pile up in the prompt for each next turn. The more tokens a model uses per turn, the sooner it crosses the 100,000-token line, and after that it pays five times the per-token rate. Heavy token use and higher-tier pricing are not separate problems; they compound each other. The roughly 3x gap Artificial Analysis measured could grow wider on long-context work than on short-context work.
Moving from Haiku 4.5: will it actually be cheaper?
Users moving from Haiku 4.5 face the same question. A per-token price about 75% lower means the cost per task falls as long as the new model uses no more than 4 times as many tokens. Beyond 4 times, the savings disappear.
Neither of the two sources we checked says how many times more tokens Haiku 5.5 uses than Haiku 4.5. Models that are smarter than the previous generation tend to think longer and write more. That is exactly why the roughly 75% price cut should not be counted as savings until you have tested the model on your own workload.
The takeaway: compare the bill per task, not the price list
Haiku 5.5 is Anthropic's answer in the price war over small models. Amid its preparations to go public, the company matched the price of OpenAI's cheapest model, showing it will not give ground on the price list.
But Artificial Analysis's figures show that this match is only on the surface. If Haiku 5.5 uses 3 times the tokens at the same per-token price, its real cost will be close to 3 times Luna's. Paying more can still be worth it if the higher token use brings higher scores and less rework. But that can only be checked task by task.
What buyers of small models should do is clear. Stop lining up per-million-token prices. Run your own workloads, then compare the bill per task and the quality of the results. Haiku 5.5 has become the clearest sign yet that you can no longer choose a model from the price list alone.
Editorial cartoon
