Fictionally Irrelevant.

Haiku 5.5, First Impressions: Cheaper, Smarter, and a Few Surprises

Cover Image for Haiku 5.5, First Impressions: Cheaper, Smarter, and a Few Surprises
Harshit Singhai
Harshit Singhai

First Impressions

Anthropic released Claude Haiku 5.5 on October 7, 2026, and my first reaction was that the "small model" label is starting to undersell it. It's the fastest model in Anthropic's lineup, aimed at classification, routing, extraction, and subagent work, but it now comes with adaptive thinking, vision, and a 1M-token context window. It also costs a lot less than Haiku 4.5 did.

The price is the first thing people notice, and it's real. But after reading through the docs and early coverage, my honest take is that Haiku 5.5 is cheaper, smarter, and full of small surprises that matter once you start building with it.

Model Input (per 1M tokens) Output (per 1M tokens) Context
Claude Haiku 4.5 $1.00 $5.00 200K flat
Claude Haiku 5.5, prompts up to 100K $0.10 $0.50 1M
Claude Haiku 5.5, prompts over 100K $0.50 $2.50 1M

Haiku 5.5 is the only current Claude model with tiered pricing by prompt length. Anthropic says the lower tier covers about 90% of requests to the previous Haiku model, which is why the headline number is a 90% cut.


Smarter Than a Small Model Should Be

The capability jump is the part I didn't expect. Anthropic reports 72.4% on the offline subset of OSWorld 2.1, a computer-use benchmark, compared with 15.7% for Haiku 4.5. Those are the company's own numbers. Still, the gap is big enough to suggest Haiku 5.5 can handle tasks that previously needed a larger model.

Adaptive thinking is on by default, so the model decides when and how much to reason. You can turn thinking off for simple requests, and effort levels let you trade reasoning depth for tokens on each task. That flexibility is the feature I'd most like to test myself, because it means one model can cover both quick classification and harder multi-step work.


Tokens Cost More Than They Used To

The first thing that caught my attention in the docs was the tokenizer. Haiku 5.5 uses the newer tokenizer that also ships with Claude 4.7 and later. Anthropic's migration guide says cost estimates based on Haiku 4.5 token counts need to be recomputed. Based on the reporting I found, the same input becomes roughly 30% more tokens. I'd check output token counts on your own traffic rather than assume the same inflation applies there.

Here's how I think about it. Take a sentence that costs 100 tokens on Haiku 4.5. On Haiku 5.5, the same sentence becomes about 130 tokens. But each token costs one-tenth as much on short prompts, so the cost works out to 130 Γ— $0.10 = $13 per million sentences, compared with 100 Γ— $1.00 = $100 on Haiku 4.5. That's about 87% cheaper. The tokenizer takes a few points off the 90% cut, but the price cut is much larger than the inflation.

Anthropic's blended estimate, which weights the two price tiers by real traffic and includes the tokenizer effect, is that the average workload costs about 75% less than on Haiku 4.5. For mixed workloads, that's the number to plan around. Short-prompt workloads will land closer to 85–90% savings.


The 100K Cliff

The price tier is the trap to watch. Once a prompt crosses 100K tokens, the higher rate applies: $0.50 input and $2.50 output, five times the short-prompt rate.

The new tokenizer makes this easier to hit. 100K tokens on the new tokenizer is roughly 77K tokens on the old one, so a prompt that used to sit comfortably under the line can now cross it.

Even above the cliff, Haiku 5.5 is still cheaper than Haiku 4.5. Long-prompt input is $0.50 against $1.00, and after the tokenizer inflation the net is about 35% cheaper. But the savings are far smaller than what short prompts get. My practical rule is to keep each request under 100K tokens where possible, especially in high-volume pipelines. The docs also don't say whether cached tokens count toward the threshold, so test that before relying on it.


Prompt Caching Got More Useful

Caching is where Haiku 5.5 becomes especially attractive for agentic and RAG workloads. Cache reads cost 10% of the base input price, about $0.01 per million tokens. Cache writes carry a premium, so caching only pays off when you reuse the prefix.

Haiku 5.5 also lowers the minimum cacheable prompt size to 512 tokens, down from 4,096 on Haiku 4.5. Smaller system prompts, tool definitions, and reference snippets can now be cached, which matters a lot for subagent fleets where each agent carries the same instructions.


How It Stacks Up Against Nova 2 Lite

Comparing Amazon Nova 2 Lite with Haiku 5.5 in terms of cost. On Bedrock, it's listed at about $0.30 per million input tokens and $2.50 per million output tokens. Haiku 5.5 is cheaper on both sides for short prompts: $0.10 input is about a third of Nova's input price, and $0.50 output is about a fifth of Nova's output price. Even after the roughly 30% tokenizer inflation on Haiku 5.5, short-prompt input still comes out cheaper. Above 100K tokens, the picture flips. Haiku 5.5 input rises to $0.50, which is more than Nova's $0.30, while output matches at $2.50. So for long-context input, Nova 2 Lite may be the cheaper option.


The Bedrock Limitation

This one caught my attention right away. Structured outputs, meaning output_config.format and strict tool use, are available on the Claude API but not on Amazon Bedrock for Haiku 5.5. Anthropic lists this as a breaking change for teams migrating from Haiku 4.5.

On Bedrock, you have two options. You can describe the output format in the prompt and validate it in code, or use a tool without strict and validate the arguments with Pydantic. The second is more reliable.

Passing a Pydantic class as response_format through LiteLLM is the risky path. Depending on how LiteLLM resolves the model's capabilities, you'll get either a "does not support response_format" error before the request leaves the proxy, or a 400 from Bedrock.


Where It Fits: Agentic RAG

Haiku 5.5 looks like a strong fit for agentic RAG, where the model decides what to retrieve, grades what comes back, and rewrites queries when results are poor. Several roles map well to a small, cheap model:

  • Query router: decide whether a question needs retrieval, a database lookup, or a direct answer.
  • Query rewriter: turn a vague question into one or more search queries.
  • Relevance grader: judge each retrieved chunk. This is a classification task, which suits a small model.
  • Subagent: handle one retrieval branch in parallel and return a summary.

A stronger reasoning model can handle final synthesis if you need it, while Haiku 5.5 takes the high-volume steps. Two design rules matter here. Keep each call's context under 100K tokens, or the price tier changes. And put stable instructions and tool definitions at the start of the prompt so they can be cached. Retrieved chunks change per query, so the cached prefix should be everything before them.


Final Read

Haiku 5.5 is the most interesting small model Anthropic has shipped so far. It's much more capable than its size suggests, it's far cheaper per token, and it's built with agentic work in mind. The surprises are in the fine print: the tokenizer inflates counts, the 100K price cliff is easy to cross, and Bedrock doesn't support structured outputs yet.

My initial verdict is that it's worth a serious look for high-volume and agentic workloads. Before you migrate, count tokens on Haiku 5.5 rather than reusing old numbers, keep prompts under 100K where you can.

I will share more details once I switch from Haiku 4.5 and start using Haiku 5.5 in Production.