Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · October 7, 2026 · concept

Claude Haiku 5.5frontier-modelsmulti-agentagent-harness

Claude Haiku 5.5 Is a Migration, Not a Model Swap

The new small Claude costs a tenth of Haiku 4.5 under 100K tokens. Five old habits now return a 400, and your prompts count about 30% heavier.

Anthropic's smallest model now beats the old one on a computer-use benchmark by more than 56 points, and it costs a tenth as much per input token. That combination usually arrives with a catch somewhere, and with Claude Haiku 5.5 the catch is not hidden in the fine print. It sits in a migration guide that labels change after change as "Breaking."

Anthropic released Claude Haiku 5.5 on October 7. The model ID is claude-haiku-5-5, with no date suffix. It has a 1M-token context window, 128K tokens of max output, adaptive thinking on by default, and an effort setting, the first Haiku to get one. Prompts up to 100K tokens cost $0.10 per million input tokens and $0.50 per million output tokens. Haiku 4.5 cost $1 and $5.

If you run multi-agent systems, that price is the headline. My position is that the headline is the least useful part of the launch. The useful part is the list of things that break when you change one string in your config, and what those breaks tell you about where this model belongs in your stack.

Why the swap is tempting, and why it fails

Small models live in the parts of an agent system nobody looks at: the summarizer that compacts context, the classifier that routes a ticket, the subagent that reads twenty files and reports back. Those calls run constantly. Cutting their price by 90% changes the economics of the whole design, and Anthropic is explicit that this is the target. The launch post positions Haiku 5.5 for summaries, compaction, database queries and classification, and as a subagent working alongside Opus 5.5 and Sonnet 5.5.

So the obvious move is to find every claude-haiku-4-5 in your codebase and replace it.

Here is what the migration guide says will happen next.

Sampling parameters return 400. Any non-default temperature, top_p or top_k fails. temperature must be 1 if you send it, top_p must be 0.99 (even 1 fails), any top_k fails, and sending both temperature and top_p fails. If your classifier pinned temperature: 0 for determinism, it now errors out.

Assistant prefill returns 400. A final assistant turn in messages fails, even with thinking off. Prefill was the classic trick for forcing JSON or skipping a preamble. The guide says to use structured outputs, or tools with enum fields for classification, and on Bedrock (which lacks structured outputs) to use tools.

Thinking budgets return 400. thinking: {"type": "enabled", "budget_tokens": N} is gone. You set thinking: {"type": "adaptive"} and control depth with output_config.effort.

The old computer-use tool returns 400. On the Claude API and Google Cloud, computer_20250124 fails. You move to computer_toolset_20260801, drop the old beta header, and change your loop to dispatch on each tool_use block's name and toolset_name instead of input.action.

Conversations must be append-only. If system, tools or earlier messages change between requests while you send thinking blocks back, the request returns a 400. The guide notes that accounts created before August 31, 2026 hit this only when a specific binding setting is configured, so older accounts and newer ones can behave differently on the same code.

None of these are exotic. Prefill and temperature: 0 are standard habits in classification code. A model-ID swap on a mature codebase will break somewhere, and the failures will land in the background jobs you watch least.

The price you see and the price you pay

The second surprise is quieter than an error code. The guide says Haiku 5.5 uses the tokenizer from Claude 4.7 and later, so "the same text counts as approximately 30% more tokens than on Claude Haiku 4.5."

Run the arithmetic on that. Under 100K tokens, input that cost $1.00 per million Haiku 4.5 tokens now costs about $0.13 for the same text ($0.10 times 1.3). That is still roughly 87% cheaper, and output works out the same way. The headline survives.

Above the line it changes. Anthropic's table prices prompts over 100K tokens at $0.50 input and $2.50 output, and the launch post's footnote calls that 50% cheaper than Haiku 4.5. With 30% more tokens for the same text, the same long prompt costs about $0.65 where it used to cost $1.00. That is closer to 35% cheaper than 50%, by my arithmetic, and the exact tokenizer increase depends on your content.

The line itself moves too. A prompt that measured 80K tokens on Haiku 4.5 measures around 104K on Haiku 5.5. Your compaction subagent, the one that eats long transcripts by design, can cross into the higher tier without a single word of input changing. That is the job Anthropic names first in the launch post, and it is the job most exposed to this.

Two more cost details are worth knowing before you budget. The migration guide says thinking tokens count toward max_tokens, so a limit tuned for Haiku 4.5 can end with stop_reason: "max_tokens" after a thinking block and before any text. And the model overview says the Batch API gives a 50% discount on input and output, which matters more for this model than for any other, because batchable background work is what it is for.

What the benchmark table is telling you

Anthropic's own table is unusually frank about where this model stops. On the OSWorld 2.1 offline subset, Haiku 5.5 scores 72.4%, against 15.7% for Haiku 4.5 and 83.9% for Sonnet 5.5. On Terminal-Bench 4.0 it scores 39.2%, where Sonnet 5.5 scores 70.6%. The launch post says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0."

Read those two rows together and you get a routing rule. Screen-driving, reading, summarizing and browsing are close enough to the bigger model to hand down. Terminal work is not. The migration guide adds a new browser use toolset (browser_toolset_20260801) for Haiku 5.5 on the Claude API and Google Cloud, which Haiku 4.5 never had, and that points the same way.

These are vendor numbers from Anthropic's own evaluations, so treat them as a starting hypothesis for your workload, not a verdict on it.

Put this into practice

The lowest-friction path is a short audit before any config change.

  1. Grep for the five breaks. Search your code for temperature, top_p, top_k, budget_tokens, computer_20250124, and any message list that ends with an assistant turn. Each hit is a candidate 400 (only temperature: 1 or top_p: 0.99 sent alone are accepted), and a trailing assistant turn always fails.
  2. Recount your real prompts. Take a sample of production prompts and count them with the model set to claude-haiku-5-5, as the guide recommends. Note how many cross 100K.
  3. Replace prefill with structure. Move JSON forcing to structured outputs or a tool with an enum field. Move preambles into the system prompt.
  4. Set effort per job. Default effort is medium. Classification and routing probably want lower; anything the model reasons through wants a higher max_tokens so thinking does not eat the answer.
  5. Handle refusals. The guide says Haiku 5.5 runs safety classifiers that can return stop_reason: "refusal" with no server-side fallback. Your loop needs a branch for it.
  6. Route by task shape. Send reading, summarizing, classifying and browser steps to Haiku 5.5. Keep terminal and code-execution steps on Sonnet or Opus until your own repeat-run numbers say otherwise.

If you only do one thing this week, do step 2. Every other decision depends on knowing what your prompts cost in the new tokenizer.

Where this falls short

Several gaps are worth naming plainly.

The migration guide does not list every available effort level; it gives medium as the example and links elsewhere. Priority Tier is not supported, so a team with a Haiku 4.5 Priority Tier commitment has to plan capacity separately. Thinking blocks only work in the account that produced them; blocks from other accounts are silently dropped and the request still succeeds, which is a debugging trap for anyone replaying conversations across environments. By default, thinking blocks also come back with an empty thinking field and only a signature unless you ask for "display": "summarized".

The browser use toolset is not listed for Bedrock in the guide, so teams on AWS should check before planning around it. And the benchmark numbers are Anthropic's own; the launch post points to the system card for methodology, and I could not find an independent reproduction at the time of writing.

The bigger limit is one the price hides. A cheap attempt is not a dependable one. If a subagent fails silently one time in five, running it at a tenth of the cost means you can afford more failures, not fewer. The savings are real only if you measure what a passing task costs, not what a call costs.

The move worth making

Haiku 5.5 will probably end up doing most of the token-heavy work in a lot of agent systems, and it should. The teams that get there cleanly will be the ones that treat October 7 as the start of a migration: an audit, a recount, a routing table, and a test that runs the same task more than once.

Open your config. Before you change the model string, search for temperature. That one grep will tell you how long this migration really is.

Sources: Anthropic, Introducing Claude Haiku 5.5; Claude Haiku 5.5 model overview; Haiku 4.5 to Haiku 5.5 migration guide.


Medium metadata

  • Title: Claude Haiku 5.5 Is a Migration, Not a Model Swap
  • Subtitle: The new small Claude costs a tenth of Haiku 4.5 under 100K tokens. Five old habits now return a 400, and your prompts count about 30% heavier.
  • Tags: Claude, AI Agents, Anthropic, LLM, Software Engineering
  • Canonical: fervorai.dev