GeethanTechGeethanTechTech, decoded daily.
LatestAISoftware & PlatformsCybersecurityStartups & VentureTech Business & Markets
All posts
AIabout 1 hour ago

Claude Haiku 5.5 costs more above 100K prompt tokens

By megan

Anthropic’s new small model offers low API rates for shorter prompts, but requests above 100,000 prompt tokens move to a higher price tier. A move from Haiku 4.5 also requires checks for token counts, thinking settings and request compatibility.

Abstract illustration of a vertical threshold separating cool grey and coral blocks, representing a price change above 100,000 prompt tokens.

Anthropic released Claude Haiku 5.5 on October 7, 2026 for high-volume tasks. Developers can call it as claude-haiku-5-5. Its one-million-token context window allows much longer requests, but the lowest advertised API rate applies only when a prompt has no more than 100,000 tokens.

That boundary matters to teams processing long documents or growing agent histories. Moving from Haiku 4.5 also involves more than changing the model ID: Anthropic’s migration guide identifies request settings and response-handling assumptions that can fail.

The price changes with prompt length

Anthropic’s Haiku 5.5 price table lists these rates in US dollars per million tokens:

For prompts up to 100,000 tokens, Anthropic lists $0.10 per million input tokens and $0.50 per million output tokens.

For prompts over 100,000 tokens, the listed rates are $0.50 per million input tokens and $2.50 per million output tokens.

The listed input and output rates are five times higher in the longer-prompt tier. Cache reads and writes have their own two-tier rates. The threshold concerns prompt length, while the applicable tier also determines the listed output rate. A one-million-token context window is therefore a capacity limit, not a promise of the lowest rate throughout that window.

Anthropic says about 90% of requests to Haiku 4.5 were within 100,000 tokens. That describes Anthropic’s previous-model traffic, not the prompt distribution of a particular application. Teams should measure their own requests, especially those near the threshold.

The same text can consume more tokens

Haiku 5.5 uses a newer tokenizer. Anthropic says the same input text produces approximately 30% more tokens than on Haiku 4.5, although the increase depends on the content. That can move a request across the price boundary and alter estimates built from old token counts. It can also make an existing max_tokens value too small for a comparable response.

Adaptive thinking is on by default. Applications that send Haiku 4.5’s thinking: {"type": "enabled", "budget_tokens": N} must change that setting; Anthropic says the old form returns a 400 error on Haiku 5.5. Clients should also select response content blocks by type rather than assuming the first block contains the answer, because it may contain thinking.

Other compatibility checks are concrete: Anthropic instructs developers to remove temperature, top_p and top_k settings, and to replace final assistant-message prefills with another approach. Its guide says incompatible values or an assistant prefill can return a 400 error. Integrations using the older computer-use tool on the Claude API or Google Cloud need the newer toolset.

Test the workload, not just the model ID

A migration test should count tokens with Haiku 5.5, sample requests on both sides of the 100,000-token boundary, and compare cost per completed task rather than relying on a per-token headline rate. It should also exercise parsers, thinking-block handling, output limits and requests that used the deprecated settings. Organizations with a Haiku 4.5 Priority Tier commitment need a separate capacity plan because Anthropic says that tier is not supported for Haiku 5.5.

Anthropic lists Haiku 5.5 as available through its API and major cloud platforms. Teams buying through a cloud provider should verify that provider’s actual billing and availability before projecting costs. Anthropic’s published benchmarks and speed claims were not independently tested for this article.

Continue with this topicAI
Browse categoryarrow_forward
Continue reading

Related posts

An open ivory tray and a closed steel-lidded tray sit beside a loose green enamel tab on a pale textured surface.
AI1 day agoBy megan

Meta and Sierra propose OAuth-based protocol for personal AI agents

Personal Agent Protocol would let an agent discover a business’s services, start a guest or account session, and work through its website, APIs, or business agent. Sierra plans to publish the v0.1 specification later in October 2026, so the announced flow is not yet an implementable standard.

Read article
Mistral Large 4 enters API preview while its model weights remain unreleased
AI2 days agoBy megan

Mistral Large 4 enters API preview while its model weights remain unreleased

Mistral’s multimodal model is available through a public preview API, but developers must wait for downloadable weights before assessing self-hosting.

Read article
GeethanTech
GeethanTech
Tech, decoded daily.
Read
Latest postsAISoftware & PlatformsCybersecurityStartups & VentureTech Business & MarketsRSS feed
Connect
mailEmaillanguageWebsiteinLinkedInIGInstagramfFacebook@Xsmart_displayYouTube
From the publisher

Read all our publications: GeethanPost | GeethanTech

A chapter in the ElegantHumanity story

Powered byElegantArc