Claude Haiku 5.5 Ships With a 75% API Price Cut
⚡ News

Claude Haiku 5.5 Ships With a 75% API Price Cut

Anthropic launched Claude Haiku 5.5 on October 7, 2026 with a ~75% average API cut, 1M context, an effort control, and a public Index score of 43.

The AI Dude · October 8, 2026 · 4 min read

Anthropic's cheapest Claude tier just got a new model and a stated average API price cut of about 75%. Claude Haiku 5.5 shipped on October 7, 2026 with a 1M-token context window and a new effort control, per the company's official launch announcement.

The same launch pairs that cut with the new effort parameter and the 1M-token context window on one Haiku SKU, the combo Artificial Analysis's X threads are stacking next to the Intelligence Index score of 43.

The cut also pushes a product decision out of the user's hands: for high-volume classification, extraction, and short tool loops, Haiku 5.5 becomes the default Claude SKU to price against, not an optional side door.

Haiku 5.5 exposes speed and depth as one effort control

Haiku 5.5 is Anthropic's small, fast Claude line refreshed for API work that needs low latency and low unit cost. The launch materials pair that positioning with two mechanical changes that matter more than the marketing name: a 1M-token context window, and an effort parameter that lets the caller dial how hard the model works on a given request.

Mechanically, effort is the interesting part. On lighter settings, the model is aimed at short, repetitive jobs where extra reasoning is wasted spend. On heavier settings, the same model ID is allowed to spend more compute (and usually more time and output tokens) on multi-step or ambiguous work. Effort stays on the same model ID and only changes how much of Haiku's budget a single call is allowed to burn.

The 1M context window is the other half of the mechanism. Long prompts, big RAG packs, and multi-file dumps can sit in one request without an immediate hop to a larger Claude tier for window size alone. That does not make the model free to fill. A million tokens of input still bills as a million tokens of input, and a high-effort pass over a long context is where the new rate card can stop looking cheap in practice.

Simon Willison covered the Haiku 5.5 launch on his blog after the announcement: simonwillison.net.

The cost the mechanism itself carries is straightforward. Effort is a spend dial with a quality promise attached. Leave it high on every call and the 75% average cut can shrink under heavy agent traffic. Leave it low on work that needs careful tool use and you save money while shipping brittle answers. Effort makes the quality-versus-spend tradeoff explicit on each request.

Artificial Analysis puts Haiku 5.5 at 43 on the Intelligence Index

The public number traveling with the launch on X and in analyst threads is Artificial Analysis's Intelligence Index score of 43 for Haiku 5.5, detailed in high-engagement October posts from the @ArtificialAnlys account at x.com/ArtificialAnlys. That Index is a composite across their published task mix. It is not a single coding contest, not a customer A/B, and not Anthropic's own scorecard.

Anthropic's headline around Haiku 5.5 is the price-and-speed story: a small model that stays useful while the average token price falls hard. Artificial Analysis's 43 is the independent figure people are screenshotting next to that story, often beside agentic benchmark charts from the same outlet. What that buys a reader is a third-party yardstick for "how smart is the cheap tier," not a guarantee that Haiku 5.5 will win your production harness.

Pitch decks and thread screenshots are already lining Haiku 5.5 up against GPT-6 Luna in the same price-and-latency band. Treat those side-by-sides as Index-and-chart comparisons under Artificial Analysis's methodology unless the post shows your own tasks. A composite 43 can move routing defaults without saying whether a refund classifier, a support draft, or a repo agent fails less often on your data.

The Index score is useful as a public anchor and weak as a purchase order. Teams that treat 43 like a pass/fail for coding agents are reading a leaderboard as if it were an SLA.

The 75% average attaches to the Haiku 5.5 API rate card Anthropic published

The price cut is attached to the Haiku 5.5 API SKU in Anthropic's launch materials, framed as an average reduction versus prior Haiku pricing rather than as a single flat new input/output pair reprinted in every secondary post. Artificial Analysis's pricing threads on X also describe tiered rates tied to how hard you run the model, which matches the effort control as a billing-relevant dial, not only a quality knob.

That is a different stakes frame from a premium-versus-volume buyer split. The live question is whether your traffic shape matches the mix inside Anthropic's "about 75%" average. Short, low-effort calls with modest output will feel the cut immediately in the invoice. Long-context, high-effort agent traces shift spend toward output tokens and toward the higher tiered rates @ArtificialAnlys described for harder runs. The model ID is cheap. The path you call it on may not be.

Haiku 5.5 is built for API traffic that wants Anthropic's cheapest Claude line with effort and 1M context on one SKU: high-volume classification, extraction, and short tool loops where public threads already price it against GPT-6 Luna. The launch story is that API rate card and the Index score of 43, not a free consumer chat tier.

Outsiders lining Haiku 5.5 up against GPT-6 Luna still cannot reproduce Anthropic's ~75% average from the Index score of 43 and the tiered rates alone.

Claude Haiku 5.5Haiku 5.5 pricingGPT-6 LunaAnthropic APIeffort parameter
Share 𝕏 / Twitter Reddit LinkedIn

Keep reading