Technology3 min read

Claude Haiku 5.5

By · Published by Everything Blog

In short

The model uses the same newer tokenizer as Claude 4.7 and later models, making it slightly more token-intensive. The model is available on Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS.

Key points

  • Claude Haiku 5.5 is designed for high-volume, latency-sensitive tasks.: Claude Haiku 5.5 is designed for high-volume, latency-sensitive tasks.
  • It supports adaptive thinking with a 1M token context window.: It supports adaptive thinking with a 1M token context window.
  • Pricing starts at $0.10 per token and includes input and output tokens with a 50% discoun…: Pricing starts at $0.10 per token and includes input and output tokens with a 50% discount.

For high-volume, latency-sensitive tasks such as classification, extraction, and routing

- 1M

- 128K

- From $0.10

- From $0.50

Overview

Claude Haiku 5.5 is built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. It supports adaptive thinking with the effort parameter, a 1M token context window, and up to 128k output tokens. It uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5. Its thinking blocks work only in the account that produced them, or in an account linked to it.

For code changes, see the migration guide. For model IDs, pricing, and limits, see the Claude Haiku 5.5 overview. For prompting guidance, see Prompting Claude Haiku 5.5.

How it compares

Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |

---|---|---|---|---|---|---|---|

Claude Fable 5.1 | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | high | Jun 2026 |

Claude Opus 5.5 | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | medium | Jun 2026 |

Claude Sonnet 5.5 | 1M | 128K | $2 / $10 | Fast | Adaptive | high | Jun 2026 |

Claude Haiku 5.5 | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | medium | Jun 2026 |

Specifications

Model IDs

Pricing

- Input

- $0.10 / MTok $0.50 / MTok

- Output

- $0.50 / MTok $2.50 / MTok

- 5m cache write

- $0.125 / MTok $0.625 / MTok

- 1h cache write

- $0.20 / MTok $1 / MTok

- Cache read

- $0.01 / MTok $0.05 / MTok

- Batch API

- 50% discount on input and output

Capabilities

- Context window

- 1M tokens

- Max output

- 128K tokens

- Max output (Batch API, beta)

- 300K tokens

- Thinking

- Adaptive

- Default effort

medium

- Comparative latency

- Fastest

- Input → output

- Text and images → text

- Reliable knowledge cutoff

- Jun 2026

- Training data cutoff

- Jun 2026

Availability

- Status

- Active (latest)

- Released

- October 7, 2026

- Retirement

- Not sooner than October 7, 2027

- Platforms

- Claude APIAmazon BedrockGoogle CloudMicrosoft FoundryClaude Platform on AWS

Good to know

- Adaptive thinking is on by default. Control thinking depth with the effort parameter.

- Omit

temperature

,top_p

, andtop_k

, since a non-default value for any of them returns a 400 error. - On the Message Batches API, Claude Haiku 5.5 supports up to 300k output tokens with the

output-300k-2026-03-24

beta header. - Query limits and capabilities programmatically with the Models API.

Resources

Behavioral differences and prompting patterns specific to Claude Haiku 5.5.

Choose a model and effort level, shape prompts, and stream output for faster responses.

Claude Haiku 5.5 decides when and how much to think. Steer depth with effort

.

1M tokens. How the window is counted and managed.

Reference

The system prompt Claude Haiku 5.5 uses on claude.ai and the Claude apps.

Safety evaluations and deployment decisions for Claude Haiku 5.5.

Full price list, including batch discounts and prompt caching rates.

How model IDs, aliases, and pinned snapshots work.

Lifecycle status and retirement commitments for every Claude model.

Was this page helpful?

Original source: platform.claude.com

Technology