Comrades, Anthropic is back with another price cut—this time targeting Haiku, the most affordable model in the Claude family.
On October 7, Anthropic officially launched Claude Haiku 5.5, claiming it to be their fastest, most capable, and lowest-cost small model to date.
Most striking is the price: as low as $0.1 per million input tokens, a 90% reduction compared to the previous generation.

Performance hasn't fallen behind. According to Anthropic's published test results, Haiku 5.5 outperforms OpenAI's GPT-6 Luna across multiple benchmarks, including computer operations, knowledge work, and agent programming, with some results showing significant margins.
Moreover, Haiku 5.5 introduces adjustable reasoning intensity, supporting a context window of up to 1 million tokens and a maximum output of 128,000 tokens.
Thus, in just 15 days, Anthropic has completed the full Claude 5.5 family update: Opus 5.5 launched on September 22, Sonnet 5.5 was released on September 28, and now Haiku 5.5 has officially joined the lineup.
This time, Anthropic focused on AI tasks with extremely high usage volumes and high price sensitivity.
Outperforms GPT-6 Luna on multiple metrics, with a computer operation success rate of 72.4%.
According to Anthropic's published benchmark results, Haiku 5.5 shows significant improvement over its predecessor and even outperforms GPT-6 Luna in some tests.

The most notable achievement came from OSWorld, a benchmark that evaluates an AI agent’s ability to operate a real computer by requiring the model to complete multi-step, cross-application tasks.
On the OSWorld 2.1 offline task subset, Haiku 5.5 achieved a 72.4% success rate, compared to 15.7% for the previous generation, Haiku 4.5, and 48.9% for GPT-6 Luna.

In knowledge work, Haiku 5.5 scored 1620 on GDPval-AA v2.1, surpassing GPT-6 Luna’s score of 1437. In the Agent programming test Terminal-Bench 4.0, the scores were 39.2% and 16.4%, respectively.
However, these data include an important detail: Haiku 5.5’s highest benchmark scores often require additional inference computation.
Haiku 5.5 introduces adjustable inference intensity for the first time in the Haiku series, allowing developers to control the amount of computational resources the model uses via the effort parameter.
VentureBeat noted that in Anthropic’s Terminal-Bench test, a score of 39.2% corresponds to maximum reasoning intensity. When switched to medium reasoning intensity, the score drops to approximately 20%, which is the default setting for Haiku 5.5.
The third-party evaluation agency Artificial Analysis also observed a similar trend. In its Intelligence Index assessment, Haiku 5.5 achieved a maximum reasoning score of 43 and a medium reasoning score of 34. Under the same evaluation, the average cost per task was approximately $0.21 and $0.05, respectively.
In other words, the model can achieve better task performance by increasing its inference budget, but the actual cost may also rise significantly.
In its own Terminal-Bench 4.0 tests, Artificial Analysis achieved a maximum inference strength of 33% and a medium strength of 15%. Due to differences in evaluation settings, these figures differ from Anthropic's official results.
This is why, when evaluating the performance of small models, it's important to consider inference intensity, task completion rate, and actual cost simultaneously.
Sonnet 5.5 still demonstrates a clear advantage on more complex agent programming tasks, achieving a Terminal-Bench 4.0 score of 70.6%.
Anthropic also explicitly stated that Sonnet and Opus remain better suited for complex agent programming tasks, while Haiku's strengths are focused on high-frequency, well-defined tasks.
Price slashed directly to 10% off, competing head-to-head with GPT-6 Luna
If performance improvements made Haiku 5.5 competitive, then pricing may be the most important highlight of this release. Anthropic has set two pricing tiers for the new model.

For requests with a prompt length of up to 100,000 tokens, the input and output prices for Haiku 5.5 are reduced by 90% compared to the previous generation. Beyond this threshold, prices are still 50% lower than those for Haiku 4.5.
Anthropic stated that 90% of requests for the previous Haiku model were within 100,000 tokens. Considering variations in request lengths and token consumption, the company expects the average cost of completing similar tasks with Haiku 5.5 to decrease by approximately 75%.
Here’s another easily overlooked detail—Haiku 5.5 has switched to a new tokenizer. According to the official developer documentation, the same text may generate approximately 30% more tokens in the new model compared to Haiku 4.5, with the exact increase depending on the content. Therefore, a 90% reduction in API pricing per unit does not mean your bill for each task will decrease by 90%.
Another notable competitor is GPT-6 Luna. OpenAI’s base pricing for Luna is also $0.10 per million input tokens and $0.50 per million output tokens, with cached read pricing matching the low-end tier of Haiku 5.5.
However, there is a significant difference in the long-context billing thresholds between the two companies.
Haiku 5.5 moves to a higher price tier after exceeding 100,000 tokens; GPT-6 Luna only triggers its long-context surcharge after input exceeds 272,000 tokens. This means Luna offers a lower cost per token for requests between 100,000 and 272,000 tokens.
Which model is more cost-effective for completing a task ultimately depends on actual token usage, inference budget, and task success rate.
Looking at the broader market, Anthropic is also putting pressure on other low-cost models.
According to VentureBeat’s listed API pricing, Google Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens; Gemini 3.8 Flash costs $0.75 and $3.75 per million input and output tokens, respectively.
Of course, different models vary in their capabilities, caching strategies, and task performance; these figures primarily reflect API price competition.
The price war among small models is further intensifying.
Eight million calls per week: Enterprises are beginning to have small models work for large models.
What is Haiku 5.5 best suited for?
Anthropic's scenarios include document summarization, information classification, database querying, context compression, and serving as a subagent within complex agent workflows.
There’s a very practical issue behind this: today’s agents are becoming increasingly complex, and a single task often requires multiple model calls. If every step is handled by a large model, costs quickly accumulate.
The financial AI company Rogo provides an example. In its workflow, a more powerful large model creates financial presentations. When processing a specific slide, the Haiku 5.5 Subagent accesses the company’s 10-K annual report, retrieves the required revenue data, and passes the results back to the main model.
For this type of task, the model needs to quickly and accurately perform information lookup and extraction. Rogo believes that Haiku 5.5 strikes the right balance between accuracy, speed, and cost for high-volume, repeated calls.
The enterprise content management platform Box also provided early test results. Haiku 5.5 showed an 11-point improvement in internal evaluation scores and approximately a 50% reduction in latency compared to Haiku 4.5.
Box stated that this makes the new model suitable for analytical work requiring scalability, such as cost reports, financial summaries, and periodic reviews. However, Box did not disclose the full evaluation criteria.
Asana's testing covered Agent tasks such as bug classification, project creation, and large portfolio searches. According to the disclosed results, Haiku 5.5 reduces task completion latency by over 30% and increases single-round Agent inference speed by up to 2.5 times compared to the currently used model.
Another example that better illustrates cost value comes from AlphaSense. The company’s Ask in Document feature handles approximately 8 million requests per week, primarily answering specific questions about one or several documents.
In internal evaluations across 400 test queries, Haiku 5.5 scored 0.84, compared to 0.76 for Haiku 4.5.
If a model is called millions of times per week, even a small cost saving per call can result in a significant cumulative cost reduction over time.
These data come from early customer feedback disclosed by Anthropic; the specific workloads, baseline models, and evaluation criteria are not entirely consistent, and further independent testing is required.
Sonnet is following suit with price cuts, while Claude 5.5 is now competing on cost.
In addition to Haiku 5.5, Anthropic has also adjusted the pricing for Sonnet 5.5. Starting October 7, the cached read cost for Sonnet 5.5 will be reduced by 50%, from $0.20 to $0.10 per million tokens.
Anthropic expects this adjustment to reduce the cost of most agent workloads by approximately another 20%, with the actual reduction depending on how extensively the application reuses cached context.
For agents that require repeatedly reading long conversation histories, tool descriptions, and task contexts, caching costs directly impact the system's long-term operational expenses.
Anthropic is introducing monthly API credits for Claude Max and Team subscribers: $100 for Max 5x users, $200 for Max 20x users, and up to $500 shared among Team users.
These credits can be used to invoke models on the Claude platform, encouraging more subscribed users to try building their own agents and applications.
On the developer side, Haiku 5.5 is available through platforms such as the Anthropic API, Amazon Bedrock, Google Cloud, and Microsoft Azure, with the API model ID claude-haiku-5-5.
Anthropic has also updated its Python and TypeScript SDKs, adding beta support for browser and computer operations.
At this point, the product roles within the Claude 5.5 series are already quite clear.
Opus 5.5 handles high-difficulty, complex reasoning tasks; Sonnet 5.5 covers everyday complex workflows and programming; Haiku 5.5 focuses on high-volume, cost-sensitive tasks requiring fast responses.
From September 22 to October 7, Anthropic completed this round of product updates over 15 days, with all three models emphasizing performance and cost efficiency.
The release of Haiku 5.5 has further highlighted a growing trend: as agents transition from demonstrations to real-world applications, developers must consider the overall system invocation costs—determining which steps in a task require more powerful models, which can be handled by smaller models, and how to allocate work efficiently between different models, all of which impact the final commercial viability.
For Anthropic, this may be Haiku 5.5’s most compelling feature: it makes economically viable tasks that were previously too costly to scale.
Competition among small models has already penetrated every execution step of agent systems.
Reference materials
https://www.anthropic.com/claude-haiku-5-5
https://venturebeat.com/technology/anthropic-launches-claude-haiku-5-5-with-90-api-price-reduction-matching-gpt-6-luna
https://x.com/claudeai/status/2107894042235166750
This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), edited by Spoonbill.
