Anthropic Launches Claude Sonnet 5.5 with 30% Faster Output and Reduced Costs

iconMetaEra
Share
AI summary iconSummary
Anthropic announced the launch of Claude Sonnet 5.5, a new model optimized for high-frequency tasks. The model delivers 30% faster output and up to 30% lower cost per task than Sonnet 5. On-chain data shows it achieved a score of 70.6% on Terminal-Bench 4.0, outperforming both Sonnet 5 and Opus 5.5. However, under maximum reasoning loads, it consumes more tokens and can cost up to 50% more. Anthropic recommends Sonnet 5.5 for well-defined tasks, while Opus is better suited for complex work.
Anthropic has released Claude Sonnet 5.5, aiming to compress capabilities close to those of the flagship Opus 5.5 into a more affordable model designed for high-frequency, everyday agent use. Its API pricing remains at $2 and $10 per million input/output tokens, but the company claims a 30%+ improvement in output speed and reduced token usage for typical tasks, lowering per-task costs by up to 30%. In official benchmarks, it achieved 70.6% on Terminal‑Bench 4.0, surpassing Sonnet 5’s 10.3% and Opus 5.5’s 66.4%; it also reached 1844 Elo on the professional task benchmark GDPval‑AA, nearly matching Opus’s 1846. However, calling it a “half-price Opus” is misleading: Artificial Analysis found that under maximum reasoning intensity, Sonnet 5.5 generates an average of ~193,000 output tokens per question—the highest token consumption among all configurations tested—resulting in a per-question cost roughly 50% higher than Sonnet 5. Real-world code review tests by CodeRabbit also showed that while Sonnet 5.5 is roughly twice as fast as its predecessor, it still misses complex issues that Opus can detect. Thus, it functions more as a new primary model suited for clear, repeatable tasks rather than a full replacement for Opus.

Article author, source: Anthropic

Anthropic is repositioning Sonnet as the everyday agent powerhouse.

Sonnet 5.5 is the second model in the Claude 5.5 series. Unlike Opus 5.5, which was released a week ago for complex open-ended tasks, Anthropic positions Sonnet 5.5 for routine work with well-defined boundaries requiring extensive repetition: fixing program bugs, reviewing code, researching information, and generating documents, presentations, and spreadsheets.

The model supports text and image inputs, offers a 1 million token context window, and is available on the Claude web and mobile apps, Anthropic API, AWS, Google Cloud, and Microsoft Azure. The API name isclaude-sonnet-5-5. Anthropic has not disclosed the model's parameter count, architecture, training data, or training compute, so it is not possible to determine whether the improvements stem from base training, post-training, inference strategies, or Agent toolchain optimizations.

The difference between it and Opus is not merely “weaker capabilities and a lower price.” Anthropic aims to establish a new model division of labor: Opus handles complex tasks with incomplete problem definitions that require ongoing judgment, while Sonnet quickly executes well-defined tasks and enables higher invocation volumes at a lower cost.

70.6% versus 10.3% is the most striking and requires careful interpretation.

Sonnet 5.5 achieved 70.6% on Anthropic’s Terminal-Bench 4.0 test, compared to Sonnet 5’s 10.3%, and even surpassed Opus 5.5’s 66.4%. This benchmark evaluates an agent’s ability to complete multi-step professional tasks in a command-line environment, reflecting coordination between code writing, terminal operations, and tool usage.

Improvements in other benchmarks were not as dramatic but still significant: CursorBench 4.0 rose from 34.1% with Sonnet 5 to 55.5%, just two percentage points below Opus 5.5’s 57.8%; Humanity’s Last Exam with tools increased from 54.9% to 64.5%; and OSWorld 2.1 computer operation jumped from 57.0% to 80.1%, nearing Opus’s 81.8%.

In the professional knowledge task GDPval-AA, Sonnet 5.5 scored 1844 Elo, while Opus 5.5 scored 1846; in the long-form knowledge evaluation AA-Briefcase, they scored 1811 and 1822 respectively. Anthropic concludes that for certain well-defined professional tasks, it is no longer necessary to default to the flagship model.

However, these figures cannot be simply combined to mean “Sonnet has surpassed Opus.” Different projects use varying levels of inference strength, and Anthropic has explicitly stated that Opus remains significantly stronger in tasks requiring sustained judgment and handling open-ended objectives.

An interesting counterexample comes from FrontierCode. Sonnet 5.5 achieved 52.1% on Xhigh inference intensity, but dropped to 46.2% when upgraded to Max. Anthropic explained that in Max mode, the model more frequently initiates multiple code review sub-agents, resulting in two timeouts or modifications to code outside the task scope, ultimately leading to failure in evaluation.

This result indicates that more reasoning and more agents do not necessarily lead to better answers. The model may lose points due to over-checking, expanding the task scope, or exhausting its time budget.

Why does the official statement say the task cost is 30% lower if no token has been discounted?

The public pricing for Sonnet 5.5 is the same as Sonnet 5: $2 per million input tokens, $10 per million output tokens, $2.5 per cache write, and $0.2 per cache read—each half the corresponding price of Opus 5.5.

Anthropic’s claim of “up to 30% cheaper per task” does not refer to a reduction in API pricing, but rather to the model requiring fewer steps and tokens to complete the same task. The company states that generation speed has improved by over 30%; Slack observed approximately a 14% reduction in output tokens during internal testing; Atlassian reported agent speeds increasing by up to 30%; and Lovable reported a reduction of about one-third in tool calls and roughly half as many Shell executions.

When tested on 2,441 financial Q&A, extraction, analysis, and prediction tasks, Sonnet 5.5 used an average of approximately 121,000 tokens per task, while Sonnet 5 used about 497,000. Since these were private tests conducted by partners, external parties cannot verify task difficulty, prompts, or full outputs, but the results collectively point to a shift: the new model performs less repeated searching, redundant reading of materials, or excessively long internal reasoning.

This is especially important for agents. In traditional chat, outputting hundreds of tokens at once has limited impact; however, long-term agents repeatedly read context, invoke tools, and correct results, where a single redundancy can amplify over dozens of rounds into significant time and cost differences.

The highest reasoning intensity reveals another truth about costs

Independent testing of Artificial Analysis awarded Sonnet 5.5's highest reasoning configuration a score of 56, just 2 points behind Opus 5.5's 58 in its overall intelligence index.

But achieving this performance came at a high cost: Sonnet 5.5 generated an average of approximately 193,000 output tokens per test, the highest level measured by the institution—about 60% higher than Opus 5.5 Max and Sonnet 5 Max, and roughly seven times that of GPT-6 Astra Max. Its estimated cost per question reached $7.60, about 50% higher than Sonnet 5.

This does not directly contradict Anthropic's claim of "task costs reduced by up to 30%." The official conclusion primarily describes typical workloads and lower inference intensity; Artificial Analysis measures scenarios where Max is enabled to pursue maximum capability.

The combined results indicate that inference strength has become a core component of the model product. Sonnet 5.5 may offer exceptional value at Low or Medium settings; however, pushing inference strength to Max to match Opus, while cheaper per token, could lose its cost advantage due to excessive output volume.

Artificial Analysis also found that Sonnet 5.5’s factual knowledge accuracy is approximately 54%, lower than Opus’s 66%; it also lags by about six percentage points in scientific reasoning tasks. The claim of being “close to Opus” primarily stems from agent terminal operations and certain knowledge work, and cannot be extrapolated to all capabilities.

In real code reviews, the speed advantage holds, and the flagship gap remains.

CodeRabbit performed a set of release-day tests using known bugs from real open-source projects.

In 13 challenging code review cases, Sonnet 5.5 with reasoning identified 6 issues, outperforming Sonnet 5’s 4; both achieved effective comment precision rates of 41.2% and 40.0%, respectively. However, Opus 5.5 in standard mode found 8 issues, and in Max mode found 10, indicating a still significant gap in complex review tasks.

In another test involving 44 real pull requests, Sonnet 5.5 averaged 6 minutes and 33 seconds per review, compared to 13 minutes and 31 seconds for Sonnet 5. The former generated 24% fewer comments and only about one-third as many trivial remarks as its predecessor; based on public pricing, the core model calls averaged approximately $0.46 for Sonnet 5.5 versus $1.16 for Sonnet 5.

However, the full error coverage score for this set of 44 PRs was not yet complete at the time of the article’s publication, so it can only be confirmed that it is faster and outputs less—whether the fewer comments missed additional real issues remains uncertain. The sample of 13 difficult cases is also small; CodeRabbit itself noted that it is sufficient to observe trends but insufficient to determine general rankings.

A more rational division of labor is: Sonnet 5.5 handles rapid routine reviews of each PR; changes involving security, core architecture, or those with high costs if errors are missed are escalated to Opus or human experts.

Visual design is beginning to become a selling point for universal work models.

Anthropic also highlighted Sonnet 5.5’s visual design capabilities. In official demonstrations, Sonnet 5 and 5.5 were tasked with creating, within a single HTML file, 400 starlings in flight, wind-sculpted dunes, and a large clock composed of 24 small bells. The new model generates faster and delivers higher quality animations, layers, and interface completeness.

It can also generate presentations based on templates. In an internal test, Anthropic provided materials including a public company’s quarterly earnings report, transcript of the earnings call, and a ten-slide presentation template; two experts concluded that the first version of Sonnet 5.5 could be sent out as-is.

These are still selected demonstration cases chosen by the vendor and do not represent the model’s ability to consistently understand every enterprise template. Data accuracy, brand guidelines, and source citations in complex charts still require manual verification, but they reflect a new direction in model competition: not only generating correct content, but also delivering interfaces and documents that closely resemble final products.

Improved security capabilities also mean that some requests will be handled by older models.

Sonnet 5.5 is the first Sonnet model to feature enterprise-grade cybersecurity protection at launch. Anthropic states that its cybersecurity capabilities are now close to those of Opus 5, so routine vulnerability fixes can still proceed normally, while higher-risk cybersecurity requests are transparently escalated to Sonnet 5.

The company also tested the model for behaviors such as misleading users, acting against user interests, facilitating high-risk abuse, and breaching environmental boundaries across approximately 1,850 scenarios. The official stated that the new model performs equally or better than Sonnet 5 on most metrics and is the least likely among all Claude models to actively probe container boundaries.

However, these are primarily internal automated evaluations by Anthropic and cannot be considered a complete proof of risks in real-world environments. The company itself acknowledges that no set of tests can identify all failure modes.

Sonnet 5.5 is the first Sonnet to incorporate an anti-distillation classifier and expands the "retain thinking content" mechanism to prevent the transfer of reasoning context between accounts. It has limited impact on general developers, but some workflows involving migrating Claude Code conversations across accounts will require adjustments.

The real change isn't being number one on the leaderboard, but that "strong enough models" are becoming the default choice.

Sonnet 5.5 did not introduce a new model architecture or comprehensively surpass Opus. Its greater significance lies in shifting a large volume of tasks previously handled by flagship models to mid-tier models that are faster and offer half the cost per token.

For customer service, code review, research, and documentation systems that run thousands of times daily, the decision to deploy is typically not based on the highest single score, but rather on how many steps each task requires, how many tokens are used, how long the wait time is, and whether errors can be escalated. Sonnet 5.5’s value is precisely focused on these metrics.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.