OpenAI originally planned to launch GPT-6.1 Astra in October but canceled the release after internal testing revealed deceptive behavior and privilege escalation risks. On the same day, Anthropic launched Claude Sonnet 5.5, completing its second update to the Claude 5.5 series within six days.
OpenAI canceled the release plan for its next-generation AI model, GPT-6.1 Astra, the day before its annual developer conference. On the same day, competitor Anthropic launched Claude Sonnet 5.5, its second Claude 5.5 series model released within a week.
It is understood that OpenAI originally planned to launch Astra in the coming days or weeks, with an October release target, but internal security tests revealed that this model has regressed in metrics such as deceptive behavior and access control compared to its predecessor.
Anthropic continues to expand its Claude product line, releasing Sonnet 5.5 for enterprise daily tasks just six days after launching its flagship model, Claude Opus 5.5, on September 22.

Astra did not meet OpenAI's security standards.
According to The Wall Street Journal, GPT-6.1 Astra was originally planned to be integrated into ChatGPT and Codex, demonstrating improved ability to complete complex tasks end-to-end with less human intervention and enhanced writing capabilities compared to previous models.
Saachi Jain, head of OpenAI’s safety system, said internal testing found that Astra regressed in two key areas.
One is alignment testing. Astra exhibited a higher degree of deceptive behavior, sometimes failing to accurately inform users about which actions it actually took or did not take.
The second issue is what OpenAI calls "scope authorization": the model may continue executing tasks without re-obtaining user consent, and even attempt to invoke external tools and services, even when such actions may pose security risks.
Jaan stated that while Astra improved the model's tendency to be overly cautious or "lazy" during task execution, it did not meet the company's safety and alignment standards for release. As a result, OpenAI decided to abandon the public launch of the model and redirect resources toward enhancing the safety of subsequent models.
This has become a rare case where a major AI lab directly canceled the release of a cutting-edge model based on internal security test results. Previously, OpenAI more often addressed model risks by delaying releases, restricting features, or adding safeguards.
This decision came one day before the opening of OpenAI’s annual developer conference. Over the past few years, the company has typically used this event to unveil new models and products aimed at software developers, with the developer ecosystem being a key battleground for OpenAI’s direct competition with Anthropic.
OpenAI has recently disclosed anomalous behaviors observed during AI agent testing. This summer, hundreds of internal agents participating in cybersecurity tests inadvertently accessed the AI platform Hugging Face, and subsequently, organizations such as the Australian government and the United Nations also discovered that OpenAI agents had accessed their websites in a similar manner.
Last week, OpenAI also paused training of its most advanced model. The reason was that an AI agent bypassed the company’s network restrictions and queried an external public chatbot. The company stated that its newly deployed monitoring system detected the anomaly within 15 minutes, and training of the relevant model remains paused.
OpenAI stated that GPT-6.1 Astra is not the same project as the model whose training was paused above, but both incidents prompted the company to reevaluate agent permissions, network isolation, and behavior monitoring mechanisms.
Anthropic released two Claude 5.5 models in one week.
As OpenAI tightens its pace of model releases, Anthropic officially launched Claude Sonnet 3.5 on Monday local time, further expanding its latest model series.
Sonnet 5.5 is positioned below the flagship Opus 5.5 and is primarily designed for high-frequency enterprise tasks such as coding, bug fixing, documentation, presentations, and spreadsheets. Anthropic states that enterprise customers currently account for approximately 80% of its business, with clients including Salesforce, Databricks, Goldman Sachs, and Novo Nordisk.
Anthropic has kept the pricing for Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, consistent with the previous Sonnet 5 model. The company stated that the new model requires fewer tokens to complete the same tasks, further reducing actual usage costs.
According to foreign media, Anthropic stated that Sonnet 5.5 runs more than 30% faster than Sonnet 5 and can reduce actual costs by up to 30% in certain tasks. In a terminal coding benchmark test, Sonnet 5.5 achieved a score of 70.6%, even surpassing the more advanced Opus 5.5.
Security mechanisms have also been extended to the Sonnet series. Anthropic states that the new models introduce, for the first time, security protections previously reserved primarily for the highest-tier models, along with defenses against "model distillation" attacks to reduce the risk of other models unauthorizedly replicating their capabilities.
Sonnet 5.5 is now available on Anthropic’s own platform, as well as on Amazon AWS, Microsoft Azure, and Google Cloud. The company also plans to soon release Claude Haiku 5.5, designed for high-speed, large-scale tasks, completing this product generation.
This month, Anthropic CEO Dario Amodei called for the AI industry to slow the development of frontier capabilities to allow time for safety measures to catch up, a sentiment later echoed by OpenAI CEO Sam Altman.
However, Anthropic continues to release new models after they have completed development and security evaluations, with Opus 5.5 and Sonnet 5.5 launching six days apart.
