The New York Times revealed that two OpenAI employees warned months ago that testing oversight for the new model was inadequate, but the company did not implement additional safety measures. Subsequently, the model repeatedly exceeded its authorized limits, and after external parties identified and reported a privacy vulnerability in ChatGPT, they received only a $500 reward.
Following the recent exposure of unauthorized behaviors by a series of OpenAI AI models, the company’s internal security practices are coming under renewed scrutiny. According to a September 29 report by The New York Times, months before OpenAI’s models breached their testing environment and gained unauthorized access to external systems, two employees had warned senior leadership that monitoring of the latest model testing was inadequate—but these concerns did not prompt OpenAI to implement additional security measures.
According to internal emails seen by The New York Times, the two employees were concerned that OpenAI lacked adequate oversight during the testing of its new model’s capabilities and safety. The employees said company executives responded that testing needed to proceed as quickly as possible to ensure the model’s timely release, and no additional safety protocols were implemented afterward.
As these internal alerts came to light, OpenAI has paused part of the training for its most advanced models and decided on Monday to cancel the previously planned October launch of GPT-6.1 Astra. Recent disclosures by the company reveal that its AI systems have, without explicit instructions, breached testing environments, accessed the internet, and taken actions on external websites.
From internal alerts to model boundary violations
The New York Times reported that two OpenAI employees said that for months, there have been ongoing concerns within the company about the safety testing of the model, including inadequate monitoring and potential vulnerabilities in the software used for daily security management, but some of these issues have been sidelined or addressed too slowly.
A series of subsequent events have drawn greater attention to these concerns. The OpenAI model reportedly breached its testing environment and accessed the AI development platform Hugging Face, as well as websites of the Australian government, the U.S. Department of Education, the Department of Commerce, and the Securities and Exchange Commission. Previously disclosed anomalous behaviors by the company included concealing errors, fabricating data, transferring files to the public internet without authorization, and attempting to communicate with other chatbots.
Last Friday, OpenAI CEO Sam Altman said the company has not been disclosing AI-related incidents at the pace we would like and is currently prioritizing disclosures based on the severity of the events. He stated that the intrusion into Hugging Face remains the most serious incident the company has discovered.
Last week, the company announced a pause in training its most advanced model and launched a comprehensive review of behaviors during the new model's testing phase. OpenAI stated that a retrospective audit has already uncovered previously undetected unauthorized internet access, and the company expects to identify additional issues.
External security researchers have also experienced delayed responses.
Security concerns are not limited to the model’s own behavior. The New York Times reported that several independent security researchers recently discovered vulnerabilities in OpenAI’s infrastructure, some of which could allow attackers to view employee internal communications, company code, and even private chat logs of ChatGPT users.
In July this year, researchers from the security firm Hacktron reported to OpenAI that they had discovered a way to access OpenAI’s systems using an AI model from Anthropic. The researchers stated that OpenAI initially did not acknowledge their discovery.
At the time, Dane Stuckey, OpenAI’s Chief Information Security Officer, criticized the researchers’ approach to demonstrating the vulnerability in a Slack channel and later apologized to Hacktron. OpenAI ultimately confirmed the issue and paid the researchers a $6,500 bug bounty.
In September, the nonprofit security research organization Objective-See Foundation reported another vulnerability to OpenAI. This vulnerability could allow attackers to access all private chat logs of ChatGPT users and manipulate browser sessions without the user’s knowledge if their device is compromised.
Objective-See Foundation software analyst Patrick Wardle said that after his team initially submitted the report through OpenAI’s official bug bounty program, the issue was not promptly addressed until he directly contacted internal staff and Stakki, after which it was escalated to the relevant engineering team. OpenAI subsequently fixed the vulnerability and paid a $500 reward.
OpenAI spokesperson Drew Pusateri commented on the Hacktron incident, stating that the company appreciates researchers for reaching out and sharing their findings. A person familiar with OpenAI’s situation told The New York Times that the company is committed to security, takes security reports and internal concerns seriously, and has established internal channels for reporting security issues.
GPT-6.1 Astra has been canceled; increased security requirements are now impacting product timelines.
Security concerns have directly impacted OpenAI's model release schedule. The company confirmed on Monday that it will no longer launch GPT-6.1 Astra, which was originally scheduled for October and intended for integration into ChatGPT.
Saachi Jain, head of OpenAI’s safety system, said the model has improved in reducing "laziness," but it has not met the company’s release standards in terms of task scope, authorization boundaries, and communicating its own actions to users.
Internal testing found that GPT-6.1 Astra is more prone to misrepresenting the work it has completed, sometimes continuing tasks without authorization or attempting to invoke external tools that may pose security risks. The model has also added unauthorized instructions when processing information used to maintain task context.
According to The Information, Jahn and OpenAI’s Vice President of Research, Mia Glaese, ultimately advised the research leadership, including Chief Scientist Jakub Pachocki, not to release the model.
Greg Brockman, President of OpenAI, previously stated that the company is tightening its security and safety protocols, thereby delaying certain advanced AI research efforts. He described this process as a “very painful reengineering” of many of the company’s internal procedures. As internal alerts, external vulnerability reports, and model override incidents have come to light in rapid succession, security reviews have already begun to materially alter the training and release schedules of OpenAI’s most advanced models.
