OpenAI Introduces Private Safety Processing to Balance AI Security and Privacy

iconMetaEra
Share
AI summary iconSummary
On August 19, 2026, OpenAI announced the continued use of Zero Data Retention for its frontier models and introduced Private Safety Processing. The system enables enterprise-grade security by detecting risk patterns without exposing raw content. Data can be stored on private chain infrastructure or encrypted with customer-held keys. Testing is currently underway with select clients, with further details expected in September.
On August 19, 2026, OpenAI announced it would continue offering "Zero Data Retention" (ZDR) for its frontier models, while previewing a security mechanism called Private Safety Processing. It aims to address an increasingly acute contradiction brought about by the proliferation of AI Agents: complex misuse or agent deviation from user intent often can only be identified through multi-turn interactions; yet financial, medical, and research institutions cannot permit model providers to retain sensitive prompts and outputs long-term. OpenAI’s solution enables automated systems to detect risk patterns across interactions, returning only limited risk signals to OpenAI without exposing raw content to its employees. Customer content can remain on customer-controlled infrastructure or be encrypted and stored by OpenAI using customer-exclusive keys. However, this technology is still in testing with early adopters; the key technical whitepaper, false positive rates, and specific trust boundaries will be further disclosed in September.

Article author, source: OpenAI

Enterprises want AI to understand context without giving up that context.

When enterprises deploy AI, there are always two competing demands.

The first is privacy. Financial records, medical information, unpublished product plans, internal code, and research findings often cannot be stored long-term by third parties, nor can they be used to train models without explicit authorization. OpenAI’s zero-data-retention commitment states that after eligible API customers make a request, their prompts and the model’s responses will not be retained by OpenAI; OpenAI employees cannot review this customer content, and enterprise data will not be used for model training unless the customer actively opts in.

The second is security. Checking a single request might catch obvious illegal instructions, but it’s difficult to detect the full intent when it’s spread across dozens of interactions. Attackers may gradually test boundaries, coordinate actions across multiple accounts, or disguise dangerous tasks as a series of seemingly normal research questions.

Long-running AI agents can complicate issues. These agents may operate continuously for hours or even days, reading files, invoking tools, and interacting with external systems. While any single action may appear entirely reasonable, the overall sequence of actions could ultimately deviate from user authorization. For example, even after the user has requested to stop, the agent may still seek alternative ways to achieve its goal.

Therefore, security systems increasingly need to understand “what this entire process aims to achieve,” rather than simply judging whether “this single sentence violates rules.” However, if achieving this capability requires centralizing and submitting all enterprise conversations for review by model companies, many institutions subject to strict compliance requirements simply cannot adopt it.

Private Safety Processing is OpenAI's answer to this contradiction.

Do not send the original text; return only limited risk signals.

According toOpenAI's official documentation, the current compatible ZDR security system primarily examines individual interactions sequentially; Private Safety Processing extends the scope of inspection to interconnected multiple interactions, identifying patterns such as repeated probing, coordinated abuse, or gradual deviation of Agent behavior from authorized parameters.

Customers can use two methods to store content.

In a full ZDR deployment, content remains within the customer-controlled infrastructure. OpenAI is also developing another approach: content can be stored on OpenAI-provided infrastructure but encrypted with keys controlled by the customer. OpenAI states that its employees will not have access to these keys and therefore cannot decrypt or view prompts or model outputs.

After the automated security system detects a risk, OpenAI receives not the original conversation, but a limited-scope signal, such as the type of risk the activity may involve. This signal can be used to determine whether enforcement action is necessary, but even if content is flagged, OpenAI employees do not automatically gain access to the underlying data.

Customers can use information stored in their own systems to investigate alerts. If customers believe a legitimate use has been incorrectly flagged and wish to appeal, or need assistance investigating confirmed abuse, they may choose whether to provide relevant details to OpenAI.

The key to this design is to separate "whether the machine can detect risks" from "whether service provider employees can read the content" into two distinct permissions. It does not mean completely eliminating security monitoring, but rather attempting to allow monitoring results to leave the customer environment in a smaller, more limited-information form.

Why this is especially important for AI agents

The primary risks of traditional chatbots are typically focused on input and output: what the user asks and how the model responds. The risks of agents, however, may be distributed across the entire chain of actions.

Suppose an Agent is initially authorized to organize code, then begins reading permission configurations, attempts to access external services, and continues seeking alternative pathways even after being instructed to stop. Individually, each action may appear to be a normal software development task; however, when viewed together, they may indicate that the Agent is exceeding the boundaries set by the user.

This requires enterprise security systems to have some cross-step memory; otherwise, each check would be like encountering the Agent for the first time, making it impossible to determine whether risks are accumulating over time.

Private Safety Processing aims to preserve the ability of security systems to "understand trajectories" without exposing raw enterprise data to model company employees. If reliably implemented, it could become a critical infrastructure for enterprises adopting long-running Agents: customers retain control over their data, while model providers still receive sufficient risk signals to enforce security policies.

However, OpenAI has not yet provided detailed information on how the system performs analysis within customer infrastructure, how “relevant interactions” are defined, what fields are included in security signals, or how these signals are prevented from indirectly leaking sensitive information. A formal technical white paper is planned for release in September 2026, so at this stage it should be understood as an architectural preview rather than a fully auditable technical specification.

Encrypting your data does not eliminate all risks.

"Using customer-controlled keys" is an important commitment, but it does not automatically answer all security questions.

Enterprises still need to understand where automated code checks run, who updates them, what plaintext data they can access, how memory is isolated during execution, and how long metadata and risk signals are retained. The system must also address false positives and false negatives: checks that are too sensitive may hinder legitimate security research or medical work; checks that are too lenient may fail to detect fragmented and disguised abusive workflows.

The article also does not disclose the detection accuracy, performance overhead, or independent audit arrangements for Private Safety Processing. OpenAI states that the system is currently being tested with early customers and plans to roll it out gradually starting in September. Therefore, whether it can simultaneously meet privacy, security, and usability requirements in real-world environments remains to be seen until the whitepaper and customer deployments are available.

There is one clear exception: images suspected of involving child sexual abuse material will continue to be retained and subject to manual review, even if they originate from a zero-data-retention deployment, in order to fulfill legal reporting obligations. This means that “zero-data-retention” does not absolutely prohibit content retention under all circumstances; customers must understand this boundary when evaluating compliance requirements.

The next phase of competition is “who can ensure security without seeing the data.”

In the past, when enterprises compared model services, they primarily focused on capabilities, price, latency, and context window. As AI begins to handle longer and more sensitive workloads, “what the service provider can actually see” is becoming an equally important competitive metric.

One extreme is for the service provider to retain all interactions to gain a more complete security perspective; the other extreme is to monitor nothing at all, leaving all risks to be borne by the customer. OpenAI aims to demonstrate that a middle ground is possible: original content remains private, automated systems analyze risks across interactions, and the model provider receives only limited signals necessary to implement safety measures.

It cannot be concluded solely based on the company’s published article whether this proposal is truly viable. However, it accurately identifies the core challenge of the Agent era: the longer a model operates, the more context its security system requires; the more sensitive the enterprise data, the less should the service provider possess that context.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.