OpenAI counters Anthropic with new customer privacy safeguards

▼ Summary
– OpenAI is previewing Private Safety Processing, an automated system that monitors for AI misuse across multiple conversations without retaining customer data.
– Anthropic’s data retention policy keeps user sessions for 30 days on “covered models” like Mythos-class, raising privacy concerns for enterprises with sensitive data.
– Private Safety Processing extends OpenAI’s existing Zero Data Retention policy by analyzing inputs and outputs across sessions to catch multi-step malicious activity, such as malware engineering.
– If triggered, the system sends a signal to OpenAI, which may then contact the customer, who can voluntarily share data for further context.
– The move intensifies competition with Anthropic, which allows human review only through a controlled, tamper-proof log, amid reports of Anthropic’s faster Q2 growth and higher revenue run rate.
As AI systems grow increasingly capable, the risks of their misuse expand in tandem, intensifying pressure on developers to implement protective measures. Companies in this space now face a delicate balancing act: honoring enterprise clients’ privacy expectations while still keeping a watchful eye on how their platforms are being used.
Spotting a chance to gain ground on rival Anthropic, OpenAI has unveiled a new privacy-focused approach to abuse detection. The organization is rolling out a preview of a service called Private Safety Processing to a select group of customers. This automated mechanism screens for potential violations without ever storing any of the client’s data.
The move stands in direct opposition to Anthropic’s recently introduced data retention policy, which has drawn criticism from some users. That policy allows the lab to hold onto customer data, including entire sessions and the conversations within them, for 30 days when it applies to what the company calls “covered models.” These encompass all Mythos-class models as well as “future models with similar capabilities,” according to Anthropic.
Announced in July, the retention policy was framed as a safety measure, giving the lab room to examine interactions for signs of wrongdoing. But it has unsettled enterprises that manage substantial volumes of sensitive information and prefer not to have it stored or scrutinized by an external AI provider.
OpenAI, like most of its peers, already provides a baseline level of privacy through Zero Data Retention (ZDR). Under ZDR, software agents embedded in the OpenAI API monitor for abuse on a session-by-session basis. This means customer information is not retained by the company, yet suspicious activity can still be flagged without requiring human oversight. Anthropic also follows ZDR in most cases, with the exception of those “covered models” such as Fable.
Private Safety Processing, OpenAI explains, expands the reach of ZDR. It functions as a form of long-horizon safety monitoring, evaluating inputs and outputs across multiple conversations rather than isolating a single exchange. The process is handled by an automated agent that, when triggered, examines interactions over time to identify patterns indicating possible misuse.
A spokesperson told TechCrunch that the new technology helps OpenAI catch malicious behavior that unfolds over several sessions. For example, a bad actor attempting to craft malware for a cyberattack might distribute their queries across many interactions to evade notice. Private Safety Processing can connect those dots and flag abuse without any human ever reviewing the user’s conversations.
When the system does detect a trigger, it may transmit a “narrowly defined signal” to OpenAI indicating a specific category of activity. From there, OpenAI determines whether “enforcement is necessary,” the company states. If action is warranted, OpenAI contacts the customer to gather additional context or collaborate on resolving the issue. The customer then decides whether to voluntarily share further data with OpenAI, the spokesperson added.
Anthropic, for its part, maintains that any human inspection of customer data occurs only through a “controlled access path” limited to “a small set of approved reviewers.” Each review session is logged in a “tamper-proof log that reviewers cannot suppress or modify,” the company says.
The rivalry between these two AI powerhouses is particularly sharp right now. Recent figures indicate OpenAI’s second-quarter growth lagged behind Anthropic’s, whose annualized revenue run rate has reportedly hit $65 billion. Anthropic investors have floated the possibility of a $2 trillion IPO, while OpenAI is also advancing its own public offering plans.
(Source: TechCrunch)




