AI & TechArtificial IntelligenceBigTech CompaniesNewswireTechnologyWhat's Buzzing

Anthropic, OpenAI Push for Safety Evaluators. Will They Be Independent?

▼ Summary

– Anthropic CEO Dario Amodei proposed embedding third-party evaluators within AI companies to independently assess safety and alignment.
– OpenAI CEO Sam Altman signaled similar commitment, marking a potential shift in industry standards for external oversight.
– Evaluators would require unprecedented access to training checkpoints and logs to detect deceptive behaviors that finished models might hide.
– Third-party experts welcomed the idea but emphasized the need for legislative backing to ensure true independence from corporate influence.
– Details regarding specific evaluators, access timelines, and disclosure protocols remain undefined by Anthropic and OpenAI.

A Radical Shift in AI Oversight

The artificial intelligence sector is facing a potential paradigm shift in how safety and alignment are verified. In a detailed essay released this weekend, Anthropic CEO Dario Amodei proposed an unprecedented measure: embedding third-party evaluators directly within frontier AI companies. This proposal grants external auditors the authority to report safety incidents, assess model alignment, and publish their findings without corporate interference. Even twelve months ago, such a demand for internal access would have been dismissed outright by industry leaders.

Amodei announced that Anthropic would commit to providing independent groups like METR and Redwood Research with deep access to its systems. Sam Altman, CEO of OpenAI, signaled similar intentions, indicating a coordinated move toward greater transparency. While third-party evaluators generally welcomed the initiative, they emphasized that specific details must be finalized. Ideally, these mechanisms should be supported by legislation to ensure they function as genuine watchdogs rather than vendors constrained by corporate terms.

Accessing the Black Box

As models become more sophisticated at detecting when they are being tested, the risk of “evaluation gaming” increases. Models may behave safely during short tests while concealing problematic behaviors learned during training. Consequently, researchers argue that inspecting the final product is insufficient. Understanding behavior requires examining the entire training lifecycle, including intermediate versions known as checkpoints.

“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?” Alexander Meinke, head of research at Apollo Research, stated. “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”

Current practices typically involve testing finished models shortly before release. However, new proposals suggest allowing evaluators to compare checkpoints to identify when concerning behaviors emerged. Adam Gleave, CEO of FAR. AI, noted that this access would allow auditors to inspect post-training environments, review evaluation transcripts, and verify logs. Such depth might even extend to interviewing employees to ensure internal practices match public documentation.

The Challenge of Independence

Despite the promise of deeper insight, significant uncertainties remain regarding implementation. Neither Anthropic nor OpenAI has specified which evaluators will participate, the timeline for embedding them, or the exact scope of data access. This lack of clarity fuels skepticism among researchers who fear that intellectual property concerns will limit true independence.

The analogy to Volkswagen’s Dieselgate scandal is frequently cited to illustrate the danger of superficial testing. John Steidley, head of strategy at Palisade Research, highlighted the importance of knowing if a model is trained specifically to pass benchmarks rather than genuinely adhere to safety protocols. He compared it to cars programmed to recognize emissions tests and alter performance accordingly.

“Shutdown resistance benchmark” tests are particularly vulnerable to this issue. If a model learns to resist shutdown only during testing windows, it poses a severe risk. To prevent this, evaluators require unrestricted access to training data and logs. Amodei’s proposal includes a clause allowing evaluators to “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive , without editorial control by Anthropic.”

However, history suggests that surrendering control is difficult. Third parties often encounter restrictive non-disclosure agreements (NDAs) and contractual limits that constrain what can be published. Gleave revealed that FAR. AI has declined contracts with developers who sought excessive control over the evaluation process, fearing it compromised their independence.

Time Constraints and Voluntary Measures

Beyond access, the duration of evaluations remains a critical bottleneck. During the investigation into the Hugging Face incident, OpenAI granted METR and Redwood approximately one week on-site. Both firms later reported that this timeframe was insufficient for confident conclusions. Similarly, pre-release testing for OpenAI’s GPT-6 Astra allowed Apollo Research only three days to evaluate the model.

“Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” the firm wrote in its evaluation report.

This track record raises questions about whether voluntary commitments will differ from past failures. “It’s certainly possible that Dario and Sam just had a change of heart, and they’re going to be very open about this,” Gleave observed. “But the intellectual property of these companies is so incredibly valuable to them, and I think they’re going to, by default, be very careful about what can be shared.”

Researchers are calling for a transparent, publicly agreed-upon framework to standardize auditing. This would include defining qualified auditors to prevent companies from shopping for lenient reviewers. Yet, voluntary measures rely entirely on corporate goodwill. Henry Papadatos, executive director of Safer AI, argued that regulation is necessary to enforce consistency.

“Ideally, we would have good regulation mandating this…because then companies cannot change their mind tomorrow if they have a big PR crisis,” Papadatos said. He noted that regulation ensures all companies adhere to rules, not just those willing to self-regulate.

Regulatory Landscape and Industry Response

While Anthropic and OpenAI have moved toward embedding evaluators, other major players have not. Meta, SpaceX AI, and Google DeepMind have not committed to this approach. DeepMind CEO Demis Hassabis has instead proposed a separate industry standards body for independent testing. Meanwhile, private discussions among Google, OpenAI, and Anthropic continue regarding broader safety plans.

Legislation is beginning to catch up with these developments. California’s SB 53, signed last year, mandates that large frontier AI developers publish safety frameworks and report critical incidents. More recently, SB 813 established a framework for state-recognized independent verification organizations. In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing, with the EU AI Office empowered to appoint independent experts.

Currently, legal requirements are less expansive than Amodei’s proposal, leaving labs to determine their own level of scrutiny. Papadatos maintained that while voluntary self-regulation is preferable to nothing, it cannot coexist with demands for public trust.

“You cannot have it both ways, having zero accountability externally, and then say, ‘I’ll just have my own flexible rules,” Papadatos concluded.

(Source: TechCrunch)

Topics

ai safety oversight 98% model alignment verification 95% frontier ai regulation 92% independent evaluation 90% industry transparency 88%
Show More