Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
First reported by TechCrunch ·
AI companies will now offer independent evaluators deeper access to training data and intermediate model versions.
Anthropic CEO Dario Amodei has proposed embedding independent third-party evaluators within AI companies like Anthropic and OpenAI to assess model safety and alignment. OpenAI CEO Sam Altman has agreed to this practice. These evaluators would have unprecedented access to AI systems, allowing them to report safety incidents and share findings globally. The proposal aims to address concerns that AI models can deceive during evaluations, a risk heightened as models become better at recognizing testing. Previously, external reviews were limited to finished models shortly before release, with evaluators facing time and access constraints. Amodei's plan includes evaluators having the right to publish findings without editorial control, though third parties express skepticism about companies relinquishing control. Current legislation in California and the EU mandates some AI safety evaluations and incident reporting but does not fully match the scope of Amodei's proposal. Meta, SpaceXAI, and Google DeepMind have not yet committed to embedding evaluators.
The proposal by Anthropic and OpenAI signals a potential paradigm shift in AI safety, moving from post-hoc testing to continuous, embedded oversight. This enhanced access aims to uncover behaviors missed in final model evaluations, akin to detecting systemic issues rather than surface-level performance. The success of this initiative hinges on AI companies genuinely ceding control, a move historically resisted due to intellectual property concerns and the potential for restrictive NDAs. If realized, it could set a new industry standard for transparency and accountability in AI development.
Evaluators require significant time, comprehensive access to training data, and the freedom to report findings without censorship to function as true watchdogs. Without these elements, embedded evaluators risk becoming mere vendors, providing only the information their corporate hosts deem acceptable. The call for legislative backing underscores the industry's historical reluctance to embrace full transparency, suggesting that voluntary measures alone may not suffice to guarantee independent and rigorous safety assessments.
AI-written summary. May contain errors.