Anthropic and OpenAI Want Outside Referees Watching Their Labs. The Catch Is Who Picks the Referee.

Anthropic and OpenAI have committed to embedding third-party safety evaluators inside their labs with access to training checkpoints and the right to publish findings, but Meta, Google DeepMind, and SpaceX AI haven't signed on,…

Server rack in a dark data center, representing AI lab infrastructure under third-party safety review

Written by Admin Alex · Fact-Checked by M.Ali · Info Verified September 2026

We review and update this article regularly as new information becomes available.

TL;DR: Anthropic and OpenAI have committed to embedding independent evaluators inside their labs, giving them access to intermediate training checkpoints and the right to publish safety findings without company sign-off. Meta, Google DeepMind, and SpaceX AI haven’t agreed to it. Groups that have done this kind of evaluation before say they’ve previously been handed one week, sometimes three days, to test a frontier model.

Close-up of a server in a data center server room

Every frontier AI lab says it takes safety seriously. Fewer of them are willing to let a stranger walk in, look at the model mid-training, and publish whatever they find, good or bad, without asking permission first.

That’s the proposal Anthropic CEO Dario Amodei put forward: embed third-party evaluators directly inside frontier AI companies, give them real access, and let them report incidents and share findings publicly, with the company having no editorial control over what gets published. Anthropic and OpenAI have both committed to the practice. Meta, Google DeepMind, and SpaceX’s AI operation have not.

What “real access” is actually supposed to mean

The proposal isn’t a light-touch audit. Evaluators under this model would get access to intermediate model versions, the checkpoints a model passes through mid-training, not just the finished product a company ships publicly. They’d also see post-training environments, evaluation logs, and get to interview employees directly to check whether a company’s safety documentation actually matches what’s happening internally. And critically, they’d have the right to publish what they find about risk levels and lab practices, independent of the company’s approval.

Several existing evaluation groups are named as the kind of organizations that would do this work: METR, Redwood Research, FAR.AI, Apollo Research, Palisade Research, and Safer AI. These aren’t new names in AI safety circles, they’ve been doing narrower versions of this work for a while already.

Here’s why “already doing this work” isn’t reassuring

The track record on how much access these evaluators actually get, when a company technically brings them in, is not great. Restrictive non-disclosure agreements have historically limited what evaluators could say or when, turning them into something closer to ordinary contractors than independent watchdogs. Apollo Research says it got three days to test GPT-6 Astra. METR and Redwood Research got one week to investigate an incident involving Hugging Face. A week or three days is not enough time to seriously probe a frontier model for the kinds of risks these evaluations are supposed to catch.

Adam Gleave, who runs one of these evaluation groups, put the tension bluntly: his organization has turned down contracts with multiple frontier developers specifically because those companies wanted too much control over how the evaluation actually happened. That’s the quiet part of this story. The labs asking for oversight are often the same ones negotiating hard over how much real oversight they’re willing to accept.

The regulatory backdrop making this less optional

This isn’t happening in a vacuum. California’s SB 53, passed in 2025, already requires frontier AI developers to publish safety frameworks and report critical incidents. A newer law, SB 813, passed this year, sets up a framework for state-recognized independent verification organizations, essentially giving California’s government a formal role in deciding who counts as a legitimate outside evaluator. The EU AI Act carries its own evaluation and incident-reporting requirements for frontier developers. Voluntary commitments like Anthropic and OpenAI’s proposal are, in part, an attempt to get ahead of rules that are coming whether labs like them or not.

Henry Papadatos of Safer AI made the stakes explicit: once real regulation exists, “companies cannot change their mind tomorrow” the way they can with a voluntary pledge made today and quietly walked back during the next PR crisis.

Bottom Line

A voluntary commitment to outside oversight is worth something, but only as much as the access behind it, and the history here is three days to test a model that could reshape entire industries. Anthropic and OpenAI deserve credit for putting their names on this before regulation forces the issue. Whether it means anything depends entirely on whether evaluators get real time and real access this time, or whether “independent” quietly becomes another marketing word.