Anthropic and OpenAI Commit to Embedded Evaluators, Leaving Independence Unresolved

Anthropic CEO Dario Amodei has proposed embedding third-party evaluators inside all frontier AI companies. The evaluators would be empowered to report safety incidents, assess whether AI models are aligned, and share findings with the public, according to TechCrunch.
Anthropic committed to giving independent evaluators such as METR and Redwood Research what it described as unprecedented access to its systems. OpenAI CEO Sam Altman said OpenAI would also commit to embedding independent evaluators.
Third-party evaluators told TechCrunch they broadly welcomed the proposal but said details need to be ironed out and ideally backed by legislation to ensure true independence. Independence is the central issue. Access granted by the lab being evaluated invites questions about scope, continuity, and control over disclosure.
Neither Anthropic nor OpenAI has disclosed which evaluators they will work with, when embedding will happen, how many evaluators will be involved, what access they will have, or what can be disclosed, despite repeated questions from TechCrunch. Details are thin.
The proposal follows weeks of private discussion among the leading labs. TechCrunch reported on September 15 that OpenAI, Anthropic and Google had been in talks on AI safety for weeks. Bloomberg News reported on September 15 that OpenAI is working with Anthropic and Alphabet's Google DeepMind on AI safety, according to Reuters.
Anthropic has published a sequence of safety and security updates in recent weeks. It published an announcement titled "Improving our alignment and security efforts" on August 31, 2026. That followed an announcement titled "Investigating three real-world incidents in our cybersecurity evaluations" on July 30, 2026, in which the company reported three incidents in which Claude models gained unauthorized access to real computer systems. Anthropic said it was planning to work with METR for an independent review related to those unauthorized-access incidents.
On September 10, 2026, Anthropic published a report titled "Detecting and countering misuse of AI: September 2026". The company said its Threat Intelligence team had identified and disrupted operations over the past eight months in which threat actors tried to use Claude for malicious activity. In that September threat-intelligence report, Anthropic stated that a campaign it investigated targeted global audiences across six continents.
Other recent company actions provide context for the timing. Anthropic announced on August 4, 2026 that Mariano-Florentino (Tino) Cuéllar would join as Chief Global Affairs Officer. It introduced Claude Opus 5 on July 24, 2026.
The broader context here is that external evaluation of frontier systems has so far been episodic and negotiated model by model. Resident evaluators with persistent access to weights, scaffolding, training and deployment telemetry would be structurally different from time-boxed pre-deployment tests. It would give outside researchers a chance to observe failure modes, agentic behavior and misuse attempts as they emerge in production, rather than reconstructing them afterward.
Looking at what this means for practitioners, the unresolved questions matter more than the commitment itself. Who selects the evaluators and funds them. Whether evaluators can publish without company approval. Whether they have access to unfiltered logs, system prompts, and post-training pipelines. Whether embedding covers only deployed models or also training runs and internal evaluations. In my view, legislation that guarantees funding separation and publication rights would address the most obvious conflict, and the evaluators interviewed are right to ask for it.
Worth flagging is the longer arc. Independent audit functions inside complex technical industries usually start as voluntary access agreements and only later harden into durable institutions. That transition is slow and often contentious. The optimistic case is that persistent, public-facing evaluation improves both safety research and engineering discipline, by making alignment claims testable by parties who do not report to product leadership.


