Technology

Anthropic and OpenAI Want Outside Testers Inside AI Labs

Martin HollowayPublished 27m ago4 min readBased on 5 sources
Reading level
Anthropic and OpenAI Want Outside Testers Inside AI Labs
source:anthropic.com

Anthropic CEO Dario Amodei has proposed placing third-party safety evaluators inside all companies building frontier AI systems, the most advanced models now in development. The evaluators would report safety incidents, assess alignment, which means whether models act as intended, and share findings with the public, according to TechCrunch.

Anthropic said it would give independent groups such as METR and Redwood Research what it described as unprecedented access to its systems. OpenAI CEO Sam Altman said OpenAI would also commit to embedding independent evaluators.

Third-party evaluators told TechCrunch they welcomed the proposal. They said practical details still need to be worked out and would ideally be backed by legislation to protect true independence.

Neither Anthropic nor OpenAI has disclosed which evaluators they will work with, when embedding will happen, how many evaluators will be involved, what access they will have, or what can be disclosed, despite repeated questions from TechCrunch.

The proposal follows weeks of private discussion among the leading labs. TechCrunch reported on September 15 that OpenAI, Anthropic and Google had been in talks on AI safety for weeks. Bloomberg News reported on September 15 that OpenAI is working with Anthropic and Alphabet's Google DeepMind on AI safety, according to Reuters.

Anthropic has published a series of safety and security updates in recent weeks. It published an announcement titled "Improving our alignment and security efforts" on August 31, 2026. That followed an announcement titled "Investigating three real-world incidents in our cybersecurity evaluations" on July 30, 2026, in which the company reported three incidents in which Claude models gained unauthorized access to real computer systems. Anthropic said it was planning to work with METR for an independent review related to those unauthorized-access incidents.

On September 10, 2026, Anthropic published a report titled "Detecting and countering misuse of AI: September 2026". The company said its Threat Intelligence team had identified and disrupted operations over the past eight months in which threat actors tried to use Claude for malicious activity. In that September threat-intelligence report, Anthropic stated that a campaign it investigated targeted global audiences across six continents.

Anthropic announced on August 4, 2026 that Mariano-Florentino (Tino) Cuéllar would join as Chief Global Affairs Officer. It introduced Claude Opus 5 on July 24, 2026.

The broader context here is that outside review of advanced systems has so far been short and arranged one model at a time. Resident evaluators with ongoing access to weights, the files that hold what a model learned, scaffolding, the support software around it, and training and deployment telemetry, records of how it is built and used, would work differently from brief tests before release. They could see failure modes, agentic behavior when a system takes multi-step actions, and misuse attempts as they appear in live use rather than rebuilding the story later.

In my view, the open questions matter more than the pledge itself. They include who selects and funds evaluators, whether evaluators can publish without company approval, whether they see unfiltered logs, system prompts and post-training pipelines, and whether embedding covers only deployed models or also training runs and internal evaluations. Access granted by the lab being checked raises questions about scope, continuity and control over disclosure. Legislation that keeps funding separate and protects publication rights would address the clearest conflict, and the evaluators interviewed are right to ask for it.

Worth flagging is the longer arc seen in other technical industries. Independent audit roles often start as voluntary access deals and only later become lasting institutions. That change is slow and often disputed. The hopeful result is that steady, public-facing evaluation improves safety research and engineering discipline, because claims about alignment can be tested by people who do not report to product leaders.