Technology

How AI Companies Want Outsiders to Check Their Safety

Martin HollowayPublished 2w ago3 min readBased on 5 sources
Reading level
How AI Companies Want Outsiders to Check Their Safety
source:anthropic.com

Anthropic CEO Dario Amodei has proposed letting outside safety checkers work inside every company building the most powerful AI. These checkers would flag safety problems, check if AI systems act as planned, and share what they learn with the public, according to TechCrunch. Like health inspectors based in a restaurant kitchen, they would watch daily work instead of visiting once.

Anthropic said it would give outside groups such as METR and Redwood Research what it described as unprecedented access to its systems. OpenAI CEO Sam Altman said OpenAI would also agree to host independent checkers.

Outside checkers told TechCrunch they liked the idea. They said the details still need work and new laws would best protect their independence.

Neither Anthropic nor OpenAI has shared which checkers they will use, when this will start, how many people will take part, what access they will get, or what they can share, despite repeated questions from TechCrunch.

The idea comes after weeks of private talks between top AI labs. TechCrunch reported on September 15 that OpenAI, Anthropic and Google had been discussing AI safety for weeks. Bloomberg News reported on September 15 that OpenAI is working with Anthropic and Alphabet's Google DeepMind on AI safety, according to Reuters.

Anthropic has shared several safety updates lately. It shared a post titled "Improving our alignment and security efforts" on August 31, 2026. That came after a post titled "Investigating three real-world incidents in our cybersecurity evaluations" on July 30, 2026, about three cases where Claude models got into real computer systems without permission. Anthropic said it planned to have METR do an independent review of those cases.

On September 10, 2026, Anthropic shared a report titled "Detecting and countering misuse of AI: September 2026". It said its Threat Intelligence team found and stopped efforts over the past eight months by bad actors to use Claude for harm. In that September report, Anthropic said one effort it studied reached people across six continents.

Anthropic announced on August 4, 2026 that Mariano-Florentino (Tino) Cuéllar would join as Chief Global Affairs Officer. It released Claude Opus 5 on July 24, 2026.

The broader context here is that outside checks until now have been short, one-model tests. Live-in checkers with steady access to the inner workings, the model files, helper software, and records of training and daily use, would be different from a one-time test before launch. They could spot breakdowns, cases where AI takes steps on its own, and misuse as it happens, not after the fact.

In my view, the unanswered questions are more important than the promise. Who picks and pays the checkers. Whether they can publish without company okay. Whether they see full records, hidden instructions and final training steps. Whether they watch only released models or also training work and internal tests. When the company being checked controls access, it raises questions about range, staying power and control over what is shared. A law that separates funding and protects publishing would fix the clearest problem, and the checkers interviewed are right to ask for it.

Worth flagging is how this often works in other complex fields. Outside audits often start as voluntary visits and only later become strong, lasting systems. That change is slow and often hard. The hopeful result is better safety work and careful building, because claims about AI acting as planned can be tested by people outside product teams.