Anthropic Gives Accenture Inside Access to Test Its Models

Anthropic named Accenture as its first embedded evaluator on Sept. 18, a role that puts Accenture staff inside Anthropic to review its models and staff. TechCrunch
Unlike outside reviewers, embedded evaluators work inside the AI company with access similar to an employee's. Anthropic That shifts testing from checking selected outputs through limited, time-boxed API access to watching systems, workflows and personnel directly.
The work will go through Faculty, the AI group Accenture acquired in January to serve as its AI division. TechCrunch Faculty staff will evaluate models, red-team them, which means deliberately trying to trick or break them, run alignment assessments, which check whether model behavior matches intended goals, and test safeguards, the controls meant to prevent misuse.
Anthropic and Accenture expect to put at least $1 billion into the project over the next five years. TechCrunch Accenture shares rose 8% after hours following the announcement.
Anthropic said more embedded evaluators will be named in the weeks ahead. It is also talking with METR and other nonprofit groups about testing parts of embedded evaluation with their own funding. The setup is therefore one commercial evaluator to start, more to follow, plus a separate nonprofit track paid for independently.
The broader context here is what inside access lets testers see. Outside testing is limited to what can be learned through interfaces, documents and short testing windows. A resident team can watch continuously for failure modes, attempts to bypass safeguards, how alignment holds when real use differs from training, known as distribution shift, and whether protections still work after models and operations change.
In my view, the practical design questions matter more than the headline number. When the tester lives inside, independence has to be built in. That includes clear separation of duties, access to versioned models and test logs, the ability to reproduce red-team results, and a clear record of how disagreements about alignment or safeguards are settled. A separately funded nonprofit track could balance a large commercial project, if both use similar methods and disclosure rules.
Looking at what this means for companies that deploy these models, embedded testing could bring lab assurance closer to everyday IT governance. If safeguard and alignment results come with inside detail, security and platform teams can connect them to threat models, access controls, incident response and decisions about when a release is ready. The long arc looks positive. Better measurement inside frontier labs should make outside use safer and more predictable, as long as the results are clear to people who were not in the room.


