Anthropic and Accenture have announced a five-year partnership to create an embedded team that will evaluate and red-team Anthropic's frontier AI models.
Each company expects to invest at least $1 billion over the period. Accenture will draw on Faculty, the AI specialist it acquired, to conduct alignment assessments and test model safeguards alongside Anthropic's internal teams and other safety partners.
This is a major investment in a category enterprise procurement will soon have to understand: independent AI evaluation.
Model assurance is moving outside the supplier's own control environment
Conventional technology due diligence asks whether a supplier has security certifications, documented development controls and a testing process.
Frontier models create a harder question. Their behaviour can change across prompts, tools, environments and deployments. A buyer cannot rely solely on a point-in-time questionnaire or the developer's own safety report.
Embedded evaluators may gain deeper access to models, staff and development processes than a conventional auditor. That could produce more meaningful evidence about alignment, misuse, cybersecurity and real-world failure modes.
Embedded is not automatically independent
Anthropic and Accenture describe the work as independent evaluation. The partnership is non-exclusive, and Anthropic says it intends to work with other evaluators.
Yet the commercial structure still deserves scrutiny. An evaluator working inside the developer and investing alongside it may have better access, but buyers should understand who defines the test scope, who owns the findings and what can be published when the results are uncomfortable.
Procurement teams buying AI assurance should test:
Whether the evaluator has unrestricted access to relevant models, logs and employees.
Who selects the scenarios and failure thresholds.
Whether the supplier can restrict publication or disclosure to customers.
How conflicts of interest and other commercial relationships are managed.
Whether a failed evaluation blocks release or merely informs it.
A new clause belongs in AI contracts
Customers should not have to wait for a public incident to learn that a model failed an important external evaluation.
High-impact AI agreements increasingly need disclosure obligations covering material evaluation findings, changes to safeguards, known failure modes and the remediation status. Buyers may also require access to an independent summary that explains the methodology and limitations without revealing sensitive model details.
The bottom line
The scale of this commitment suggests external model evaluation is becoming infrastructure rather than a niche research activity. It also creates a new market of evaluators, standards and assurance reports that procurement will need to compare.
An AI supplier saying “we tested it” is no longer enough. The next question is who tested it, what access they had and whether the buyer gets to see what failed.
Sources
Accenture and Anthropic partnership announcement, 18 September 2026.
Reuters: Anthropic and Accenture invest in model evaluation, 18 September 2026.

