Autonomy

Anthropic Is Paying the Firm That Will Evaluate It. It Said So in the Announcement.

CRAZE CRAZE Summary 3 things to know
  • Anthropic named Accenture its first embedded evaluator, funding the firm that red-teams and safety-tests its own frontier models.
  • Accenture is also Anthropic's largest Claude deployment and reseller, so the evaluator's invoice goes to the evaluated.
  • Anthropic admits no funding standards exist, says government or pooled money should eventually pay evaluators, and promises more soon.
Emon Editorial | · 4 min read
Anthropic Is Paying the Firm That Will Evaluate It. It Said So in the Announcement.

On September 18, Anthropic announced that Accenture would embed evaluators inside the company to red-team its frontier models, run alignment assessments, and test safeguards. Both companies said they expect to invest at least $1 billion over five years.

The announcement is the first concrete implementation of a commitment Anthropic CEO Dario Amodei made six days earlier in his “We Must Pace the Frontier” essay: place independent evaluators inside AI labs with access comparable to employees.

Anthropic's blog post also contains a sentence that describes the structure it just created. “Long-term, we think funding should come from pooled or government sources,” it says. “As neither exists today, we plan to work with different evaluators under different funding arrangements”.

Anthropic will fund Accenture's work directly. The evaluator‘s invoice goes to the company being evaluated.

The Evaluator Is Also the Largest Customer

Accenture is not an arms-length third party. In December 2025, Anthropic and Accenture formed the Accenture Anthropic Business Group, a dedicated practice built around Claude. Roughly 30,000 Accenture professionals are being trained on Claude, and tens of thousands of its developers use Claude Code — what Anthropic has called its largest ever deployment.

The two companies also run a joint Claude center of excellence and co-develop offerings for regulated industries including financial services, healthcare, and the public sector.

The Next Web summarized the arrangement: “Accenture is a customer, a reseller, an implementation partner, and now the evaluator”.

Anthropic does not treat this as a conflict. It presents Accenture’s enterprise deployment experience as a qualification, arguing that “understanding of how enterprises use AI in practice informs their safety approach”. TechCrunch noted that the choice of Accenture “surprised many AI watchers,” since discussions around embedded evaluators had focused on AI safety research organizations like METR, Redwood Research, and Apollo Research.

Anthropic says more evaluators will be announced in coming weeks and that it is in dialogue with METR about piloting embedded evaluation using METR‘s own funding.

Anthropic Is Paying the Firm That Will Evaluate It. It Said So in the Announcement.
Anthropic will pay Accenture to embed evaluators inside Anthropic, while Accenture remains Anthropic's largest Claude deployment.

The Market Read It as a Growth Story

Accenture shares jumped nearly 9% after hours on the announcement, reversing a 4.73% decline during the regular session after a Guggenheim downgrade. The move reflected investor interest in AI evaluation as a new revenue category for the consulting giant.

The combined $2 billion figure is not a contract paid to Accenture. It represents the two companies’ expected investments in building evaluation capacity over five years.

The Standards That Don‘t Exist

Anthropic’s blog post concedes that “there are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find”.

The AI Evaluator Forum — which released its own list of minimum conditions the same day — specifies that evaluation institutions should not have significant commercial dealings with the companies being evaluated, should not receive compensation tied to findings, should have access equal to senior insiders, and should be protected from retaliation.

Accenture meets the first condition in neither direction. It has substantial commercial dealings with Anthropic, and Anthropic funds its work directly.

Anthropic‘s framing is that independent evaluators “do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility”. That is a statement of responsibility retention. It is not a statement about what happens when an evaluator with employee-level access finds something that would delay a release — inside a company whose larger business is selling that release to clients.

The Pattern Keeps Repeating

The organizations technically capable of auditing frontier models are the ones the industry wants to acquire or contract. Hugging Face volunteered to audit AI labs; Nvidia is buying Hugging Face. Accenture has the staffing scale to embed a standing team with employee-level access — a problem nonprofits the size of METR cannot solve.

Anthropic says the arrangement is non-exclusive and that more evaluators are coming. The test will be whether any of them arrive self-funded. If none do, the “ecosystem of evaluators” Anthropic describes will consist of firms whose invoices are paid by the companies they assess.


P.S. Faculty, the Accenture unit leading the evaluation, was acquired by Accenture in March 2026. It had previously worked with leading labs including OpenAI and Anthropic on model safety. The evaluator evaluating Anthropic is a unit of a company that was already helping Anthropic deploy Claude.


Frequently Asked Questions

Q: What did Anthropic announce?

A: On September 18, Anthropic named Accenture as its first embedded evaluator. Accenture will place teams inside Anthropic to red-team frontier models, run alignment assessments, and test safeguards. Both companies expect to invest at least $1 billion over five years.

Q: Why does it matter that Anthropic funds the evaluator?

A: Anthropic's own blog post states that paying your evaluator directly “is not the right long-term arrangement,” and that funding should come from pooled or government sources — which don‘t exist yet.

Q: What is Accenture’s relationship to Anthropic?

A: Accenture is Anthropic‘s largest Claude Code deployment, with roughly 30,000 professionals being trained on Claude. The two companies run a joint business group and a Claude center of excellence.

Q: How did the AI safety community react?

A: TechCrunch reported that the choice “surprised many AI watchers,” since embedded evaluator discussions had focused on research organizations like METR, Redwood Research, and Apollo Research.

Q: What standards exist for embedded evaluators?

A: None. Anthropic concedes there are “no standards for what information embedded evaluators should have access to, or how they should report what they find.”

Q: What conditions did the AI Evaluator Forum propose?

A: No significant commercial dealings with the evaluated company, no compensation tied to findings, access equal to senior insiders, minimized NDAs, and protection from retaliation. Accenture meets neither of the first two conditions.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article