In partnership with

Anthropic has announced a new model for evaluating frontier artificial intelligence from inside the company developing it.
The September 18 announcement names Accenture’s specialist AI business, Faculty, as the initial partner.
Anthropic describes the arrangement as independent evaluation, although Anthropic will directly fund the first partnership.
The proposal combines deeper access with unresolved questions about how independence will work in practice.

AI made PMs faster. Multiplayer mode is still broken.

A PM can summarize research, draft a PRD, and mock up a prototype before lunch. The hard part starts when the team has to decide what actually gets built.

Jira Product Discovery gives product teams one place to capture insights, prioritize ideas with consistent frameworks, and build living roadmaps stakeholders can rally around.

And because it’s connected to Jira, the context behind every decision stays with the work—so developers and their agents know not just what to build, but why.

AI helps PMs move faster. Jira Product Discovery helps the whole team build with confidence.

The planned work includes model evaluation, red-teaming, alignment assessments, and safeguard testing.
Anthropic says embedded evaluators would work inside the company with access comparable to an employee.
That access could let evaluators observe training decisions, deployment choices, and internal discussions.
It could also let them follow risks that are difficult to see from a short external assessment.

Anthropic and Accenture each expect to invest at least one billion dollars in capacity over five years.
Those figures are company expectations, not an independently audited measure of completed spending.
The announcement does not specify how many evaluators will work on the program or when the first assessment will appear.
It also does not set a public timetable for reporting findings.

Anthropic acknowledges that embedded evaluation is new and that many operating details remain unsettled.
The company says there are no established standards for the information evaluators should receive.
There is also no settled system for how evaluators should report what they find.
These gaps are part of the design problem rather than minor administrative details.

The funding relationship creates a direct accountability question.
Anthropic says its safety responsibility remains unchanged, while the evaluator receives funding from Anthropic.
That arrangement may provide access and resources, but it can also create pressure around scope, timing, and publication.
An evaluation cannot be fully independent merely because the evaluators sit outside the company’s reporting structure.

200+ Proven Ways to Make Money With AI in 2026

The next wave of millionaires will be people who figured out how to make AI work for them.

The window to get ahead is still open. But not for long.

Here are 200+ proven ways to make money with AI in 2026.

Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.

The strongest version of the proposal would make those pressures visible.
It would define who selects evaluators, what evidence they can inspect, and when access can be limited.
It would also specify how disagreements are recorded when an evaluator challenges a company decision.
Public reporting rules would show whether findings can be delayed, shortened, or withheld.

Accenture’s enterprise experience gives the arrangement a practical perspective on how organizations use AI.
That perspective may help evaluators examine deployment safeguards beyond laboratory benchmarks.
It does not by itself establish technical independence or scientific validity.
Those questions require methods that other researchers can inspect and, where possible, reproduce.

The announcement is non-exclusive, according to Anthropic.
The company says it plans to work with other evaluators, while Accenture may work with other AI developers.
Anthropic is also speaking with METR and other nonprofit evaluators about pilot work using their own funding.
That wider ecosystem could reduce reliance on a single firm if standards become shared and public.

The proposal should therefore be judged by records produced over time.
Useful evidence would include test protocols, evaluator independence rules, incident escalation paths, and published limitations.
It would also include examples of findings that changed training, deployment, or monitoring decisions.
An announcement creates a governance mechanism, but it does not yet demonstrate its performance.

The Associated Press reported support for embedded evaluators alongside criticism of their independence and standards.
That tension makes explicit test conditions and reporting obligations necessary.

Embedded evaluation could connect external scrutiny to decisions that ordinary audits cannot observe.
Its credibility will depend on whether access, funding, and publication rights remain visible to outsiders.
The next meaningful evidence will come from published assessments and documented responses to their findings.

The Future of AI in Marketing. Your Shortcut to Smarter, Faster Marketing.

This guide distills 10 AI strategies from industry leaders that are transforming marketing.

  • Learn how HubSpot's engineering team achieved 15-20% productivity gains with AI

  • Learn how AI-driven emails achieved 94% higher conversion rates

  • Discover 7 ways to enhance your marketing strategy with AI.