Anthropic has published a measurement framework for the process of building frontier AI systems.
The proposal tracks how much AI performs AI research, how agents are monitored, and how compute is allocated.
It is a disclosure about development operations, not a claim that a model is independently building its successor.
AI made PMs faster. Multiplayer mode is still broken.
A PM can summarize research, draft a PRD, and mock up a prototype before lunch. The hard part starts when the team has to decide what actually gets built.
Jira Product Discovery gives product teams one place to capture insights, prioritize ideas with consistent frameworks, and build living roadmaps stakeholders can rally around.
And because it’s connected to Jira, the context behind every decision stays with the work—so developers and their agents know not just what to build, but why.
AI helps PMs move faster. Jira Product Discovery helps the whole team build with confidence.
The first measure is called the Anthropic R&D Automation Index.
Anthropic says it catalogues AI research and development tasks, rates their automation, and aggregates those ratings.
The company uses levels that separate human work, collaboration, and AI led work.
As of August 2026, Claude led 26 percent of measured AI research and development work.
The company says more than 90 percent of measured work was at least collaborative.
It also says no measured subset was fully autonomous.
The trend is notable because the same index placed AI led work below 1 percent in February.
That comparison reflects Anthropic’s internal method and its own task categories.
It should not be read as a universal measure across AI laboratories.
Anthropic says cross lab comparisons need common definitions and independent checks.
The second measure concerns oversight of agents working on research and engineering tasks.
Anthropic reported approximately 30,000 agents on its most used internal platform during August.
The company says all actions pass through an online monitor before execution.
It also says all actions enter an offline monitor after they occur.
Online review is intended to block dangerous actions that could cause immediate harm.
Offline review is intended to identify slower patterns that may need further examination.
Anthropic analyzed more than one billion decisions from research and engineering agents during August.
It reported that 0.002 percent of decisions were blocked by the online monitor.
That rate is about one blocked decision in 47,000 decisions.
The company also reported roughly 100,000 offline transcript flags each week.
About 50 high priority flags per week were escalated to human review.
These are company measurements, and the post does not establish how missed behavior was measured.
The third measure concerns compute allocation.
Anthropic examined a snapshot covering July 13 through July 20.
It reported that about 6 percent of AI research and development compute went to safety work.
It reported that about 12 percent of compute for AI driven AI research went to safety work.
The company calls these conservative estimates because safety and capability work can overlap.
The figures also exclude safeguards classifiers, which Anthropic treats as a separate category.
200+ Proven Ways to Make Money With AI in 2026
The next wave of millionaires will be people who figured out how to make AI work for them.
The window to get ahead is still open. But not for long.
Here are 200+ proven ways to make money with AI in 2026.
Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.
The methodology contains limits that change how the numbers should be interpreted.
Anthropic used internal records, sampled work activity, and its own models to organize tasks.
The company acknowledges that its judging model could make errors similar to the evaluated model.
Its compute categories also rely on best effort workload labels.
Those limits do not make the measurements useless, but they prevent simple league tables.
Anthropic says it plans to embed independent third party evaluators inside the company.
It wants those evaluators to review safety practices, incidents, and the published measurements.
That proposal is not the same as an existing external audit.
The difference matters because disclosure and verification answer separate questions.
The useful development is the choice of measurable operating signals.
Public reporting can show how automation, oversight, and safety resources change together.
It cannot show the full safety of a model from production statistics alone.
For now, the data support closer scrutiny of AI development speed and monitoring capacity.
The next test is whether other laboratories publish comparable records using shared definitions.
Learn AI in 5 minutes a day
You don't have to scroll every AI thread, track every new tool, or watch every demo.
The Rundown AI breaks it all down for you — the latest AI news, tools, and tutorials in one free 5-minute email every morning.
Trusted by 2M+ professionals at Apple, Google, and NASA.





