In partnership with

Anthropic has published a measurement framework for describing how artificial intelligence participates in frontier model development. The September 17 paper focuses on three questions. How much research is performed by AI? How are agents overseen? How is computing divided between capability and safety work?

For product teams moving at AI speed.

AI makes it easier to ship anything, even bad ideas. The hard part is knowing which ideas are worth building.

Jira Product Discovery brings your ideas, customer insights, and priorities into one place, so your team can decide what to ship and move forward with confidence.

Capture ideas, prioritize with evidence, and build living roadmaps your team can rally around—all while staying connected to delivery in Jira, so everyone can see what’s being built and why.

Better product decisions in the AI era.

The first measure classifies AI involvement on a five-level scale. At one level, humans perform the work. At higher levels, systems collaborate, lead tasks, or operate without human participation. Anthropic says that, as of August 2026, Claude led 26 percent of its measured AI research and development work. The company says more than 90 percent was at or above collaboration.

Those figures describe Anthropic’s internal measurement system. They are not a universal measure of every laboratory’s automation. The paper says the index required a map of AI research tasks, an automation rating, and weights based on the importance of each task. Anthropic used work records and Claude to organize the task hierarchy.

The second measure concerns oversight. Anthropic reports approximately 30,000 research and engineering agents on its most-used internal platform at one time in August. Online monitors check activity in real time. Offline monitors inspect activity afterward. The two layers address different timing problems, from quickly reversible intervention to slower signals that may need later analysis.

Anthropic also reports a blocked-action rate for monitored activity, but that figure depends on the company’s platform, classifier definitions, and sampled period. The measurement paper says monitoring cannot yet guarantee a complete picture of agent behavior. That admission matters because a low recorded rate may reflect detection limits as well as low-risk activity.

The third measure tracks compute allocation. Anthropic says about six percent of its AI research and development computing was assigned to safety work in August. The company describes this as a conservative estimate. Work with equal capability and safety purposes was counted as research and development rather than safety.

Stop Dabbling With AI and Start Earning With It

Ready to stop using AI as a search engine and start using it as an income engine?

The Hustle's "200+ AI-Powered Income Ideas" is your free playbook for turning the most overhyped technology of our time into actual cash.

Inside you'll find:

  • 200+ real, vetted ways to generate income with AI, spanning freelance services, digital products, content creation, and emerging markets

  • Actionable strategies built for non-engineers, so your technical background (or lack of one) won't hold you back

  • Ideas aligned with where the market is actually heading, not where it was two years ago

Subscribe free today and unlock the full guide. The people already cashing in aren't smarter than you, they just started earlier.

The method combines accelerator monitoring with workload labels and a classifier. Anthropic sampled about 14 percent of nearly 10,000 research runs for one part of the calculation. It also classified some agent sessions through transcripts, team information, or conservative defaults. These choices make the estimate inspectable, but they also create points where another laboratory might classify the same work differently.

That transparency is useful because the reported percentages are not directly comparable with benchmark scores. A benchmark measures task performance under fixed conditions. Anthropic’s index measures task ownership inside one organization. The two types of evidence answer different questions. One concerns what a model can do in a test. The other concerns how much work a lab assigns to models in practice.

Repeated measurement could reveal whether automation grows, stalls, or shifts between research categories. It could also expose whether safety work receives a stable share of computing as capabilities change. Those comparisons would require consistent definitions, independent review, and enough disclosure to distinguish task counts from task importance. Without that structure, a percentage can sound precise while remaining difficult to compare.

The paper’s wider argument is about public visibility. Anthropic says other developers could publish similar measures using shared definitions and third-party checks. The Associated Press reported the same headline findings and noted that Claude was not operating fully autonomously for the measured research work. That independent account helps separate the company’s claim from the basic description of its limits.

The evidence does not establish recursive self-improvement, a cross-lab scale, or a forecast for when human oversight becomes unnecessary. It establishes a more modest fact. One frontier laboratory has released internal measurements for AI participation, oversight, and compute, while acknowledging that comparison standards remain unfinished.

1,000+ Claude Prompts Top Professionals Actually Use at Work

Claude can be your analyst, editor, and strategist.

But most professionals are using it to fix grammar.

These 1,000+ Claude prompts take it from grammar tool to your most powerful AI work assistant.

Sign up for Superhuman AI and get:

  • 1,000+ ready-to-use Claude prompts to get real work done in minutes — researched, tested, and used by professionals at Google, Microsoft, and NASA

  • Superhuman AI newsletter (4 min daily) so you keep learning new AI tools and skills to stay ahead in your career — the prompts are just the beginning