Anthropic’s new measurement framework separates three questions: how much research AI leads, how agents are monitored, and where computing resources go. The company presents internal snapshots and methods for other labs to use.
AI-led does not mean autonomous
Anthropic says Claude led 26% of its measured AI research and development in August 2026, with humans supervising. No measured category reached fully autonomous operation. Its prototype index uses Claude to classify work and is weighted using an approximation of staff time.
Agent oversight figures cover its most-used internal platform, not every deployment. Monitoring coverage describes which actions pass through checks; it does not establish that every dangerous action is detected.
Keep the compute denominator attached
For July 13–20, Anthropic estimates safety received about 6% of AI-R&D compute and 12% of AI-driven AI-R&D compute. Those are different denominators. The estimates exclude safeguards classifiers and classify equally mixed safety/capability work conservatively.
The appendix cautions that workload labels can be wrong, one week cannot establish a trend, and resource use does not measure safety effectiveness. These are company-reported measurements, not an independent audit.
NextWith.ai perspective: ask what would change the conclusion
Our assessment is that a useful transparency report should make its numbers open to challenge. A percentage without a stable definition can rise because behavior changed, because the sample changed, or because someone moved the boundary between categories.
For readers comparing future reports, we would start with three questions. Is the same activity being measured? Is the observation period comparable? Can an outside reviewer reproduce the classification? If the answers are unclear, a tidy chart may conceal more than it explains.
Oversight requires a different test again. Imagine a monitoring system that examines every action but misses a particular kind of failure. Complete coverage would describe its reach accurately while saying little about that blind spot. This is an illustrative example, not a finding about Anthropic’s system.
The next meaningful signal would be a repeat measurement with consistent definitions and an explanation of any revisions, alongside external scrutiny of what the checks miss. That would help distinguish a change in the underlying work from a change in how it is counted.