Anthropic says its Claude AI system now “leads” 26% of the company’s AI research-and-development work, a figure that sounds dramatic but comes with an essential qualifier: humans remain in the loop.
In Anthropic’s use of the term, an AI-led task is one Claude can carry through mostly from end to end after receiving a high-level prompt, while a human supervises the process. The company also says AI handles at least large portions of more than 90% of its R&D work under close human direction. That broader number includes the 26% share where Claude is considered the lead.
The distinction matters. Neither figure means that Claude independently runs a quarter of the company’s research operation, makes final decisions without people, or is autonomous across a defined part of the R&D pipeline. Anthropic explicitly says Claude is not fully autonomous for any measured subset of that work. Still, the disclosure presents a concrete picture of how deeply a frontier AI developer says it has embedded its own model into the process of building and improving AI.
The numbers arrive alongside Anthropic’s proposal for three measurements intended to communicate the pace and character of advanced AI development: the degree to which AI conducts AI R&D, the quality and intensity of oversight for AI agents, and the amount of computing power assigned to AI R&D. The central idea is straightforward: claims about powerful AI are easier to evaluate when companies publish understandable indicators rather than relying on broad assurances.
What “leads” means in this measurement
It is tempting to translate “Claude leads 26%” into “Claude does 26% of the work alone.” Anthropic’s definition is narrower and more cautious than that.
For a task to be AI-led in this framing, Claude must be able to complete most of it from a high-level instruction, with a person supervising. A high-level prompt is an objective or direction rather than a complete, click-by-click procedure. The point of the metric is to describe how much initiative the model can take on a task, not to erase the human reviewer from the workflow.
That makes the 26% figure a measure of task-level automation under supervision. It says something about the model’s ability to execute substantial pieces of work with less granular human steering. It does not, on its own, reveal the difficulty, duration, stakes, or downstream impact of each task. A percentage of tasks can be useful, but it is not automatically a percentage of total research value, labor hours, or critical decision-making.
The company’s second figure is wider: AI performs large chunks of work under close human direction on more than 90% of its research. That could cover work where people provide more detailed instructions, review intermediate steps more actively, or retain tighter control over how the task proceeds. Put together, the two figures describe a spectrum rather than a binary choice between “human work” and “AI work.”
- AI-led work: Claude completes most of a task from a broad prompt while a human supervises.
- Closely directed AI work: Claude contributes major portions, but people guide it more directly.
- Fully autonomous work: A system operates without the relevant human oversight. Anthropic says this does not apply to any measured slice of its AI R&D.
That last category is especially important in a safety discussion. An AI agent is generally a system that can pursue multi-step goals and take actions, rather than merely return a single answer. More agentic behavior can be useful, but it also changes what needs to be checked: not just whether an answer is correct, but what the system did while pursuing an instruction, whether its actions were authorized, and whether a human had meaningful opportunity to intervene.
A proposed yardstick for AI-led AI research
Anthropic derived its R&D statistic through the first of its proposed measurements, focused on “AI-led AI R&D.” It combined an index of the amount of research and development performed by Claude with an automation rating scale developed by Epoch AI. The result is a chart intended to show Claude’s automation level over time, beginning in August 2025.
An automation level in this context is not a claim that a system has crossed a clean threshold into independence. It is an attempt to express the degree to which the model can carry out research-related work with less step-by-step human input. This is useful because raw capability demonstrations can be hard to compare. A model may perform impressively on a benchmark yet still require extensive prompting, correction, or review to be dependable in an actual R&D process.
Anthropic’s stated goal is for other frontier model developers to be able to reproduce the approach using their own internal data, with third-party validation. Reproducibility is a meaningful part of the proposal. A company’s internal dashboard can provide information, but it also reflects that company’s definitions, task classifications, and reporting choices. Independent validation would not eliminate judgment calls, but it could make the metrics more credible and more comparable.
There are also practical questions any such framework will need to answer. What precisely counts as a task? How are tasks divided when one larger project contains many AI and human steps? What qualifies as “most” of a task? How is supervision recorded when a person checks only final output versus monitoring an agent’s intermediate actions? The supplied figures do not resolve those methodological details, so they should be treated as company-reported indicators rather than a complete audit of AI labor or autonomy.
Even so, publishing an operational definition is more informative than using vague language about systems becoming increasingly capable. It gives observers something specific to interrogate: a stated boundary between supervised AI-led work, closely managed assistance, and full autonomy.
Why oversight needs its own measurement
The second proposed measurement concerns AI-agent oversight. Anthropic suggests tracking how much agent activity is monitored, the time required for that activity to be reviewed, and the frequency with which agent behavior is flagged.
This moves beyond the familiar question of whether a chatbot produces a convincing response. If an agent can work through a sequence of actions, oversight has at least three dimensions:
- Coverage: How much of the agent’s activity is actually monitored?
- Review speed: How long does it take before a person evaluates the work or actions?
- Flag rate: How often does monitoring identify behavior requiring attention?
Each dimension tells a different story. High monitoring coverage could mean that supervisors have strong visibility, but it does not prove that they can keep up with the volume of activity. Fast review could be valuable, but only if the review is sufficiently rigorous. A high flag rate may reveal problematic behavior, or it may indicate that a monitoring system is sensitive and actively catching issues. A low flag rate may mean the system is behaving well, but it could also mean problems are not being detected. Metrics need context, not just a dashboard color.
That is why measurements of oversight should not be confused with safety guarantees. They can expose whether monitoring exists and whether it is finding concerns, but they cannot by themselves settle whether the rules, review processes, and escalation paths are adequate. Their value is that they make the question more concrete.
For readers following AI’s effects on games, software, and digital creative work, the practical takeaway is similar. The relevant question is increasingly not simply whether an AI tool was used. It is what authority the tool had, what work it could initiate, what it could change, and who checked those changes before they mattered. That same focus on agent guardrails is visible in efforts to frame AI systems around bounded responsibilities, including an experimental AI-agent approach built around guardrails.
Compute is a signal of development pace, not a verdict
Anthropic’s third proposed indicator is the amount of compute devoted to AI R&D. Compute refers to the computational resources used to train, run, and experiment with AI systems. In this proposal, tracking the share or level of compute going toward AI R&D is meant to help indicate how aggressively frontier development is progressing.
Compute is a useful signal because advanced AI work depends on substantial computational resources. A rising commitment to R&D compute can suggest that a company is expanding experimentation or placing more resources behind development. It is not, however, a direct measure of model quality, real-world reliability, safety, or the arrival of a particular capability. More compute does not automatically explain what a model can do, and less visible compute does not prove that development has slowed in every relevant sense.
Its main role is therefore contextual. If AI-led R&D rises, oversight data changes, and compute commitments increase at the same time, the public and outside evaluators could have a clearer basis for assessing the overall tempo of development. A single measure is easy to overread; several measures can provide a more useful, if still incomplete, picture.
Transparency does not settle the regulation question
Anthropic’s proposal is built around transparency and third-party review, not a description of a binding external regulatory system. The company has committed to allowing third-party evaluators to examine its development practices. Other major AI companies have also expressed support for slowing or managing development in principle. The harder and unresolved question is whether self-regulation will be supplemented or replaced by enforceable outside rules.
That uncertainty is heightened by the political backdrop. President Donald Trump has largely minimized AI risks, while prominent technology figures have expressed agreement with calls for AI safety. Agreement on general principles is easier than agreement on who sets the standards, what must be disclosed, how compliance is checked, and what happens when a company fails to meet the standard.
Anthropic’s figures should therefore be read in two ways. First, they are a notable claim about present-day AI use inside an AI lab: Claude reportedly has a leading, supervised role in 26% of the company’s AI R&D tasks and contributes substantial work across more than 90%. Second, they are an argument for a reporting framework that could let outsiders track changes over time.
Whether that framework becomes broadly adopted, independently validated, and paired with meaningful accountability remains open. For now, the most valuable part of the disclosure may be its insistence that debates about fast-moving AI need definitions that distinguish assistance from supervised task leadership, and supervised task leadership from actual autonomy.






