Technology
Danish Kapoor
Danish Kapoor

Claude’s role in Anthropic turned out to be bigger than expected

Anthropic shared remarkable data about the extent to which Claude is used in AI research and development within the company. According to the company’s new blog post, in 26 percent of Claude Anthropic’s artificial intelligence R&D work, it can carry out the work largely end-to-end, starting from a high-level command, and human employees supervise this process. Anthropic defines this way of working as “AI-led”, that is, work guided by artificial intelligence. According to the information provided by the company, Claude’s role is not limited to this 26 percent. AI systems perform at least large portions of the work under close human direction in more than 90 percent of Anthropic’s research activities.

These rates indicate that Claude was somehow involved in much of the research work done by the company’s employees. However, Anthropic underlines that Claude did not operate completely autonomously in any part of the measured AI R&D work. In other words, although the system can complete some tasks largely on its own from start to finish, human oversight is not removed from the process. At the 26 percent “leadership” rate defined by the company, final control remains with human employees. This distinction also reveals that looking only at task completion rates is not enough when measuring the effectiveness of artificial intelligence tools in research processes.

Anthropic proposes three ways to measure AI progress

Anthropic created the data in question as part of the first of three measurement methods aimed at making the speed of artificial intelligence development more understandable and comparable. In this approach, called “AI-led AI R&D”, it is measured how much of the artificial intelligence research and development work is carried out by Claude. For this, the company combined its own usage data with the automation rating scale prepared by Epoch AI. Thanks to this method, a chart was created showing the automation levels Claude has achieved since August 2025. According to Anthropic, a similar measurement could be applied by other advanced AI model developers using its own data, and the results could be verified by independent third parties.

The second area of ​​measurement proposed by the company focuses on the extent to which AI agents are audited. Here, it is recommended to track data such as how much of the agent’s activities are monitored by people, how long it takes for the transactions to be checked, and how often the system behaviors are marked as problematic. Anthropic thinks such indicators could provide more concrete information about the actual workings of increasingly autonomous AI agents. In addition, he argues that instead of simply measuring whether an agent has completed a certain task, the level of supervision under which this task was carried out should also be recorded. Thus, it is aimed to create a security and surveillance framework that can be more directly compared between companies.

The third measurement involves tracking the computational capacity allocated to artificial intelligence research and development activities. Anthropic states that disclosing how much computing power companies allocate to R&D efforts can be used to understand the speed at which advanced models are being developed. If these data are evaluated together with automation level and human control measurements, a broader picture can emerge about the development pace of the sector. The company is of the opinion that higher transparency can help the public better monitor artificial intelligence development processes and intervene earlier in developments that may get out of control.

Anthropic CEO Dario Amodei has also recently argued that more measurements should be made regarding artificial intelligence security and monitoring of advanced models. According to the source text, OpenAI has previously stated that it is positive about the idea of ​​slowing down the development of artificial intelligence when necessary. Anthropic, on the other hand, is committed to allowing development applications to be reviewed by third-party evaluators. Despite this, it remains unclear how the binding control mechanisms to be implemented throughout the sector will be shaped. The source text states that US President Donald Trump evaluates the risks of artificial intelligence in a more limited framework; Therefore, there is not yet a clear picture of the extent to which the regulatory approach will go beyond companies’ own internal controls.

The 26 percent rate announced by Anthropic is one of the concrete examples showing that current artificial intelligence systems are moving away from being auxiliary tools that only write code or perform singular research tasks. However, the company’s own data reveals that Claude does not operate as a fully independent investigator, with human guidance and supervision still decisive. For this reason, it is not correct to interpret the announced figures directly as “one quarter of the jobs are left entirely to artificial intelligence”. A more accurate assessment is that Claude is able to complete more and more tasks from high-level instructions, while human supervision remains a key element of the process. Whether Anthropic’s proposed measurement standards will be adopted industry-wide will depend on how much detailed data AI companies are ready to share.

Danish Kapoor