Anthropic's R&D Automation Index

usefulAnthropicreport2026-09read ai llmssoftware engineeringphilosophy

Synopsis — AI-drafted from Dan's notes

Anthropic published internal numbers on how much of its own AI research its models now carry. As of August 2026, Claude is rated as leading roughly a quarter of that work (26%), and more than nine tenths of it sits at or above the level the index calls collaboration. The baseline was under one percent in February of the same year.

The method matters more than the headline. Tasks are rated against an automation scale developed by Epoch AI, then weighted by how much person-time they represent. The scale is third-party. The judging is not: Anthropic’s own models do the rating, and the publication says so directly, including the admission that the judge model could make the same class of error as the model it is grading.

The rest is oversight arithmetic. Roughly thirty thousand agents run concurrently, over a billion of their decisions were analysed across August, and two thousandths of one percent were blocked in real time. About a hundred thousand transcripts are flagged for offline review each week, of which roughly fifty reach a human. Safety is given about six percent of AI R&D compute, and about twelve percent of the compute for AI-driven AI R&D, on estimates the publication calls deliberately conservative.

Where I land

I need to sit with this one for a bit. The shape of the objection is clear enough to write down, though.

It is important to have some external judging (another person, at least) because of your blind spots. Having an external reviewer on both the process and the outputs is important. And that measure should matter more than time saved. At the same time, I think we are all trying to figure out how to apply this technology at scale, scientifically.

Anthropic having external reviewers is fine, and adopting that into their workflow is the right move: just as long as they do not chase the percentage of adoption as the only metric worth chasing.

I intend to hold myself to the same standard. My thinking is to share my thoughts here and then have that step occur in the wild (on Mastodon, for instance) as a form of feedback from the external world. I picked it because it has lots of people with tech-based values like my own. However, I have to release it before that happens.

Connections