Anthropic has published the first hard numbers on how fast AI is building AI inside a frontier lab. As of August 2026, its model Claude leads 26 percent of the company's own AI research and development, up from under 1 percent in February. More than 90 percent of that work now involves the model as a collaborator or better. None of it is fully autonomous. For a contractor, the number that matters is not the 26 percent. It is the six months.
The report is Measurements for understanding the pace of AI development inside frontier labs, published this week by the Anthropic Institute and written by Marina Favaro and Phillie Wright, with research direction from Jack Clark. It says nothing about small businesses, occupations, or anything that happens outside an AI lab. This post is about why it matters anyway, and about the parts of it you should hold loosely.
What did Anthropic actually publish?
A measurement framework with the company's own numbers attached, built so that other labs could report the same way. Three measurements, each with a stated method and a stated limitation.
| Measurement | What it tracks | The headline number |
|---|---|---|
| AI R&D Automation Index | How much of Anthropic's research and engineering work Claude performs, on a six-level scale (AL0 no AI involvement, AL3 AI collaborates, AL4 AI leads, AL5 fully autonomous) | AL4 on 26% of AI R&D in August 2026, up from under 1% in February. AL3 or above on more than 90%. AL5 on none. |
| Oversight of AI agents | How automated monitors perform across roughly 30,000 internal agents — a real-time check before every action, and a slower review after | Over a billion decisions in August; 0.002% blocked before execution, about 1 in 47,000. Around 100,000 transcripts a week flagged for after-the-fact review; about 50 a week reach a human. |
| Compute allocation | What share of computing power goes to safety work versus capability work, over one week | About 6% of AI R&D compute went to safety, July 13–20, 2026. About 12% of the compute for AI-driven AI R&D. |
The stated purpose is transparency. In the report's words, "As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows."
The line the report itself draws under the chart: "Claude is not operating fully autonomously for any measured subset of AI R&D work."
Why should a contractor care what happens inside a frontier lab?
Because of the two clocks.
The first clock runs inside the lab. A system that led under 1 percent of its own development in February led 26 percent by August. Whatever you think of the method — and there are real questions, below — the direction is not in dispute, and Anthropic is one of a handful of companies whose systems answer when a homeowner in Waxahachie types "who should I call about a slab leak" into an AI assistant. The recommendation layer that decides whether your business gets named is built by the class of system that is now writing a quarter of its own research.
The second clock runs on a job site, and it is much slower. MIT's Industrial Performance Center spent three years watching more than twenty companies try to put generative AI to work, and its report — which we translated for contractors here — describes a technology that is "far more useful on some tasks than others," and in many companies "a hammer in search of a nail." The manufacturers it studied bought their AI from outside vendors and doubted it could meet their quality standards. The most common way AI arrived in most workplaces was in the background, embedded in software the company already used, with no decision to adopt it at all. And the much-discussed shift to AI agents, the researchers found, is "an incremental change from robotic process automation" rather than a new transformation.
So: the systems that decide who gets recommended are compounding on a six-month clock. The systems that run a contracting business are moving on a multi-year one. That gap is not a threat. It is the window. The engines are getting better at answering the question faster than most local businesses are getting better at being the answer, which means the businesses that make themselves findable, verifiable, and consistent now are competing against a field that mostly has not started.
Which of these numbers should you hold loosely?
Anthropic lists its own limitations, and they are substantial enough to repeat.
The model graded its own work. The automation index was produced by having Claude rate roughly 15,000 of the company's own tasks. When staff rated the same areas independently, the model matched them exactly 59 percent of the time — better than the 35 percent rate at which two humans matched each other, and within one level of the humans 97 percent of the time. That is a reasonable result and also a system judging itself, which the report acknowledges could produce correlated errors.
The task list is frozen. The basket of tasks was fixed to a July 2026 baseline. Any new kind of work that has appeared since is not in the measurement. The index can therefore describe how much of last summer's job the model does, not how much of this autumn's.
One week of compute is a snapshot, not a trend. The safety-allocation figure covers July 13 to 20. The report is explicit that safety research uses less compute by nature — the design work is time-intensive, not compute-intensive — so the 6 percent understates effort even before questions about labeling, and a single week cannot show direction.
Oversight has been running for months, not years. The monitors have not had time to encounter every pattern of behavior, and the report expects the methods to change significantly.
None of that undercuts the direction of the R&D number. It does mean the precise figures are best read as the first data points on a chart Anthropic has committed to keep updating, and it means a lab reporting on itself will need outside verification before the numbers should settle anything.
What should a contractor do with a report about a frontier lab?
Three things, none of which require understanding automation levels.
Be the answer before the engines get better at asking. Every one of these systems researches a business before recommending it — its website, its Google profile, the directories and license registries that vouch for it. We tracked 700 AI visibility checks in one Texas market and found the same company scoring 92 percent on one engine and zero on another, in the same window, on identical work. The difference was whether each engine's pipeline could see and verify the business. That is fixable, and it is worth more every month the engines improve.
Start from a problem you already have. The MIT finding that separated AI deployments that scaled from ones that got shelved was a pre-existing, documented problem the tool was pointed at. That holds whether the tool is a proposal drafter or an AI receptionist. A faster clock in the lab does not change the order of operations in your office.
Keep a person on the edge cases. The most useful detail in Anthropic's oversight numbers is not the 0.002 percent block rate. It is that about fifty transcripts a week still reach a human. Even the company that builds these systems runs them with someone on the exceptions. So should you.
The pace of AI development is now measurable, and the first measurement says a quarter. It will be higher next time. The businesses that will be fine are not the ones that predicted the number. They are the ones that were findable when the engines came looking.
Find out whether they can find you.
The AI Visibility Audit checks whether ChatGPT, Gemini, Perplexity, and Google's AI name your business when a customer asks who to call — and what to fix if they don't.
Get My AI Visibility Audit — $27Sources
- Favaro, M., and Wright, P. "Measurements for understanding the pace of AI development inside frontier labs." Anthropic Institute, September 2026. Report. All figures for the automation index, agent oversight, and compute allocation are Anthropic's self-reported measurements for July and August 2026.
- Epoch AI. The automation-level scale (AL0–AL5) used by the index. Toward an O*NET for AI R&D.
- Armstrong, B., Shah, J., et al. Humans in the Loop. MIT Industrial Performance Center, April 2026. Our contractor translation is at MIT's 10 Levers for Generative AI, Applied to Contractors.
- Direct quotations are from the Anthropic report and the MIT report. The two-clocks framing and every contractor application are ours.