Anthropic's Pace-of-AI Measurements: What 26% AI-Led R&D Means for Contractors

Anthropic just published the first hard measurements of how fast AI is building AI inside a frontier lab: Claude now leads 26% of the company's own R&D, up from under 1% in February. What that pace means for a business at the far end of the pipe — and why the gap is the opportunity.

Anthropic has published the first hard numbers on how fast AI is building AI inside a frontier lab. As of August 2026, its model Claude leads 26 percent of the company's own AI research and development, up from under 1 percent in February. More than 90 percent of that work now involves the model as a collaborator or better. None of it is fully autonomous. For a contractor, the number that matters is not the 26 percent. It is the six months.

The report is Measurements for understanding the pace of AI development inside frontier labs, published this week by the Anthropic Institute and written by Marina Favaro and Phillie Wright, with research direction from Jack Clark. It says nothing about small businesses, occupations, or anything that happens outside an AI lab. This post is about why it matters anyway, and about the parts of it you should hold loosely.

What did Anthropic actually publish?

A measurement framework with the company's own numbers attached, built so that other labs could report the same way. Three measurements, each with a stated method and a stated limitation.

MeasurementWhat it tracksThe headline number
AI R&D Automation IndexHow much of Anthropic's research and engineering work Claude performs, on a six-level scale (AL0 no AI involvement, AL3 AI collaborates, AL4 AI leads, AL5 fully autonomous)AL4 on 26% of AI R&D in August 2026, up from under 1% in February. AL3 or above on more than 90%. AL5 on none.
Oversight of AI agentsHow automated monitors perform across roughly 30,000 internal agents — a real-time check before every action, and a slower review afterOver a billion decisions in August; 0.002% blocked before execution, about 1 in 47,000. Around 100,000 transcripts a week flagged for after-the-fact review; about 50 a week reach a human.
Compute allocationWhat share of computing power goes to safety work versus capability work, over one weekAbout 6% of AI R&D compute went to safety, July 13–20, 2026. About 12% of the compute for AI-driven AI R&D.

The stated purpose is transparency. In the report's words, "As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows."

Share of Anthropic's own AI R&D that Claude leadsAUTOMATION LEVEL 4 — AI LEADS END TO END, HUMAN SUPERVISES · EPOCH AI SCALE · ANTHROPIC, SEPT 2026FEB 2026under 1%AUG 202626%← six monthsAL3 OR ABOVE (AI COLLABORATES OR BETTER), AUG 2026over 90%AL5 — FULLY AUTONOMOUS, NO HUMAN IN THE LOOP0%Source: Anthropic Institute, "Measurements for understanding the pace of AI development inside frontier labs" (2026). Self-reported.

The line the report itself draws under the chart: "Claude is not operating fully autonomously for any measured subset of AI R&D work."

Why should a contractor care what happens inside a frontier lab?

Because of the two clocks.

The first clock runs inside the lab. A system that led under 1 percent of its own development in February led 26 percent by August. Whatever you think of the method — and there are real questions, below — the direction is not in dispute, and Anthropic is one of a handful of companies whose systems answer when a homeowner in Waxahachie types "who should I call about a slab leak" into an AI assistant. The recommendation layer that decides whether your business gets named is built by the class of system that is now writing a quarter of its own research.

The second clock runs on a job site, and it is much slower. MIT's Industrial Performance Center spent three years watching more than twenty companies try to put generative AI to work, and its report — which we translated for contractors here — describes a technology that is "far more useful on some tasks than others," and in many companies "a hammer in search of a nail." The manufacturers it studied bought their AI from outside vendors and doubted it could meet their quality standards. The most common way AI arrived in most workplaces was in the background, embedded in software the company already used, with no decision to adopt it at all. And the much-discussed shift to AI agents, the researchers found, is "an incremental change from robotic process automation" rather than a new transformation.

So: the systems that decide who gets recommended are compounding on a six-month clock. The systems that run a contracting business are moving on a multi-year one. That gap is not a threat. It is the window. The engines are getting better at answering the question faster than most local businesses are getting better at being the answer, which means the businesses that make themselves findable, verifiable, and consistent now are competing against a field that mostly has not started.

Which of these numbers should you hold loosely?

Anthropic lists its own limitations, and they are substantial enough to repeat.

The model graded its own work. The automation index was produced by having Claude rate roughly 15,000 of the company's own tasks. When staff rated the same areas independently, the model matched them exactly 59 percent of the time — better than the 35 percent rate at which two humans matched each other, and within one level of the humans 97 percent of the time. That is a reasonable result and also a system judging itself, which the report acknowledges could produce correlated errors.

The task list is frozen. The basket of tasks was fixed to a July 2026 baseline. Any new kind of work that has appeared since is not in the measurement. The index can therefore describe how much of last summer's job the model does, not how much of this autumn's.

One week of compute is a snapshot, not a trend. The safety-allocation figure covers July 13 to 20. The report is explicit that safety research uses less compute by nature — the design work is time-intensive, not compute-intensive — so the 6 percent understates effort even before questions about labeling, and a single week cannot show direction.

Oversight has been running for months, not years. The monitors have not had time to encounter every pattern of behavior, and the report expects the methods to change significantly.

None of that undercuts the direction of the R&D number. It does mean the precise figures are best read as the first data points on a chart Anthropic has committed to keep updating, and it means a lab reporting on itself will need outside verification before the numbers should settle anything.

What should a contractor do with a report about a frontier lab?

Three things, none of which require understanding automation levels.

Be the answer before the engines get better at asking. Every one of these systems researches a business before recommending it — its website, its Google profile, the directories and license registries that vouch for it. We tracked 700 AI visibility checks in one Texas market and found the same company scoring 92 percent on one engine and zero on another, in the same window, on identical work. The difference was whether each engine's pipeline could see and verify the business. That is fixable, and it is worth more every month the engines improve.

Start from a problem you already have. The MIT finding that separated AI deployments that scaled from ones that got shelved was a pre-existing, documented problem the tool was pointed at. That holds whether the tool is a proposal drafter or an AI receptionist. A faster clock in the lab does not change the order of operations in your office.

Keep a person on the edge cases. The most useful detail in Anthropic's oversight numbers is not the 0.002 percent block rate. It is that about fifty transcripts a week still reach a human. Even the company that builds these systems runs them with someone on the exceptions. So should you.

The pace of AI development is now measurable, and the first measurement says a quarter. It will be higher next time. The businesses that will be fine are not the ones that predicted the number. They are the ones that were findable when the engines came looking.

The engines are getting better at asking

Find out whether they can find you.

The AI Visibility Audit checks whether ChatGPT, Gemini, Perplexity, and Google's AI name your business when a customer asks who to call — and what to fix if they don't.

Get My AI Visibility Audit — $27

Sources

  • Favaro, M., and Wright, P. "Measurements for understanding the pace of AI development inside frontier labs." Anthropic Institute, September 2026. Report. All figures for the automation index, agent oversight, and compute allocation are Anthropic's self-reported measurements for July and August 2026.
  • Epoch AI. The automation-level scale (AL0–AL5) used by the index. Toward an O*NET for AI R&D.
  • Armstrong, B., Shah, J., et al. Humans in the Loop. MIT Industrial Performance Center, April 2026. Our contractor translation is at MIT's 10 Levers for Generative AI, Applied to Contractors.
  • Direct quotations are from the Anthropic report and the MIT report. The two-clocks framing and every contractor application are ours.
Related posts
Insight

MIT's 10 Levers for Generative AI, Applied to Contractors

September 11, 2026
Insight

AI Visibility by Trade: Why Roofers, Plumbers, HVAC Companies, and Electricians Get Recommended Differently

August 28, 2026
Start here

See where you
actually stand.

The AI Visibility Audit runs your company through all six layers and delivers instant results. $27. Specific to your business, trade, and market.

Get the AI Visibility AuditExplore the Method
Get the audit