MIT's 10 Levers for Generative AI, Applied to Contractors

MIT studied more than twenty companies over three years to find what separates generative AI that makes work better from AI that only makes it faster. Their ten levers, translated for contractors — with the numbers, the caveats, and a ten-question audit that scores your shop.

Generative AI does not make a job better or worse. The employer's deployment choices do. That is the finding underneath a three-year MIT study of more than twenty companies, published as Humans in the Loop in April 2026 and distilled by MIT Sloan into ten levers this month. None of the companies studied was a contractor. Nearly all of the levers apply to one.

This is a translation. The researchers — Ben Armstrong of the MIT Industrial Performance Center, Kate Kellogg of MIT Sloan, and Julie Shah of MIT's Interactive Robotics Group — interviewed executives, managers, and employees across healthcare, finance, retail, and manufacturing from 2023 to 2025. What follows is what they found, what it looks like on a job site, and where the study's own limits mean you should hold the conclusions loosely. At the end is a ten-question audit that scores how a shop is actually using these tools.

What question was MIT actually asking?

Not whether AI works. Whether it makes jobs better, or just faster — and what separates the two.

The report opens with a number: in January 2026, roughly half of American workers reported using AI. What that has done to the quality of their jobs, the authors write, "remains largely unclear." So they went and looked, company by company, at what generative AI was deployed to do, how the roles of the people around it changed, and which experiments scaled versus which got quietly shelved.

Three patterns showed up everywhere. The companies were pointing AI at three kinds of problem:

The problemWhat it isWhere it shows up in a contracting business
BottleneckSkilled people buried in simple tasks that block higher-value workThe estimator answering scheduling texts; the owner rewriting the same proposal language for the fortieth time
CafeteriaA process that needs input from several experts, integrated into one outputA bid that needs the super's take on sequence, the PM's on subs, and the office's on terms before it goes out
Learning curveThe extra time a novice needs to do complex work in a new domainA second-year tech troubleshooting a unit he has never seen, from a manual he has never read

Every use case in the study maps to one of those three. It is worth deciding which one you are actually trying to solve before buying anything, because the tools and the risks differ for each.

Three problems, every companyMIT IPC · 20+ COMPANIES · 2023–2025 · TRANSLATED TO A CONTRACTING BUSINESS01 · BOTTLENECKSkilled people, simple tasksThe estimator answeringscheduling texts all day.AI job: speed the near-routine so the skill isspent where it counts.02 · CAFETERIAOne output, many expertsA bid that needs the super,the PM and the office first.AI job: predict what eachexpert would say, fromwhat they have said before.03 · LEARNING CURVENovices, complex workA second-year tech on aunit he has never seen.AI job: perform as ifexperienced — and, doneright, become experienced.Source: MIT IPC, Humans in the Loop (April 2026). Contractor examples are ours, not the report's.

What are the ten levers, and what do they look like on a job site?

The report lays them out as three principles that govern how you use the tools, then seven outcomes worth aiming at. Here is each one, with the contractor version underneath.

The ten leversTHREE PRINCIPLES THAT GOVERN USE · SEVEN OUTCOMES WORTH AIMING ATPRINCIPLES01Gather evidence before scaling02One size does not fit all03Learn when to trustOUTCOMES04Minimize drudgery05Promote learning06Preserve teamwork07Better interfaces08Keep investing in expertise09Maintain accountability10Create new workThe principles decide whether the outcomes are reachable. Skip 01 and every outcome becomes a guess.Source: Armstrong, Kellogg & Shah, MIT IPC (2026); MIT Sloan Ideas Made to Matter (Sept 3, 2026).

1. Gather evidence before scaling

The most effective deployments in the study started with a problem the company had already defined and already knew was worth solving. The less successful ones, in the Sloan summary's words, automated tasks based on time savings alone "without asking what it meant for quality."

On a job site: Before you put AI on proposals, write down what a bad proposal costs you now — the rework, the scope fights, the bids you lost to a vague line item. Then decide what evidence would prove the tool helped. Faster is not evidence. Fewer change orders is.

2. One size does not fit all

Workers doing the same job used the tools in vastly different ways, and the report treats that as a feature. Variation produces evidence about what works, for whom, and when — and it improves job quality, because people can use the tool how they want and skip it when they do not trust it.

On a job site: Let the estimator who loves the tool go deep and the one who distrusts it work his way. Then compare their outputs. You will learn more from the difference than from a mandate.

3. Learn when to trust

Willingness to trust is one of the strongest predictors of using automation well. The problem with generative AI is that it is a black box, so people cannot calibrate that trust on their own. The report's answer: if the tool cannot be made transparent, the employer has to build the practices that tell workers when to lean on it and when not to.

On a job site: Make a short list. AI drafts the follow-up email — send it. AI drafts the scope of work — a person reads every line. AI suggests a fix for a unit — the tech confirms against the manual before touching anything. Trust by task, written down, is what calibration looks like in a shop.

4. Minimize drudgery

AI proved most effective at removing routine work so people could spend time on problem-solving. The report adds a useful observation: the people closest to the routine tasks are the ones best placed to spot what should be automated.

On a job site: Ask the office manager and the lead tech what they do every day that a machine could do. Their list will be more accurate than yours.

5. Promote learning

This is the lever with the sharpest warning in the report. It cites a study in which undergraduates doing a research task with ChatGPT showed far less brain activity and far less ability to remember their work than peers using a search engine or their own memory. The implication for the learning-curve problem is direct: a tool that helps an inexperienced worker perform as if experienced may not be helping them become experienced.

On a job site: Your second-year tech who fixes the unit with AI's help and cannot explain what was wrong is a liability in eighteen months. Build the tool so it shows the reasoning, not just the answer, and make the explanation part of the job.

6. Preserve teamwork

AI can let one person finish what used to take three. The report is careful about the cost: the collaboration it removes was also where mentoring, collective learning, and trust between colleagues happened. Turning a cooperative task into a solo one can make the job less desirable and the team less capable.

On a job site: The pre-bid huddle where the super corrects the estimator's sequence is not wasted time. It is how the estimator learns to sequence. Keep it, even if the tool could skip it.

7. Design better interfaces

Most companies buy AI rather than build it, so the interface — how the tool is configured, what it shows, when it interrupts — is the lever they actually control. Good design builds what the report calls situational awareness while managing mental load.

On a job site: The difference between a tool that dumps a summary and one that flags the three things that changed since yesterday is the difference between a tool people use and one they turn off. You cannot change the model. You can change the setup.

8. Continue to invest in domain expertise

The report predicts short-term reductions in entry-level roles in the fields where AI is strongest, and in the same breath insists that "breakthroughs will still require experienced people to interpret and test what AI produces." Its manufacturing section describes what it calls a bipolar workforce: a growing share of young workers and a concentration of technical experts about to retire. AI that helps a novice technician troubleshoot from the manuals could raise the floor faster than experience alone. It does not replace the person who wrote the manual.

On a job site: That is your shop. The master plumber is sixty-one. The apprentice is twenty-three. AI is a bridge between them only if the master is still in the loop on anything that matters.

9. Maintain accountability

AI output can look credible while hiding an error. Making a named person accountable for the result, the report argues, "builds an incentive to actually learn the material and raises the cost of making an error." Its example is airport security, which adds random secondary screening precisely so the human does not switch off.

On a job site: Every AI-drafted document has an owner whose name is on it. When the tool gets the load calc wrong, someone is accountable for having sent it. That is not blame. It is the only thing that keeps the human reading.

10. Create new work

The report ends the list on the point most articles about AI skip: new technology does not only remove roles, it invents them. Redesigning jobs around both what the business needs and what people want to learn was associated with better adoption and more engaged careers.

On a job site: The office manager who now runs the AI proposal system, reviews its output, and trains new hires on it has a job that did not exist two years ago and is harder to leave. That is the outcome worth aiming for.

Ten questions · about four minutes

Score your shop against the ten levers

Check a box only if the answer is an unqualified yes — something you could point to today, not something you could probably arrange. Your score updates as you go. Your two lowest levers are where to start.

Running score0/ 10

Get your band, your lowest levers, and a copy of your answers by email.

Joins the 6 Signal list. Unsubscribe any time.

What made an AI deployment stick, and what got it shelved?

This is the most useful section of the report for anyone about to spend money, and it barely made the summary. After the experimentation phase, the working group asked why some applications scaled and others were quietly abandoned. Three features showed up in the ones that stuck.

Feature of what scaledWhat it meant in practiceThe contractor test
A pre-existing, well-documented problemThe winners addressed a challenge the organization already knew was costing it — doctors drowning in notes, nurses losing information at shift handoffCan you state the problem in one sentence, with a number, before you name a tool?
A constellation of technologies plus a humanOne company paired an LLM to read and categorize inquiries with an older rules-based bot to send consistent replies — the LLM for flexibility, the bot for reliability, because they would not let the LLM compose the answerIs the AI doing the part where variation is fine, and something predictable doing the part where it is not?
PersistenceA life-sciences firm's first attempt failed and was nearly shelved until an internal program supported iterating on itHave you budgeted for the second attempt, or does the first bad output kill the project?

The report's line on why the first feature matters is worth keeping: when frontline users "have a shared interest and incentive to solve a problem, that might overcome their resistance to technological change." Nobody resists a tool that fixes the thing they complain about.

Scale versus shelvesTHE THREE FEATURES THAT SEPARATED WHAT STUCK FROM WHAT WAS ABANDONED01 A PROBLEM YOU ALREADY HADDocumented, costed, and complained about before the tool showed up.02 AI PLUS SOMETHING RELIABLE, PLUS A PERSONThe LLM where variation is fine. A rules bot where it is not. A human on the edge cases.03 A BUDGET FOR THE SECOND TRYThe first version usually disappoints. The ones that scaled were allowed to iterate.Source: MIT IPC, Humans in the Loop, "Stage 3: Scale and Shelves."

One more finding from the same section deserves a contractor's attention. A real estate company piloted its AI tools with its highest-performing brokers first, on the reasonable theory that the best people would give the best feedback. It did not work. Experienced brokers with deep contact lists needed different things from the tool than newer brokers still building their knowledge — and the newer brokers turned out to have the most to gain. If you pilot AI with your best estimator, you may learn nothing about what it would do for your third-best one.

What do the numbers say, and which ones should you trust?

The report is unusually honest about forecasts, including the ones made by people who build these tools.

Forecasts swing by a factor of five. Usage is the number that is real.SHARE OF US JOBS OR WORKERS · EACH BAR MEASURES SOMETHING DIFFERENT — THAT IS THE POINT2013 forecast: jobs "vulnerable" by 203047%Same forecast, small methodology change9%Recent estimates: jobs with a substantial share of tasks exposed to generative AI50–70%Measured: American workers who reported using AI, January 2026≈50%Source: MIT IPC, Humans in the Loop (2026), citing Frey & Osborne (2013), subsequent revisions, and 2026 survey data.

A widely cited 2013 paper predicted that 47 percent of the US labor market was vulnerable to automation by 2030. A follow-up study changed the forecasting method slightly and got 9 percent. The report notes that current predictions about generative AI vary just as widely, from claims that it will "disrupt" half of white-collar jobs within five years to estimates that it will have only a modest effect on productivity. Its own language for what it found on the ground is more careful: AI is a "jagged frontier," far more useful on some tasks than others, and in many companies "a hammer in search of a nail."

Two other figures matter for anyone reading this from a truck.

First, when Anthropic released data on how its own model was being used, more than a third of usage was for computer and mathematical tasks — a category that covers about 3 percent of the workforce. The loudest AI success stories come from the smallest slice of the labor market. Software is where these tools are strongest; your business is not software.

Second, the report cites analysts projecting that generative AI "threatens white collar jobs, often occupied by skilled workers, and that middle-skill jobs in the skilled trades may become beneficiaries of this wave of technological change." That sentence is the trades' position in this whole story, stated by researchers with no reason to flatter you.

Where does the study stop applying to a contractor?

Three places, and the report names all of them itself.

The companies studied were "primarily large, established organizations that are among the leading firms in their industries." Not a ten-truck HVAC company. The dynamics of a Fortune 500 pilot — dedicated data scientists, internal sandboxes, hackathons — do not exist in a shop where the owner is also the estimator.

The people interviewed were mostly managers, "not the frontline workers or individual contributors whose jobs have been affected." A manager's account of how the crew uses a tool is not the crew's account.

And the manufacturers in the study were skeptical that generative AI could meet the quality standards of a production environment. Most of their AI came from outside vendors. Their software expertise to modify it was limited. That description fits a contractor better than anything else in the report — and it means the levers you can actually pull are the management ones, not the engineering ones.

Which is the good news. Nine of the ten levers are practices, not products. Gather evidence, calibrate trust, keep the expert in the loop, name an owner. None of it requires a data scientist. All of it requires deciding, before the tool arrives, what better means.

What should a contractor do with this?

Pick one problem from the three-problem table above. State what it costs you now, in a number. Choose the evidence that would prove a tool fixed it. Pilot with the person who has the most to gain, not the most to show off. Keep an experienced person on every output that matters, and put a name on every document the tool drafts. Budget for the second attempt.

Then, when the tool is working, ask the question the report ends on: what new work has this created, and who is going to grow into it?

If you scored the audit above, your two lowest levers are where to start. Everything else in this post can wait until those are running without you.

Before you deploy anything

Find out whether the AI engines can find you first.

Half of your future customers will ask ChatGPT, Gemini, or Google's AI who to call before they ever see your website. The AI Visibility Audit shows whether you are in the answer.

Get My AI Visibility Audit — $27

Sources

  • Armstrong, B., Shah, J., et al. Humans in the Loop: The Evolution of Work in Early Experiments With Generative AI. MIT Industrial Performance Center, April 2026. Full report (PDF).
  • Murray, S. "10 levers for shaping generative AI that truly improves worker performance." MIT Sloan Ideas Made to Matter, September 3, 2026. Article.
  • Researcher pages: Ben Armstrong, MIT IPC; Kate Kellogg, MIT Sloan; Julie Shah, MIT AeroAstro.
  • All contractor examples and the ten-question audit are ours, not the researchers'. Direct quotations are from the report and the Sloan article.
Related posts
Insight

Anthropic's Pace-of-AI Measurements: What 26% AI-Led R&D Means for Contractors

September 18, 2026
Insight

AI Visibility by Trade: Why Roofers, Plumbers, HVAC Companies, and Electricians Get Recommended Differently

August 28, 2026
Start here

See where you
actually stand.

The AI Visibility Audit runs your company through all six layers and delivers instant results. $27. Specific to your business, trade, and market.

Get the AI Visibility AuditExplore the Method
Get the audit