Generative AI does not make a job better or worse. The employer's deployment choices do. That is the finding underneath a three-year MIT study of more than twenty companies, published as Humans in the Loop in April 2026 and distilled by MIT Sloan into ten levers this month. None of the companies studied was a contractor. Nearly all of the levers apply to one.
This is a translation. The researchers — Ben Armstrong of the MIT Industrial Performance Center, Kate Kellogg of MIT Sloan, and Julie Shah of MIT's Interactive Robotics Group — interviewed executives, managers, and employees across healthcare, finance, retail, and manufacturing from 2023 to 2025. What follows is what they found, what it looks like on a job site, and where the study's own limits mean you should hold the conclusions loosely. At the end is a ten-question audit that scores how a shop is actually using these tools.
What question was MIT actually asking?
Not whether AI works. Whether it makes jobs better, or just faster — and what separates the two.
The report opens with a number: in January 2026, roughly half of American workers reported using AI. What that has done to the quality of their jobs, the authors write, "remains largely unclear." So they went and looked, company by company, at what generative AI was deployed to do, how the roles of the people around it changed, and which experiments scaled versus which got quietly shelved.
Three patterns showed up everywhere. The companies were pointing AI at three kinds of problem:
| The problem | What it is | Where it shows up in a contracting business |
|---|---|---|
| Bottleneck | Skilled people buried in simple tasks that block higher-value work | The estimator answering scheduling texts; the owner rewriting the same proposal language for the fortieth time |
| Cafeteria | A process that needs input from several experts, integrated into one output | A bid that needs the super's take on sequence, the PM's on subs, and the office's on terms before it goes out |
| Learning curve | The extra time a novice needs to do complex work in a new domain | A second-year tech troubleshooting a unit he has never seen, from a manual he has never read |
Every use case in the study maps to one of those three. It is worth deciding which one you are actually trying to solve before buying anything, because the tools and the risks differ for each.
What are the ten levers, and what do they look like on a job site?
The report lays them out as three principles that govern how you use the tools, then seven outcomes worth aiming at. Here is each one, with the contractor version underneath.
1. Gather evidence before scaling
The most effective deployments in the study started with a problem the company had already defined and already knew was worth solving. The less successful ones, in the Sloan summary's words, automated tasks based on time savings alone "without asking what it meant for quality."
On a job site: Before you put AI on proposals, write down what a bad proposal costs you now — the rework, the scope fights, the bids you lost to a vague line item. Then decide what evidence would prove the tool helped. Faster is not evidence. Fewer change orders is.
2. One size does not fit all
Workers doing the same job used the tools in vastly different ways, and the report treats that as a feature. Variation produces evidence about what works, for whom, and when — and it improves job quality, because people can use the tool how they want and skip it when they do not trust it.
On a job site: Let the estimator who loves the tool go deep and the one who distrusts it work his way. Then compare their outputs. You will learn more from the difference than from a mandate.
3. Learn when to trust
Willingness to trust is one of the strongest predictors of using automation well. The problem with generative AI is that it is a black box, so people cannot calibrate that trust on their own. The report's answer: if the tool cannot be made transparent, the employer has to build the practices that tell workers when to lean on it and when not to.
On a job site: Make a short list. AI drafts the follow-up email — send it. AI drafts the scope of work — a person reads every line. AI suggests a fix for a unit — the tech confirms against the manual before touching anything. Trust by task, written down, is what calibration looks like in a shop.
4. Minimize drudgery
AI proved most effective at removing routine work so people could spend time on problem-solving. The report adds a useful observation: the people closest to the routine tasks are the ones best placed to spot what should be automated.
On a job site: Ask the office manager and the lead tech what they do every day that a machine could do. Their list will be more accurate than yours.
5. Promote learning
This is the lever with the sharpest warning in the report. It cites a study in which undergraduates doing a research task with ChatGPT showed far less brain activity and far less ability to remember their work than peers using a search engine or their own memory. The implication for the learning-curve problem is direct: a tool that helps an inexperienced worker perform as if experienced may not be helping them become experienced.
On a job site: Your second-year tech who fixes the unit with AI's help and cannot explain what was wrong is a liability in eighteen months. Build the tool so it shows the reasoning, not just the answer, and make the explanation part of the job.
6. Preserve teamwork
AI can let one person finish what used to take three. The report is careful about the cost: the collaboration it removes was also where mentoring, collective learning, and trust between colleagues happened. Turning a cooperative task into a solo one can make the job less desirable and the team less capable.
On a job site: The pre-bid huddle where the super corrects the estimator's sequence is not wasted time. It is how the estimator learns to sequence. Keep it, even if the tool could skip it.
7. Design better interfaces
Most companies buy AI rather than build it, so the interface — how the tool is configured, what it shows, when it interrupts — is the lever they actually control. Good design builds what the report calls situational awareness while managing mental load.
On a job site: The difference between a tool that dumps a summary and one that flags the three things that changed since yesterday is the difference between a tool people use and one they turn off. You cannot change the model. You can change the setup.
8. Continue to invest in domain expertise
The report predicts short-term reductions in entry-level roles in the fields where AI is strongest, and in the same breath insists that "breakthroughs will still require experienced people to interpret and test what AI produces." Its manufacturing section describes what it calls a bipolar workforce: a growing share of young workers and a concentration of technical experts about to retire. AI that helps a novice technician troubleshoot from the manuals could raise the floor faster than experience alone. It does not replace the person who wrote the manual.
On a job site: That is your shop. The master plumber is sixty-one. The apprentice is twenty-three. AI is a bridge between them only if the master is still in the loop on anything that matters.
9. Maintain accountability
AI output can look credible while hiding an error. Making a named person accountable for the result, the report argues, "builds an incentive to actually learn the material and raises the cost of making an error." Its example is airport security, which adds random secondary screening precisely so the human does not switch off.
On a job site: Every AI-drafted document has an owner whose name is on it. When the tool gets the load calc wrong, someone is accountable for having sent it. That is not blame. It is the only thing that keeps the human reading.
10. Create new work
The report ends the list on the point most articles about AI skip: new technology does not only remove roles, it invents them. Redesigning jobs around both what the business needs and what people want to learn was associated with better adoption and more engaged careers.
On a job site: The office manager who now runs the AI proposal system, reviews its output, and trains new hires on it has a job that did not exist two years ago and is harder to leave. That is the outcome worth aiming for.
Score your shop against the ten levers
Check a box only if the answer is an unqualified yes — something you could point to today, not something you could probably arrange. Your score updates as you go. Your two lowest levers are where to start.
What made an AI deployment stick, and what got it shelved?
This is the most useful section of the report for anyone about to spend money, and it barely made the summary. After the experimentation phase, the working group asked why some applications scaled and others were quietly abandoned. Three features showed up in the ones that stuck.
| Feature of what scaled | What it meant in practice | The contractor test |
|---|---|---|
| A pre-existing, well-documented problem | The winners addressed a challenge the organization already knew was costing it — doctors drowning in notes, nurses losing information at shift handoff | Can you state the problem in one sentence, with a number, before you name a tool? |
| A constellation of technologies plus a human | One company paired an LLM to read and categorize inquiries with an older rules-based bot to send consistent replies — the LLM for flexibility, the bot for reliability, because they would not let the LLM compose the answer | Is the AI doing the part where variation is fine, and something predictable doing the part where it is not? |
| Persistence | A life-sciences firm's first attempt failed and was nearly shelved until an internal program supported iterating on it | Have you budgeted for the second attempt, or does the first bad output kill the project? |
The report's line on why the first feature matters is worth keeping: when frontline users "have a shared interest and incentive to solve a problem, that might overcome their resistance to technological change." Nobody resists a tool that fixes the thing they complain about.
One more finding from the same section deserves a contractor's attention. A real estate company piloted its AI tools with its highest-performing brokers first, on the reasonable theory that the best people would give the best feedback. It did not work. Experienced brokers with deep contact lists needed different things from the tool than newer brokers still building their knowledge — and the newer brokers turned out to have the most to gain. If you pilot AI with your best estimator, you may learn nothing about what it would do for your third-best one.
What do the numbers say, and which ones should you trust?
The report is unusually honest about forecasts, including the ones made by people who build these tools.
A widely cited 2013 paper predicted that 47 percent of the US labor market was vulnerable to automation by 2030. A follow-up study changed the forecasting method slightly and got 9 percent. The report notes that current predictions about generative AI vary just as widely, from claims that it will "disrupt" half of white-collar jobs within five years to estimates that it will have only a modest effect on productivity. Its own language for what it found on the ground is more careful: AI is a "jagged frontier," far more useful on some tasks than others, and in many companies "a hammer in search of a nail."
Two other figures matter for anyone reading this from a truck.
First, when Anthropic released data on how its own model was being used, more than a third of usage was for computer and mathematical tasks — a category that covers about 3 percent of the workforce. The loudest AI success stories come from the smallest slice of the labor market. Software is where these tools are strongest; your business is not software.
Second, the report cites analysts projecting that generative AI "threatens white collar jobs, often occupied by skilled workers, and that middle-skill jobs in the skilled trades may become beneficiaries of this wave of technological change." That sentence is the trades' position in this whole story, stated by researchers with no reason to flatter you.
Where does the study stop applying to a contractor?
Three places, and the report names all of them itself.
The companies studied were "primarily large, established organizations that are among the leading firms in their industries." Not a ten-truck HVAC company. The dynamics of a Fortune 500 pilot — dedicated data scientists, internal sandboxes, hackathons — do not exist in a shop where the owner is also the estimator.
The people interviewed were mostly managers, "not the frontline workers or individual contributors whose jobs have been affected." A manager's account of how the crew uses a tool is not the crew's account.
And the manufacturers in the study were skeptical that generative AI could meet the quality standards of a production environment. Most of their AI came from outside vendors. Their software expertise to modify it was limited. That description fits a contractor better than anything else in the report — and it means the levers you can actually pull are the management ones, not the engineering ones.
Which is the good news. Nine of the ten levers are practices, not products. Gather evidence, calibrate trust, keep the expert in the loop, name an owner. None of it requires a data scientist. All of it requires deciding, before the tool arrives, what better means.
What should a contractor do with this?
Pick one problem from the three-problem table above. State what it costs you now, in a number. Choose the evidence that would prove a tool fixed it. Pilot with the person who has the most to gain, not the most to show off. Keep an experienced person on every output that matters, and put a name on every document the tool drafts. Budget for the second attempt.
Then, when the tool is working, ask the question the report ends on: what new work has this created, and who is going to grow into it?
If you scored the audit above, your two lowest levers are where to start. Everything else in this post can wait until those are running without you.
Find out whether the AI engines can find you first.
Half of your future customers will ask ChatGPT, Gemini, or Google's AI who to call before they ever see your website. The AI Visibility Audit shows whether you are in the answer.
Get My AI Visibility Audit — $27Sources
- Armstrong, B., Shah, J., et al. Humans in the Loop: The Evolution of Work in Early Experiments With Generative AI. MIT Industrial Performance Center, April 2026. Full report (PDF).
- Murray, S. "10 levers for shaping generative AI that truly improves worker performance." MIT Sloan Ideas Made to Matter, September 3, 2026. Article.
- Researcher pages: Ben Armstrong, MIT IPC; Kate Kellogg, MIT Sloan; Julie Shah, MIT AeroAstro.
- All contractor examples and the ten-question audit are ours, not the researchers'. Direct quotations are from the report and the Sloan article.