Your Company Already Collects the Data. Then It Throws It Away.
The construction industry has a reputation for being data-poor. It is the opposite. A single mid-sized job generates receipts, delivery tickets, timecards, photos, texts, RFIs, rental invoices, inspection reports, and change directives daily — a paper trail most industries would envy.
Then almost all of it evaporates. FMI estimated that 95.5 percent of the data captured in engineering and construction goes unused. The follow-up study with Autodesk put a price on the residue that does get used badly: bad data may have cost the global industry $1.85 trillion in 2020, including roughly $88 billion in rework traceable to decisions made on inaccurate or incomplete information.
The instinctive response is "we need a system." Most contractors have bought several. The failure repeats because the diagnosis is wrong.
Why the data dies where it lands
Walk one receipt through a typical company. A foreman buys fittings at the supply house at 7 a.m. The receipt goes in the truck. If it surfaces at all, it surfaces at month-end, gets keyed into accounting against whichever job the bookkeeper guesses, and becomes a line in a general ledger nobody reads by job. The information existed at every step. What never existed was anyone whose job was to move it, structure it, and connect it to the estimate that job was bid against.
That missing role is the entire mystery of the 95.5 percent. Structuring field data was clerical work, the clerical hours were unaffordable, so the industry rationally let the data rot. Every "go digital" initiative that ignored the labor economics — most of them — produced software that field crews correctly experienced as unpaid data entry.
What changed
The labor economics changed. Reading a receipt, matching an invoice to a purchase order and a job, coding a transaction for review, transcribing a voice note into a structured daily log, comparing this week's hours to the estimate line — this is precisely the work current AI does well, cheaply, and without being asked twice. The clerical constraint that justified throwing away 95 percent of your operational record is gone.
What has not changed: someone still has to decide which questions the data should answer. Collection without questions produces dashboards; questions without collection produce guesses. The order matters — questions first.
The five-number test
If you could only run the company on five numbers, refreshed weekly, which five? A defensible set for most contractors:
- Estimated vs. actual hours, per active job. The earliest honest margin signal you can get — weeks before accounting sees it.
- Completed work not yet billed. Change orders performed, milestones passed, retainage due. This is your money, parked in other people's accounts.
- True cash position, net of what you are floating. Payment delays are structural in this industry; pretending otherwise is how solvent companies die.
- Lead response time. The revenue system's equivalent of a leak detector.
- Margin by job type. Not by job — by type. This is the number that should be steering what you bid next year.
Notice what the list excludes: vanity aggregates, industry benchmarks, anything you would not act on. Five numbers a week, believed and acted on, beat forty charts nobody trusts.
Where this goes
The four-system frame from last week says the leaks live in handoffs — and every handoff is a place where data currently dies. The next posts in this track follow the consequences into money specifically: what financial information arriving late actually costs, and how the flow from field activity to job cost to billing gets rebuilt with AI doing the collection and a human doing the deciding.
The uncomfortable truth of the 95.5 percent is also the good news. You do not have a data collection problem. You have twenty years of already-collected evidence about how your company actually performs — and for the first time, structuring it costs almost nothing.