Problem
This contractor prices a job by setting a production rate for each type of work: labor hours per unit, say hours per foot of storm pipe. Set it too low and they underbid and lose money. Too high and they lose the job. They had years of job-cost history and no systematic way to check a new estimate against how that same work had actually gone.
Objective
One measurable question: does an LLM reading field notes beat plain statistics run on the same history?
A deterministic core did all the math: pull comparable jobs, compute the real spread of rates, classify the new estimate, and refuse to answer when history was thin. A thin LLM layer sat on top doing only what code cannot, reading unstructured field notes. That layer was the thing under test.
Baseline
How often the predicted range contained the real outcome
scroll to see the full diagram →
The misses split almost evenly: 19.5% of jobs landed above the predicted range, 18.5% below. I had expected a skew toward overruns. An even split means random execution variance, not budgets set too low. Widening the range until it covered 80% would have made it too wide to act on. The variance is not in the data, so no model was going to find it.
Why the LLM Idea Failed
The plan was to tag jobs with conditions that slow work down, such as groundwater, rock, weather and breakdowns, then price each one. That turns "rock is bad," which every estimator knows, into "rock costs an extra X% on this work," which they don't have.
Two ways the signal failed
scroll to see the full diagram →
Field notes describe a whole crew-day, not one line of the estimate, so a single rain note got attached to every type of work done that day. That leak produced the only statistically significant result in the project, and it vanished once notes were tied to specific work codes. The one surviving signal, groundwater, got weaker as more data was added, which is what noise does.
Location explains it. This is Florida civil work: sandy limestone, where the water table is the adversary and rock barely appears. "Rock slows digging" is true in general and not true here.
What the Data Did Support
So I dropped that idea and built what the history did support, which needed no LLM: a check that flags work codes whose budgeted rate is chronically optimistic. One sidewalk code ran about eight times its budgeted labor rate and came in over on 96% of jobs. Several grading and demolition codes ran 80 to 90%. That is a problem with the estimating template, not bad luck.
One trap on the way: my first evaluation reported zero overruns anywhere, because it included unfinished jobs, whose part-way hours make them look under budget. The fix was to stop inferring completion and read the real job-status field.
Decision
Same tool, two questions
scroll to see the full diagram →
59% of all estimate lines overrun. 67% of the ones the tool flagged do. A 1.1x lift does not justify a number an estimator would rely on, so the per-line risk score was recommended against in writing. The same flags used as a review filter catch 64% of the excess hours in half the lines, so the systemic check and the triage view shipped.
Where the LLM Earned Its Place
Exactly one job: reading notes on repeat-rework hotspots and classifying why work was redone. Damaged by another trade, damaged by weather, or our own defect. Other-trade damage can be billed back, so that reading has money attached.
The guardrail is worth copying. The model cites each note by index and never reproduces its text, and we look the real note up ourselves, so fabricated evidence is impossible by construction. The worst it can do is cite a note that doesn't support the call, which a reviewer sees. Every row ships marked "candidate pending verification." The evidence filter left six candidates worth roughly $65k in labor, a lead list for a project manager rather than a number to book. The client verifies each candidate before anything is billed, and what they confirm is the real figure.
That classifier also shipped broken, on a helper imported from a prototype folder that was never packaged: fine in development, missing from the installed tool. The tests stub out the LLM call for speed, so the one path they cannot cover is exactly where it broke. I packaged the helper and added a live smoke test to CI that exercises the real call once per run. It has not regressed since.
Outcome
Three analyses that could have disagreed did not. The calibration backtest, the note mining and the bid-time evaluation all landed in the same place: the systemic signal is real and worth surfacing, and per-job prediction is capped by variance that is not in the data.
Delivered:
- A memo naming which budgeted rates to fix, and which ideas to stop chasing
- A backcharge worklist a project manager can act on
- A data-quality punch list
- Two ship / don't-ship calls, each backed by a measurement