All projects Client work

Testing Whether an LLM Improved Bid Estimates

A contractor wanted an LLM to read field notes and sharpen their labor estimates. I built it, measured it against plain statistics, and it lost. I shipped the part that held up and recommended against the rest.

Client
Heavy-civil contractor, Florida
Question
Does the LLM beat plain statistics?
Answer
No
Stack
Python · SQL · LLM evaluation
Code
Client-owned. Public: workflow-agent

Problem

This contractor prices a job by setting a production rate for each type of work: labor hours per unit, say hours per foot of storm pipe. Set it too low and they underbid and lose money. Too high and they lose the job. They had years of job-cost history and no systematic way to check a new estimate against how that same work had actually gone.

Objective

One measurable question: does an LLM reading field notes beat plain statistics run on the same history?

A deterministic core did all the math: pull comparable jobs, compute the real spread of rates, classify the new estimate, and refuse to answer when history was thin. A thin LLM layer sat on top doing only what code cannot, reading unstructured field notes. That layer was the thing under test.

62%
real coverage of a nominal 80% range
reported as-is
96%
of jobs one work code ran over budget
~8x the budgeted rate
1.1x
lift from the per-line risk score
recommended against shipping
3
independent methods, same conclusion
the agreement is the result

Baseline

How often the predicted range contained the real outcome

At a nominal 80 percent prediction interval, leave-one-out coverage was 73 percent and forward-chaining coverage was 62 percent, both below target. Decomposing the forward-chaining run, 18.5 percent of actuals fell below the band and 19.5 percent above it: the under-coverage is near-symmetric, not skewed high. SHARE OF ACTUALS THAT FELL INSIDE THE INTERVAL Trained on all other jobs 73% Trained only on earlier jobs how it would run at bid time 62% nominal 80% WHERE THE MISSES WENT, FORWARD-CHAINING 18.5% below the band 62.0% inside 19.5% above the band near-symmetric, not skewed high

scroll to see the full diagram →

Forward chaining trains only on jobs that started earlier, so it reflects how the tool would have behaved at bid time. An 80% range covered 62% of actual outcomes, and the ranges were wide. A weak baseline is still the real baseline.

The misses split almost evenly: 19.5% of jobs landed above the predicted range, 18.5% below. I had expected a skew toward overruns. An even split means random execution variance, not budgets set too low. Widening the range until it covered 80% would have made it too wide to act on. The variance is not in the data, so no model was going to find it.

Those are two different problems with two different fixes, and most of this project depends on keeping them apart.

Why the LLM Idea Failed

The plan was to tag jobs with conditions that slow work down, such as groundwater, rock, weather and breakdowns, then price each one. That turns "rock is bad," which every estimator knows, into "rock costs an extra X% on this work," which they don't have.

Two ways the signal failed

Left panel: the weather driver was significant at p 0.01 using day-level notes, but collapsed to p 0.69 once notes were linked to specific cost codes, showing the original result was an attribution artifact. Right panel: the water effect fell from plus 0.53 to plus 0.21 within-code z as the sample grew to 55 codes and 380 jobs, never reaching significance. The one strong signal was a data error weather, p-value (lower = stronger) 0.01 looks real notes tied to the day 0.69 no effect notes tied to the work One rainy-day note was attached to every kind of work done that day. The surviving signal faded as data grew groundwater, effect size +0.53 promising a few work codes +0.21 not significant 55 work codes, 380 jobs An effect that shrinks as data grows was never really there.

scroll to see the full diagram →

Extraction was not the problem: the tags matched the notes on the samples I checked. The tagged conditions simply do not predict cost in this data. Both failures appeared only after the test was tightened.

Field notes describe a whole crew-day, not one line of the estimate, so a single rain note got attached to every type of work done that day. That leak produced the only statistically significant result in the project, and it vanished once notes were tied to specific work codes. The one surviving signal, groundwater, got weaker as more data was added, which is what noise does.

Location explains it. This is Florida civil work: sandy limestone, where the water table is the adversary and rock barely appears. "Rock slows digging" is true in general and not true here.

What the Data Did Support

So I dropped that idea and built what the history did support, which needed no LLM: a check that flags work codes whose budgeted rate is chronically optimistic. One sidewalk code ran about eight times its budgeted labor rate and came in over on 96% of jobs. Several grading and demolition codes ran 80 to 90%. That is a problem with the estimating template, not bad luck.

One trap on the way: my first evaluation reported zero overruns anywhere, because it included unfinished jobs, whose part-way hours make them look under budget. The fix was to stop inferring completion and read the real job-status field.

Decision

Same tool, two questions

As a per-line predictor, flagging lifted the overrun rate from 59 percent to 67 percent, a 1.1 times lift over the base rate. As a triage filter, about half the flagged lines accounted for 64 percent of the total excess hours. The triage use shipped; the per-line scorer did not. Does a flag predict that this line will overrun? 59% any planned line 67% a line the tool flagged 1.1x lift. Barely better than the base rate. Does a flag concentrate where the damage is? 48% share of lines flagged 64% share of excess hours caught Useful as triage. This is what shipped.

scroll to see the full diagram →

As a predictor it barely beats guessing, because the biggest blowups are execution variance. One job ran a code many times over budget on a planned rate that matched history perfectly. As a filter for where to look first, it concentrates most of the damage into about half the lines.

59% of all estimate lines overrun. 67% of the ones the tool flagged do. A 1.1x lift does not justify a number an estimator would rely on, so the per-line risk score was recommended against in writing. The same flags used as a review filter catch 64% of the excess hours in half the lines, so the systemic check and the triage view shipped.

Where the LLM Earned Its Place

Exactly one job: reading notes on repeat-rework hotspots and classifying why work was redone. Damaged by another trade, damaged by weather, or our own defect. Other-trade damage can be billed back, so that reading has money attached.

The guardrail is worth copying. The model cites each note by index and never reproduces its text, and we look the real note up ourselves, so fabricated evidence is impossible by construction. The worst it can do is cite a note that doesn't support the call, which a reviewer sees. Every row ships marked "candidate pending verification." The evidence filter left six candidates worth roughly $65k in labor, a lead list for a project manager rather than a number to book. The client verifies each candidate before anything is billed, and what they confirm is the real figure.

That classifier also shipped broken, on a helper imported from a prototype folder that was never packaged: fine in development, missing from the installed tool. The tests stub out the LLM call for speed, so the one path they cannot cover is exactly where it broke. I packaged the helper and added a live smoke test to CI that exercises the real call once per run. It has not regressed since.

Outcome

Three analyses that could have disagreed did not. The calibration backtest, the note mining and the bid-time evaluation all landed in the same place: the systemic signal is real and worth surfacing, and per-job prediction is capped by variance that is not in the data.

Delivered:

  • A memo naming which budgeted rates to fix, and which ideas to stop chasing
  • A backcharge worklist a project manager can act on
  • A data-quality punch list
  • Two ship / don't-ship calls, each backed by a measurement
The build took days and the measurement took weeks. The measurement is also what caught the leak in my own test harness, not just the weakness in the model.