Problem
This brand sells on its own storefront, on Amazon in four countries, and ships everything from a third-party warehouse. All four systems keep their own record of what a product is, and they are supposed to line up on one product code. Nothing makes them. There is no integrity check between a company's separate software vendors and no job that fails, so when two systems disagree, all four keep running perfectly, each on its own version of the truth.
Four systems, four sets of identifiers, one link nothing enforces
scroll to see the full diagram →
What It Costs
One product was left off the spreadsheet that loads items into the warehouse. It went live, sold, and built up stock. A month later it showed up in a report holding 537 units and real revenue, unknown to the warehouse supposedly holding it. It surfaced only when someone tried to create a shipment and could not find it, 34 days after the mistake.
Another product had nine hundred reviews and displayed two, because the reviews sat on a parent product ID while the listing carried a separate variant ID. Every vendor checked their own side, found it correct, and said so. Nobody was responsible for looking across the boundary. That thread ran fifty days and closed unresolved.
Objective
The engagement started open-ended: where could AI help here? I analyzed a year of the operations team's correspondence, with their consent, and found no repetitive pile worth automating. The work was one-off diagnosis, nearly every case unlike the last.
So I recommended against the general-purpose agent that had been the obvious pitch. It needed standing read access to five vendors, a permanent liability in exchange for speeding up work that was never the bottleneck. One problem did survive, showing up in about a fifth of operational threads: this one. Deterministic, single domain, and buildable on a pipeline that already pulled two of the four systems every night.
Decisions
A nightly job pulls all four systems, matches them to one spine, and reports disagreements in six categories: missing, conflicting identifier, conflicting status, orphan, broken variation, and a listing left live after the product was retired. It writes to nothing and holds no credential that can change a product anywhere. It reports the disagreement and a person decides which side is wrong.
Two choices did most of the work:
- Normalize identifiers before comparing them. Barcodes differ between systems by a leading zero far more often than they differ in substance. After normalizing, exactly 15 were genuinely different.
- Pull Amazon from the bulk listings report, not the catalog endpoint. That endpoint only returns products that have sold, which is the exact blind spot the check exists to cover.
Tuning the First Run
Run one produced a few hundred findings against 433 products, as expected. A check like this is worthless until someone who knows the business explains which oddities are intentional, and that calibration is the difference between a report someone reads and one they filter to a folder.
Three rule corrections, all from client review
scroll to see the full diagram →
Two more categories were suppressed in the logic rather than by hand. Amazon variation parents and multi-item bundles are absent from the warehouse by design and would have produced around eighty permanent false alarms. A mute list is something a person has to maintain forever, and nobody does.
Outcome
A few hundred findings on the first run became nineteen worth a human's attention: three rule corrections from client review, two categories suppressed in the logic, and no mute list to maintain. Two of the nineteen: a product the warehouse knows by a name that appears in no alias map, and two products selling on both the storefront and Amazon that exist in no product master at all.
It ships in non-failing mode on purpose. It emails, it does not break the pipeline, and it stays that way until the backlog is small enough that red means something new happened. A check that is red on day one gets muted within a week.
One caveat printed on the report itself: I reconstructed what it would have said the day that product went missing, and it names it the next morning. That reconstruction uses today's data rather than a stored snapshot of that day, so it is a plausibility argument, not a backtest.