Problem
A company that wants to use AI has to decide where to point it, and the people best placed to answer cannot. An operations lead knows what their week feels like. Which parts of it a current system could absorb is a different question, and asking directly gets you the tasks they find annoying rather than the tasks that cost the business.
The inbox is where the work lands and where it stalls, which makes it a better record than anyone's recollection. It is also the least pleasant thing a consultant can ask for, so the terms came first.
The Terms
One mailbox, read-only, a fixed twelve-month window, agreed with the client before anything ran. The credential could not send, modify, label or delete anything. The client was told plainly, in advance, that message text goes to a model provider. Raw mail is destroyed when the engagement closes and only the derived table survives.
Objective
Turn a mailbox into something countable: a table of work threads, each labeled with what went wrong and what the operator had to do about it, so a recommendation rests on a number and a list of threads rather than on an impression formed while reading.
How the Mailbox Gets Read
Six stages, each one resumable and each one writing to the same table:
- Index. The Gmail API lists every thread in the window with its message metadata. No bodies yet.
- Exclude. A sender distribution built from that metadata alone decides what is bulk and what is work.
- Fetch and clean. Bodies are retrieved for survivors, then quoted chains and signatures are stripped so each message is what that person actually wrote rather than a copy of the thread below it.
- Extract. Each cleaned thread goes to Claude through the Batch API against a fixed seventeen-field schema. Free text where the answer is a phrase, such as what the work was and what actually broke. Enums where it has to be countable: what the thread needed from the operator (knowledge, authority, attention or nothing), how it ended, what was at stake. One row per thread, matched back by thread identifier and never by position.
- Derive the taxonomy. The free-text phrases from every thread are collapsed into categories, which is what turns many ways of describing the same work into a countable few.
- Aggregate. Failure phrases cluster into recurring problems, ranked by how many of the operator's messages they consumed and how long they stayed open, each carrying its thread citations.
Every extraction also returns a noise verdict, and a thread the model calls a notification, a pitch or personal goes back onto the exclusion list with its body purged. Sender metadata could not catch those; reading the thread could.
Email Is Untrusted Input
Message bodies are written by people outside the client's control, so everything reaching a model is delimited and labelled untrusted, with instructions to treat it as data and to report anything that tries to address the model directly. Bodies are also scanned for injection-shaped patterns: text telling a reader to ignore earlier instructions, lines impersonating a system role, shell commands.
Nothing matching is removed. Stripping suspicious text would corrupt the corpus to defend against a hypothesis, and a thread containing an attempt is a finding worth seeing. The real exposure was never the batch extraction, where the output shape is constrained and the blast radius is one row. It is a person later reading an exported slice inside a tool that can run commands.
What survived each stage
scroll to see the full diagram →
Decisions
Four choices did most of the work:
- The Gmail API rather than a Takeout export. Native thread identifiers, server-side date and sender filtering, and no export queue. The alternative depended on a header surviving an mbox export, which was the highest-risk unknown in the original plan.
- Exclude before fetching, then again after reading. 460 threads were ruled out before a body was retrieved at all. A further 2,538 went once the extractor had read them, which is what took the corpus from 3,828 to 1,290. The cost of that design is real: a few hundred threads reached the model before anything knew they were out of scope.
- Derive the taxonomy from the whole corpus, not a sample. A thirty-thread sample shows you the common cases, which are the ones the client can already describe. These come from the corpus twice over: 1,259 distinct process phrases collapsed into twenty kinds of work, and 482 distinct failure phrases collapsed into the thirty-two problems below.
- The thread is the unit, not the message. A thread is a piece of work. Counting messages would count how much people write.
Spend model attention only on what SQL cannot compute. Chase counts and response times were cut from the extraction schema because the timestamps already hold them exactly and for free. A proposed seventeen additional fields came down to eleven. Every field a model fills is a field that can be wrong.
One gate before spending anything. Model choice here was a quality question, not a cost one: batched, the gap between the two tiers was about thirteen dollars. Two fields carry the deliverable, and both need a model to name a mechanism rather than restate a symptom. I ran a sample under each tier and read the output, checking whether the cheaper one still made the distinctions the schema was built around. It did, so the corpus went through it in batches.
What the Corpus Said
Thirty-two recurring problems, each carrying the threads that evidence it and a proposed intervention. The distribution of those interventions is the finding.
What the 32 problems actually called for
scroll to see the full diagram →
Two findings shaped the recommendation more than the list did.
The biggest time sinks are the least automatable. Platform bugs and broken syncs between systems account for 230 of the COO's messages, more than any other pair of problems. In about a third of those threads the COO supplied something no system had: context, history, or a decision only they could make. In another 40 percent all they supplied was attention, which is the part a monitor can take, and that is what ten of the thirty-two proposals are.
The COO writes nothing at all in 359 of the 1,290 threads. These are not newsletters, which had already been excluded. Two thirds are internal and the rest are vendors and agencies writing in. Most of the internal volume concentrates on a handful of senders, which makes it a conversation about when to use cc rather than a system to build, and the cheapest change available in the entire assessment.
The Same Corpus, a Different Question
A second pass asked something else entirely: not what goes wrong, but what systems exist and how they connect. It reconstructed 38 platforms and 26 flows between them, with nothing supplied by the company. Every platform, its purpose and its failures were inferred from the mail and cited back to the threads that named them.
A platform estate nobody had written down
scroll to see the full diagram →
This is the part I would reuse on any engagement. A company of this size usually has no current architecture document, and the real estate lives in whoever has been there longest. Their mail already contains it. Two days of reading gives you the platform list, who owns each one, which are being replaced, and which the warehouse can actually see, which is the first thing anyone needs before proposing what to build next.
One person's mailbox has gaps that the business does not, so a platform going quiet is weak evidence that it was dropped. Status was decided on explicit migration language rather than the date of last mention, and where the mail could not settle it, five of the thirty-eight are marked unclear instead of guessed at.
Outcome
The deliverable was a reviewable table: every problem with its thread citations, a proposed intervention, what it needs first, what must stay human, and its blast radius if it goes wrong. Every row traces back to the threads that produced it, so the client could argue with any of them.
I recommended against the general-purpose assistant that had been the obvious pitch. It needed standing read access to five vendors, a permanent liability in exchange for speeding up work that was never the bottleneck. One problem did justify building: the cross-system product data that no single vendor could see across. That became the next engagement.