All projects Client work

Insights from a COO's Inbox

The client asked where AI could help their operations. Nobody can answer that from memory, so I turned twelve months of their COO's mailbox into a counted table of recurring problems. Two of the thirty-two called for AI.

Client
Consumer electronics brand
Engagement
Paid discovery, three days
Access
One mailbox, read-only, fixed window
Stack
Python · DuckDB · Gmail API · Claude

Problem

A company that wants to use AI has to decide where to point it, and the people best placed to answer cannot. An operations lead knows what their week feels like. Which parts of it a current system could absorb is a different question, and asking directly gets you the tasks they find annoying rather than the tasks that cost the business.

The inbox is where the work lands and where it stalls, which makes it a better record than anyone's recollection. It is also the least pleasant thing a consultant can ask for, so the terms came first.

The Terms

One mailbox, read-only, a fixed twelve-month window, agreed with the client before anything ran. The credential could not send, modify, label or delete anything. The client was told plainly, in advance, that message text goes to a model provider. Raw mail is destroyed when the engagement closes and only the derived table survives.

Objective

Turn a mailbox into something countable: a table of work threads, each labeled with what went wrong and what the operator had to do about it, so a recommendation rests on a number and a list of threads rather than on an impression formed while reading.

1,290
threads of real work, read and labeled
from 4,288 in the window
32
recurring problems, each with its evidence
472 threads cited
2
of the 32 called for an AI assist
the 2nd and 6th biggest problems
62%
of threads needed neither knowledge nor authority
attention is the automatable part

How the Mailbox Gets Read

Six stages, each one resumable and each one writing to the same table:

  • Index. The Gmail API lists every thread in the window with its message metadata. No bodies yet.
  • Exclude. A sender distribution built from that metadata alone decides what is bulk and what is work.
  • Fetch and clean. Bodies are retrieved for survivors, then quoted chains and signatures are stripped so each message is what that person actually wrote rather than a copy of the thread below it.
  • Extract. Each cleaned thread goes to Claude through the Batch API against a fixed seventeen-field schema. Free text where the answer is a phrase, such as what the work was and what actually broke. Enums where it has to be countable: what the thread needed from the operator (knowledge, authority, attention or nothing), how it ended, what was at stake. One row per thread, matched back by thread identifier and never by position.
  • Derive the taxonomy. The free-text phrases from every thread are collapsed into categories, which is what turns many ways of describing the same work into a countable few.
  • Aggregate. Failure phrases cluster into recurring problems, ranked by how many of the operator's messages they consumed and how long they stayed open, each carrying its thread citations.

Every extraction also returns a noise verdict, and a thread the model calls a notification, a pitch or personal goes back onto the exclusion list with its body purged. Sender metadata could not catch those; reading the thread could.

Email Is Untrusted Input

Message bodies are written by people outside the client's control, so everything reaching a model is delimited and labelled untrusted, with instructions to treat it as data and to report anything that tries to address the model directly. Bodies are also scanned for injection-shaped patterns: text telling a reader to ignore earlier instructions, lines impersonating a system role, shell commands.

Nothing matching is removed. Stripping suspicious text would corrupt the corpus to defend against a hypothesis, and a thread containing an attempt is a finding worth seeing. The real exposure was never the batch extraction, where the output shape is constrained and the blast radius is one row. It is a person later reading an exported slice inside a tool that can run commands.

What survived each stage

Of a mailbox holding 34,259 threads, 4,288 fell in the twelve-month window and were indexed from metadata alone. Four rounds of exclusion left 1,290 threads in scope and purged the bodies of the other 2,998. 472 of those threads are cited as evidence in a finding. The work reduced to 32 recurring problems. 34,259 THREADS IN THE MAILBOX · 4,288 IN THE WINDOW 4,288 Indexed in the 12-month window metadata only, no message bodies 1,290 In scope after four rounds of exclusion 2,998 dropped, their bodies purged 472 Cited as evidence in a finding 473 citations, one thread counted twice 32recurring problems, each carrying the threads that evidence it

scroll to see the full diagram →

Counts read from the pipeline's own run table. Exclusion ran four times: senders and bulk categories first, then subject rules, then twice more as the extractor's noise verdicts came back. Every round after the fetch purged the bodies it had already retrieved, 3,684 messages in all.

Decisions

Four choices did most of the work:

  • The Gmail API rather than a Takeout export. Native thread identifiers, server-side date and sender filtering, and no export queue. The alternative depended on a header surviving an mbox export, which was the highest-risk unknown in the original plan.
  • Exclude before fetching, then again after reading. 460 threads were ruled out before a body was retrieved at all. A further 2,538 went once the extractor had read them, which is what took the corpus from 3,828 to 1,290. The cost of that design is real: a few hundred threads reached the model before anything knew they were out of scope.
  • Derive the taxonomy from the whole corpus, not a sample. A thirty-thread sample shows you the common cases, which are the ones the client can already describe. These come from the corpus twice over: 1,259 distinct process phrases collapsed into twenty kinds of work, and 482 distinct failure phrases collapsed into the thirty-two problems below.
  • The thread is the unit, not the message. A thread is a piece of work. Counting messages would count how much people write.

Spend model attention only on what SQL cannot compute. Chase counts and response times were cut from the extraction schema because the timestamps already hold them exactly and for free. A proposed seventeen additional fields came down to eleven. Every field a model fills is a field that can be wrong.

One gate before spending anything. Model choice here was a quality question, not a cost one: batched, the gap between the two tiers was about thirteen dollars. Two fields carry the deliverable, and both need a model to name a mechanism rather than restate a symptom. I ran a sample under each tier and read the output, checking whether the cheaper one still made the distinctions the schema was built around. It did, so the corpus went through it in batches.

What the Corpus Said

Thirty-two recurring problems, each carrying the threads that evidence it and a proposed intervention. The distribution of those interventions is the finding.

What the 32 problems actually called for

Of the 32 interventions proposed across the corpus, 10 are a system of record, 10 are monitoring or an alert, 4 are a configuration change, and 2 are an AI assist. Two each are an integration, an automation and a written procedure. Every one of the 32 keeps a human on the exceptions; none is fully autonomous. WHAT THE 32 PROBLEMS ACTUALLY CALLED FOR 10 A system of record 10 Monitoring or an alert 4 A configuration change 2 An AI assist 2 An integration 2 An automation 2 A written procedure All 32 keep a person on the exceptions. None of them runs unattended.

scroll to see the full diagram →

Categories recorded on each proposal. The engagement was commissioned to find where AI belonged. Twenty of the thirty-two answers were a place to record something or a check to watch it.

Two findings shaped the recommendation more than the list did.

The biggest time sinks are the least automatable. Platform bugs and broken syncs between systems account for 230 of the COO's messages, more than any other pair of problems. In about a third of those threads the COO supplied something no system had: context, history, or a decision only they could make. In another 40 percent all they supplied was attention, which is the part a monitor can take, and that is what ten of the thirty-two proposals are.

The work here is one-off diagnosis, spread thin. The two problems that do call for a model are among the biggest in the corpus. The other thirty need somewhere to write things down first.

The COO writes nothing at all in 359 of the 1,290 threads. These are not newsletters, which had already been excluded. Two thirds are internal and the rest are vendors and agencies writing in. Most of the internal volume concentrates on a handful of senders, which makes it a conversation about when to use cc rather than a system to build, and the cheapest change available in the entire assessment.

The Same Corpus, a Different Question

A second pass asked something else entirely: not what goes wrong, but what systems exist and how they connect. It reconstructed 38 platforms and 26 flows between them, with nothing supplied by the company. Every platform, its purpose and its failures were inferred from the mail and cited back to the threads that named them.

A platform estate nobody had written down

Thirty-eight platforms and twenty-six flows between them were reconstructed from the mailbox alone. The existing data pipeline reads four of the thirty-eight. By status, 21 are settled and current, 9 are mid-migration either arriving or leaving, 5 cannot be determined from the mail, and 3 are retired. 38 PLATFORMS, 26 FLOWS BETWEEN THEM, NONE OF IT DOCUMENTED 4 read by the data pipeline 34read by nothing THE SAME 38, BY WHAT THE MAIL SAYS ABOUT THEIR STATUS 21 Settled and current 9 Mid-migration 5 Unclear from the mail 3 Retired Nine platforms changing at once is the backdrop to most of the integration failures.

scroll to see the full diagram →

Counts from the reconstruction. Fulfilment, affiliates and support were all changing operator at the same time, which explains most of the integration failures in the problem list.

This is the part I would reuse on any engagement. A company of this size usually has no current architecture document, and the real estate lives in whoever has been there longest. Their mail already contains it. Two days of reading gives you the platform list, who owns each one, which are being replaced, and which the warehouse can actually see, which is the first thing anyone needs before proposing what to build next.

One person's mailbox has gaps that the business does not, so a platform going quiet is weak evidence that it was dropped. Status was decided on explicit migration language rather than the date of last mention, and where the mail could not settle it, five of the thirty-eight are marked unclear instead of guessed at.

Outcome

The deliverable was a reviewable table: every problem with its thread citations, a proposed intervention, what it needs first, what must stay human, and its blast radius if it goes wrong. Every row traces back to the threads that produced it, so the client could argue with any of them.

I recommended against the general-purpose assistant that had been the obvious pitch. It needed standing read access to five vendors, a permanent liability in exchange for speeding up work that was never the bottleneck. One problem did justify building: the cross-system product data that no single vendor could see across. That became the next engagement.