Problem
An LLM agent with database access is the most useful thing I run and the one I trust least. It reads a warehouse every morning and writes up whatever moved. It is also a probabilistic system holding a shell and a network, where the failure mode is not a crash you can see but a quiet data leak you learn about later.
Objective
I run a fleet of these on my own hardware against real client data, so the design target was not what an agent is permitted to do but what it is structurally incapable of doing.
Decision: Separate What It Can Reach From What It Does
Every agent is two files. A policy says what it can reach: which databases attach to the run, which tools are allowed. A role says what it should do: model, turn limit, timeout, task. A role cannot grant itself anything, only name a policy, so adding an agent is a behavior change rather than a security change. The default is deny: a policy naming nothing produces an agent with nothing. My briefing agents run that way and have never had database access.
What one agent run can touch
scroll to see the full diagram →
The easier option was to put agent runs on the shared network with everything else and rely on credentials to keep them in their lane. Credentials get edited. A network route that was never created cannot be edited by mistake.
Database access is a real Postgres role, not a convention. The auditor
login carries 21 SELECT grants and zero
INSERT, UPDATE or DELETE, enforced a
layer below anything the agent or I can reach while it runs.
The Weak Point, Named
These runs are unattended, on a schedule, at night. Nobody is there to approve a tool call, so permission prompting is off. That makes the tool allowlist the control between the agent and the shell rather than one of several, and my own spec says so in those words. I would rather write that down than find it during an incident.
The five layers behind it:
- No added capabilities on the container
- No Docker socket, so a run cannot start other containers
- Input mounted read-only
- Credentials resolved only from host environment variables, and rejected at parse time if a password is ever written in plaintext
- No session persistence, so nothing a run learns survives it
Exposure, Counted
Where 47 containers actually sit
scroll to see the full diagram →
The number that matters is how many things are actually listening, and it drifts from the diagram the moment you stop looking. When I counted for this write-up, my own inventory notes were stale by eleven containers.
Every Run Leaves Evidence
A run writes report.json to its own directory, and that path
is the index: a log shipper parses service, role and run ID straight out of
it into Loki, while four metrics per run go to Prometheus. There are 1,049
of these archived, unbroken since February, which means the platform can be
judged on its own record instead of on my description of it.
Every run since February, and what happened to it
scroll to see the full diagram →
Those 58 non-zero exits are two unrelated stories, and only grouping them separates the two. Half a month of them were a real client-side problem, which is what led me to split the drift check into its own stage. The next twelve nights were that split backfiring: the new stage was the first command in the pipeline to contain a quoted argument, and the runner was building its arguments by splitting on spaces, so the test stage collected nothing and exited before it checked anything. Both look identical to a monitor that only watches whether a stage finished.
Guards, Not Elegance
Twelve minutes after midnight
scroll to see the full diagram →
Not elegant. Something upstream does something inconvenient on a schedule, so you measure it and put a guard in front of it. Dokploy exposes no exclusion list for the prune that I could find, so the alternative was forking the orchestrator and owning that fork forever. 41 of 47 containers currently hold three weeks of continuous uptime, and that floor is the night the guards were still being written.
What It Runs
The main workload is a client ETL pipeline: eight stages that a small dependency engine resolves into tiers and runs in parallel where it can, pulling five external APIs through a staged warehouse into the client's reports. Only one of those stages is allowed to stop a delivery.
Eight stages, and which failures are allowed to stop a report
scroll to see the full diagram →
That shape came from a mistake. One data-health check was originally doing two jobs: catching bugs in my pipeline, and surfacing drift in the client's process. The client's backlog kept it permanently red, which trains everyone to ignore it. Splitting the client-side findings into their own stages made red mean we broke something again.
Alongside it, a security scanner sweeps images, secrets and host hardening on its own schedule, and two backup jobs dump their databases weekly, then restore them into throwaway containers and match row counts, so the backup is verified rather than assumed.