Daniel Nguyen
← All work

Guardrailed production data sync for coding agents

An agent skill that lets a coding agent pull a fresh, scoped copy of production data into a local database to fix a production bug, with every risky step failing closed.

TTMI
Less manual setupagent fetches matching data itself
Area
Retail ERP
When
2026
My role
Self-initiated, built end to end
  • Bash
  • Python
  • PostgreSQL
  • GitHub Actions
  • AWS Batch
  • S3

How it fits together

Guardrailed production data sync for coding agentsThe agent only names a scope and a target; scripts check the target, get a fresh dump through CI, validate it and restore it in one transaction.Coding agentpersonOpt-in + target checkscheck or alarmMachine-wide lockcheck or alarmCI job (own identity)externalCloud scoped dumpworkerShort-lived linkdata storeUntrusted-dump checkscheck or alarmScoped restoreserviceLocal DB (loopback)data storescope + target aliasacquiredispatch fresh dumpstart jobpublishdownload, no keysformat + allowlist okone txn, FK check
The agent only names a scope and a target; scripts check the target, get a fresh dump through CI, validate it and restore it in one transaction.
  • Person
  • Check or alarm
  • External
  • Worker
  • Data store
  • Service

The problem

TTMI runs a retail ERP used by ~50 stores. Many production bugs only reproduce with real data: years of ledger history, orders stuck in unusual states, soft-deleted rows that still matter. Local databases were empty or stale, and getting a fresh copy meant asking someone with cloud access. Coding agents doing the fix work kept stalling on missing data.

Handing the agent a production dump is the obvious answer and also the obvious red flag. Production data on a laptop is a data-protection risk. An agent with database tools can restore into the wrong database. Cloud credentials on every machine widen the blast radius. A half-applied dump leaves a database that looks ready and is wrong.

Nobody assigned this. I started it on my own initiative because the setup step kept costing me time, and I set one goal: the safe path should be the easy path, and every dangerous path should fail closed.

What I did

The model decides what, the scripts decide how

The skill splits the work in two. The model decides whether the problem really is missing data, as opposed to a code, schema or permission bug, and which business module’s data it needs. Deterministic scripts do everything risky: the checks, the dispatch, the download, the validation and the restore. Safety never depends on the model reading prose carefully. The trade-off is flexibility. The agent picks from a fixed set of module scopes, and a custom scope needs explicit table-name rules taken from the task, never guessed.

No cloud credentials on the laptop

The laptop never connects to the production database or object storage with keys of its own. The helper starts a CI job under the developer’s own CI identity. The job runs the scoped dump in the cloud and hands back only a short-lived download link, which the helper never prints. Every run forces a fresh dump. A refresh takes minutes and depends on CI being up. In return, access follows existing repository permissions, and revoking someone needs no key rotation.

Make the wrong target hard to reach

Every run must name a target alias explicitly. The agent must ask the human if it is missing, never infer it. Each alias maps to one local connection. The helper refuses to continue unless the host is a loopback address and the database name in the config matches the one expected for that alias. After connecting, it checks the name a second time against what the server itself reports. There is no remote escape hatch. A machine-wide lock stops two agent sessions from restoring at the same time. Automatic restores also need an opt-in flag that the machine owner sets in a private config, checked before anything is dispatched.

Restore only the scope, all or nothing

The restore replaces data only in the tables present in the dump and leaves every other table alone. It runs in one transaction, checks the affected foreign keys inside it, and rolls back on any failure, so the previous local data stays intact. It never relaxes constraints to make a restore pass. The cost is that a scoped restore can leave cross-module references stale. When the foreign-key check fails, the agent stops and reports the exact failing stage instead of claiming the database is ready.

Treat the dump as untrusted input

Before anything touches the database, the helper checks the archive format and compression integrity, accepts table names only if they match a strict allowlist pattern, and rejects any plain-SQL statement outside the expected data-load shape. If the dump’s schema has drifted from the local one, column changes are derived from a temporary shadow database built from the dump itself. A malformed dump fails early with a precise error.

Result

A coding agent that hits missing data can now fetch a fresh, matching slice of production by itself, which removes a manual setup step. Wrong-target and broken-dump runs are designed to fail closed: refused before dispatch, or rolled back with the old data intact. The skill is packaged with an integrity manifest so another developer can install it with their own CI identity. I have no time-saved number, so the outcome stays qualitative.

What I’d do differently

The biggest gap is personal data: the client side does not mask anything, so customer and employee details land on the laptop as they are, and I would pseudonymise those fields in the export job. The config template also ships with auto-restore switched on, which makes that flag the real consent for agent-driven runs; it should ship off. Downloaded dumps have no checksum either, so I would have the CI job publish one and verify it before restoring.