The problem
Some bugs in our retail ERP only showed up with production-shaped data. A report would be wrong for one store, or a calculation would break on a combination of records nobody had in their local database. To fix these bugs, engineers needed realistic data on their own machines.
The obvious answer, giving engineers direct access to the production database, is a security and operational risk. Long-lived credentials end up on laptops, and a heavy dump running on the wrong host can slow down the system for users. We needed a way to hand out a scoped dataset on demand, without standing access, without running dump work next to user traffic, and without leaking secrets into job logs.
I co-built this with a teammate. I worked on the export logic and the workflow that triggers it, and my teammate owned the batch infrastructure.
What I did
Run the export as a batch job, away from the API
The export runs in its own container as a job on AWS Batch with a narrowly scoped role. Heavy dump work never touches the hosts that serve users, and every export is a recorded job with its own logs. The trade-off is more moving parts than a script on a server: a job definition, an image and a role to maintain. We accepted that for the isolation and the audit trail.
Scoped presets instead of dumping everything
Engineers start an export from a manual CI workflow and pick a preset for the area they need. Each preset exports only the slice of data required to reproduce problems in that area. Smaller exports are faster, cheaper to store and expose less data. The cost is that presets need updating when the data model changes, and sometimes an engineer needs a slice no preset covers yet.
Keep credentials out of the child process
The dump tool runs as a subprocess. Before starting it, the export removes the database connection secret from the environment the subprocess inherits, and passes only what the tool needs. Tests assert this behaviour, so it does not depend on someone remembering the rule. This made the export code a little more awkward, since it has to build the subprocess environment explicitly.
Validate before writing, verify after
The export checks the destination before uploading anything. After the upload it confirms the file exists in object storage, and if any step fails it deletes partial files. An engineer never gets a link to an empty or half-written snapshot. The extra checks add a few calls to every run, which is cheap next to a debugging session on broken data.
Short-lived presigned downloads
When the job succeeds, the workflow creates a time-limited presigned link to the file in S3 and posts it to the run summary. Engineers download the snapshot with that link and restore it locally. Nobody needs storage permissions of their own. The trade-off is that a link can be forwarded while it is valid, so its lifetime is capped and can be shortened per run.
Result
Engineers can reproduce production-only issues locally from a scoped snapshot. They trigger it themselves from CI, and they never need production database credentials. The secret-handling behaviour is covered by automated tests, which makes it a guarantee we check on every change.
What I’d do differently
I would define each preset together with the engineers who debug that area, since they know which records a bug needs better than the schema does. I would also put the export history, who ran which preset and when, on one easy-to-read page so access reviews stay quick.