The problem
Our retail ERP serves ~50 stores, ~500 employees and 4 brands from one Python/Django codebase on PostgreSQL. The API ran on AWS App Runner, a fully managed container service. That was easy to start with, but it gave us little say over how a release goes out or how to get back when one goes wrong. Background jobs also needed their own long-running worker service, which fit poorly on a platform built around web requests.
So we wanted to move the API and the workers to ECS. The hard part was doing it while the system stayed live. Users kept working, database migrations kept shipping, and we needed a safe way back if the new platform misbehaved. For a period, the old and new platforms would both talk to the same database.
I built the deploy pipeline and the AWS setup for the move. The decisions below are the ones that shaped it.
What I did
Run migrations once, from the image being released
Each release builds a container image, pushes it to the registry, and then runs database migrations as a separate one-off task started from that exact image. Only after the task succeeds does the new version deploy. This means the schema change and the code that expects it always come from the same build, so they cannot drift apart between build and deploy. The trade-off is a slower pipeline with one more step that can fail, and a failed migration now blocks the release. I wanted that block.
Keep schema changes backward compatible during the overlap
While both platforms served traffic, two versions of the code ran against one database. We agreed to allow only additive changes in that window: new columns and tables were fine, while renames, drops and type changes waited until the old platform was gone. This slowed down schema work for a while. In return, either platform could take traffic at any moment, and the actual switch became a routine release.
Separate images for the API and the workers
I split the container build into stages that produce a lean API image and a heavier worker image. The workers need libraries for things like report generation that the API never uses. Keeping them apart made the API image smaller and faster to start, at the cost of maintaining two image definitions and making sure both get rebuilt on every release.
Blue/green releases with an error-rate alarm and automatic rollback
On App Runner a release replaced the running version in one go. On ECS I set up blue/green deployments: the new version starts next to the current one, receives a small share of traffic first, and has to stay healthy for a bake period before it takes over. An error-rate alarm watches the new version, and if it fires, traffic goes back to the old version without anyone pressing a button. The trade-off is slower releases and double capacity during each rollout. I also added a separate read-only check that inspects the platform configuration before a release without building, migrating or changing anything, so setup problems show up early.
Make failures visible on both platforms
A rollback alarm only helps if someone can find out why it fired. I integrated Sentry for error tracking and centralized API and worker logs in CloudWatch Logs, across App Runner and ECS. A filter drops expected client errors, such as not-found or validation errors, before events reach Sentry, so the error stream shows real failures. I kept performance tracing off to avoid per-request overhead and relied on logs for timing. The filter rules need upkeep as the code changes, but every real server error is still captured.
Result
The production API and background workers now run on ECS, and the old App Runner setup is retired. Every release goes through blue/green with a short test period, and a bad version rolls back automatically on its error rate. When something does go wrong, errors land in Sentry and the matching logs are in one place, so investigations start from evidence.
What I’d do differently
I would turn the backward-compatibility rule into an automated check on new migrations from day one, so it does not depend on reviewers remembering it. I would also document the database side of a rollback as clearly as the traffic side, because the alarm can move traffic back but cannot undo a schema change.