The problem
The ERP at TTMI serves 50 stores, 500 employees and 4 brands. Some of its work takes a long time: inventory costing, posting large batches of documents, data migration jobs. All of it ran inside HTTP requests, where it held request threads until it finished.
Moving that work to a queue sounds simple. The hard part is what can go wrong in between:
- AWS SQS delivers at least once. The same message can arrive twice, sometimes at two workers at the same time.
- A worker can die halfway through a job. Someone has to notice and pick it up again.
- A request can enqueue a job and then roll back its own transaction. The job would then run against data that never existed.
For ERP work these are real risks. A costing job that runs twice, or a posting job that runs half a time, puts wrong numbers in the books.
What I did
I wrote the worker platform and deployed it as a separate worker service on ECS. I chose to build the claiming and lease logic myself instead of adopting a full task framework, because I wanted exact control over who owns a job at any moment. That also means owning poison-message and dead-letter handling.
The job table is the source of truth
Each job has a row in PostgreSQL. A worker claims it with a single conditional update: set it to running only if it is still pending. If the update touches one row, that worker won. If it touches zero, another worker got there first and this one drops the message. There is no select-then-write gap for a race to slip into.
The trade-off is that the queue becomes only a delivery mechanism. Status, attempts and ownership live in the database.
Leases, reclaim and a heartbeat
A claim comes with a lease that expires. If a worker dies, its lease runs out and another worker can reclaim the job, with the attempt counter increased.
While a job runs, a heartbeat thread does two things: it extends the SQS message’s visibility so the queue does not hand it to someone else, and it renews the lease in the database. If the heartbeat finds the lease was lost, the job stops instead of racing a second owner. At startup the worker checks that the heartbeat interval is shorter than both the lease and the visibility timeout, because a misconfiguration there would quietly break every guarantee above.
Idempotency keys as a unique constraint
Each active job has a fingerprint protected by a unique constraint. If two requests try to create the same job, the second insert fails on the constraint, and the code treats that as a normal race outcome instead of an error.
Enqueue only after commit
The message is published only after the transaction that created the job commits. If that transaction rolls back, there is no job and no message, so there are no orphans.
Retries, and separate images
Failed jobs retry with backoff up to a delivery cap. The API and the worker build from separate container images, so the heavy dependencies the workers need stay out of the API service.
Result
Costing and several heavy document workflows now run on the dedicated worker service instead of inside requests. Unit tests cover the failure cases directly: claim races, reclaim after a crash, heartbeat behaviour, lease loss and idempotent cleanup. Duplicate deliveries and dead workers do not lead to double execution.
The queue simulation on the homepage is a toy model of this design. It shows workers claiming jobs, leases expiring and duplicates being ignored, with made-up timings.
What I’d do differently
“Exactly once” in practice means at-least-once delivery plus idempotent, lease-guarded execution, and I would explain it that way to the team from the start. Because I built the queue layer myself, I would design the dead-letter path and its alerting on day one, instead of treating it as a follow-up.