The problem
The ERP at TTMI covers accounting and inventory for 50 stores, 500 employees and 4 brands. Every stock document (a receipt, a transfer, a sale) changes what an item is worth. The system values stock at weighted average cost per warehouse, so a change to one document can shift the average for every later movement in that warehouse, and those numbers flow into accounting entries.
Two things were wrong.
First, recalculation ran inside the HTTP request. It walked a long history of stock documents and wrote results back row by row. A user who saved a document could hold a request thread for a long time, and the code was hard to reason about once the data grew.
Second, several people edit documents that touch the same financial rows at the same moment. Without clear transaction and lock rules, that leads to lost updates (one save silently overwrites another) and deadlocks (two transactions each wait for a row the other holds).
What I did
I worked with the accounting team, who owned the business rules and gave me reference figures to check against. I implemented the costing engine and the shared locking conventions used by the inventory and accounting write paths.
Set-based write-back through a staging table
The engine computes new prices, bulk-loads them into an indexed temporary table and applies them to documents, movements and accounting entries with one set-based update. This replaced a loop that updated each row on its own.
The trade-off is more hand-written SQL and less ORM portability. In return, the database does one large join instead of a long series of small round trips.
Parallel lanes split by warehouse
Average cost in one warehouse does not depend on another warehouse, so I split the work into lanes by warehouse and run them concurrently. The rule only works if lanes never write the same rows. If that guarantee breaks, parallelism brings lock contention straight back, so the partitioning rule is something I check before changing anything in it.
Move costing to a background job
Recalculation now runs on the async job platform I built (see the async jobs case study). Saving a document commits the change and enqueues a recalculation after commit.
The cost is a short window where figures are eventually consistent. Users need to see that a recalculation is queued or running, so job status is visible to them.
Lock in a fixed order, and only what changes
Every write path that touches financial rows follows the same pattern: one atomic transaction; lock the parent document first; then lock only the child rows that actually change, always sorted the same way. Fixed ordering removes the deadlock where two users take the same two rows in opposite order.
One database detail shaped this. PostgreSQL cannot place a row lock on the nullable side of an outer join, so “lock everything I read” queries failed. When joined relations are nullable, the code locks only the primary table. Pessimistic locking lowers concurrency on hot rows, and I accepted that because it makes correctness easy to argue in review.
One precision for money, and tests that pin behaviour
I unified decimal precision across the financial models in one schema change, instead of widening fields one at a time as bugs appeared. That was a broad change up front, and it ended rounding drift between modules.
Before each refactor I added characterization tests that pin the engine’s output to the accountants’ reference results. Any change that moves a number fails the suite.
Result
Costing no longer blocks API request threads; it runs as a background job with visible status. Its behaviour is locked by tests built on the accountants’ reference cases. The same locking conventions now protect inventory, accounting and quality-control write paths.
What I’d do differently
Most of the correctness work came down to where the transaction and lock boundaries sit, and I learned that by fixing problems one write path at a time. I would write the locking convention down as a short team document first and review new write paths against it. I would also measure recalculation time before and after, because today I can only describe the speed-up qualitatively.