An inventory quantity is a running total of every receipt, sale, reservation, and correction that ever touched a SKU. Store only the total and you throw the history away, then spend years reconciling systems that disagree about it. Store the history as events and every total becomes a view you can rebuild.
The oversell that ruins a Monday rarely looks like a bug. The warehouse counted 40 units on Friday, the storefront sold 44 by Sunday night, and now someone in customer service is choosing which 4 customers get an apology instead of a parcel.
Every system involved did exactly what it was written to do. The root cause is a modeling decision so common it feels like physics: inventory lives as a number in a table, several systems update that number on different schedules, and the number is wrong in the gaps between updates.
This post makes the architectural case for the alternative. Inventory is a stream of facts, the table is a projection of that stream, and once you separate the two, oversell and undersell stop being mysteries and become measurable projection lag.
A quantity column is a cache with no invalidation story
Start with why the single number fails, because it fails for database reasons rather than integration reasons. The default isolation level in PostgreSQL is Read Committed, and its documentation spells out that 2 successive SELECT commands can see different data even inside a single transaction.
Now run the standard availability check. Two checkouts read the same row, both see 1 unit remaining, both pass the check, and both decrement. The table now reads -1, or worse, reads 0 while 2 confirmation emails go out.
The textbook fix is SELECT FOR UPDATE, which serializes every checkout for that SKU on a single row lock. That trades the oversell for lock contention on exactly the product everyone wants at exactly the moment they all want it.
The hottest SKU in the sale is, by definition, the row with the most write contention. A single quantity column turns your best-selling product into your database's worst bottleneck.
The Azure Architecture Center's event sourcing pattern names this precisely: under CRUD, updates are read-modify-write cycles with row-level locking, so concurrent writes to the same entity degrade performance and become a bottleneck under load. Appending an immutable event has no read-modify-write cycle to contend on.
A stream is a table in disguise
The theoretical grounding here is old and well documented. The Kafka Streams core concepts describe the stream-table duality: a stream can be viewed as a table and a table as a stream, because a stream is the changelog of a table, and replaying that changelog from the beginning reconstructs the table.
Read that sentence against your inventory schema and the conclusion is uncomfortable. Your quantity column was always derived state, a materialized fold over receipts, sales, and corrections. The difference is that you materialized it eagerly, kept no changelog, and now cannot rebuild it when 2 systems disagree.
Accounting internalized this centuries ago. The ledger is append-only, the balance is derived, and when a balance looks wrong you audit the entries rather than argue about the total.
The event sourcing pattern is the software formalization: store the full series of actions in an append-only store, treat that store as the system of record, and materialize domain objects from it. The Azure documentation lists the 2 wins that matter for inventory, write throughput under contention and a complete audit history.
What goes in the log
An inventory event is a fact about the physical or contractual world, written in past tense, immutable once appended. A correction is a new event, never an edit to an old one.
- StockReceived: units arrived at a location from a purchase order or transfer.
- StockCounted: a human or scanner observed an actual quantity, which may disagree with the projection.
- ReservationPlaced and ReservationReleased: units promised to an order, or unpromised when checkout abandons.
- StockPicked and StockShipped: units left the shelf, then left the building.
- ReturnRestocked: inspected units re-entered sellable stock.
- StockDamaged and StockQuarantined: units left sellable stock without leaving the building.
- TransferDispatched and TransferReceived: units in motion between locations, which is where single-number models quietly lose track.
Every event carries the SKU, the location, the quantity delta, the actor that observed it, and the time the fact occurred rather than the time it was recorded. That last distinction is what lets you rebuild an honest as-of view when a store's POS syncs 6 hours late.
Who may write which events is the same ownership question I argued in One Event Backbone Beats Point-to-Point Integration: the WMS and store systems own on-hand facts, the OMS owns reservations, and a consumer that wants a change sends a command to the owner. A stream does not change who may write. It changes what they write, from destructive updates to appendable facts.
Ordering matters at the entity level and only there. Partition the log by SKU, or SKU plus location, so every event for one item is read in the order it was written, which is the guarantee Kafka's design makes per partition, alongside retention that is configuration rather than a side effect of consumption.
Projections: every consumer gets its own answer
The second half of the argument is that there was never one inventory number to begin with. Different consumers ask different questions of the same facts, and forcing them all through one column is how the column ends up correct for nobody.
This is the CQRS pattern applied narrowly: segregate reads from writes into separate models so each can be optimized independently. The write side is the event log with its ownership rules. The read side is a set of projections, each built by folding the log into exactly the shape one consumer needs.
Commerce platforms already model inventory this way internally, whatever their APIs suggest. Shopify's inventory model tracks 8 named states, including incoming, available, committed, reserved, damaged, safety_stock and quality_control, defines on_hand as the sum of 6 of them, and describes order placement as a movement that decrements available while incrementing committed. The primitive is the movement between states, and the numbers are sums over movements.
| Consumer | Question it asks | Projection it reads | Staleness it tolerates |
|---|---|---|---|
| Storefront and checkout | Can we promise this unit right now? | Available-to-promise per sell channel | Sub-second, and the promise itself goes through the owner |
| OMS allocation | Which node should fulfill this order? | On-hand minus committed, per location | Seconds |
| Store operations | Can an associate pick this locally? | On-hand at 1 location | Seconds to a minute |
| Replenishment | What do we reorder and when? | Net position including incoming transfers | Nightly |
| Finance | What is inventory worth at close? | On-hand valued at cost, as of period end | Batch, but must be exactly reconstructable |
Each row is a separate consumer of the same log, with its own store, its own freshness contract, and its own rebuild path. When the ATP math turns out wrong, the fix is corrected fold logic plus a replay of the affected window, not a reconciliation script patching a shared column at 2am.
The question "what is the inventory number" has no answer. The answerable question is "what does this consumer need to believe about inventory, and how stale may that belief be."
Writers append facts, the log keeps them in order per SKU, and each consumer folds the same history into the view it needs.
Oversell and undersell are both projection failures
With the model in place, the 2 classic inventory failures stop being separate incidents and become the 2 directions of one measurement: how far a projection sits from the facts.
Oversell is a projection running ahead of reality. The storefront's ATP believed in units that a store had already sold offline, or that a transfer had already left, or that 2 channels each counted the same safety stock as theirs.
Undersell is the projection running behind reality, and it is the quieter and longer-lived failure. Returned units sit inspected and shelved but never restated as available, reservations from abandoned checkouts never release, and per-channel buffers stack until a third of real stock is invisible to every channel.
Oversell generates apology emails, so it gets found and fixed. Undersell generates nothing at all, which is why it survives for years. Nobody files a ticket about an item that showed sold out.
An event log gives both failures the same diagnosis path. Compare each projection against the log, and against StockCounted events specifically, then publish the divergence as a metric per projection. Oversell risk and undersell waste become numbers on a dashboard instead of quarterly surprises.
The strongly consistent core stays small
None of this makes the actual promise to a customer eventually consistent. Displaying availability tolerates staleness. Committing to it does not.
The commit is a command to the reservation owner, evaluated against the owner's own state, which appends ReservationPlaced or rejects. That single decision point is the serialization you actually need, scoped to 1 entity and 1 owner instead of imposed on every read in the estate.
Everything downstream of that decision, confirmation, allocation, the storefront counter ticking down, is projection and may lag by design. The engineering discipline is keeping the synchronous core to that 1 decision and refusing to let other consumers creep into it.
When the table is still the right call
The honest part of the argument is its boundary. The Azure documentation is blunt that event sourcing is a complex pattern with significant trade-offs, costly to migrate to or from, and that for most systems and most parts of a system, traditional data management is sufficient.
Inventory earns the pattern because it sits at the intersection of the 3 conditions that justify it: multiple legitimate writers, real write contention on hot entities, and audit questions that only history can answer. Most of your estate does not sit there.
- Single warehouse, single writer, modest concurrency: a table with row locking is correct and simpler. Take the lock and move on.
- Catalog attributes, content, settings: 1 owner, low contention, CRUD is fine.
- Multi-location, multi-channel, flash-sale traffic, franchise or marketplace writers: this is the stream's home ground.
Event-sourcing the whole platform because inventory needed it is how the pattern gets a bad name. Scope it to the aggregate that has the problem.
The peak calendar is the forcing function
There is a reason to have this argument in September rather than January. Everything that breaks a quantity column breaks hardest in the 4th quarter: flash-sale contention on hero SKUs, store and web channels selling from the same pool at full speed, and transfers in flight everywhere as allocation chases demand.
The same calendar shapes what is achievable before Black Friday. Steps 1 and 2 below are additive and shadow-only, safe to run through October, and the shadow diff is itself a peak-readiness artifact: it tells you exactly how wrong your availability numbers run under load, per SKU, before customers find out.
Cutover, moving real reads and the reservation path, waits for the January lull. Spending November arguing with a divergence metric is preparation. Spending it debugging a new promise path is self-harm.
Getting there without a rewrite
The migration is incremental, and each step pays for itself before the next one starts.
- Make existing writers emit events alongside their current updates, through an outbox so the event and the update commit together. No consumer changes yet.
- Build 1 projection, ATP, and run it in shadow. Diff it against the legacy quantity column daily and chase every divergence to a missing event type.
- Move 1 consumer, storefront availability display, onto the projection behind a flag. Watch the divergence metric, not the vibes.
- Route the reservation command path through the owner, and let the legacy column become just another projection that finance keeps until it trusts the new valuation view.
- Add projections per consumer and retire the direct database reads one at a time, which is the same one-at-a-time consumer migration that worked for the event backbone.
Reconciliation never fully retires. The nightly compare against physical counts and financial totals is how you find the event types you forgot existed, and a clean month of reconciliation is the actual definition of done.
FAQ
Do we need Kafka for this?
No. You need the properties: an append-only log, ordering per SKU, retention independent of consumption, and replay. An append-only table in PostgreSQL delivers all 4 at moderate volume, and a broker earns its operational cost when fan-out and throughput grow past that.
What happens when a physical count disagrees with the log?
The count is an event, not an embarrassment. Append StockCounted with the observed quantity, let projections apply the correction, and track the size of these corrections per location, because that divergence metric is your shrinkage and sync-lag detector.
Isn't replaying years of events too slow to be practical?
Rebuilds run from a snapshot plus the tail, not from the beginning of time. Snapshot each projection on a schedule, and keep a compacted changelog per SKU where only the latest state per key matters, so a full rebuild is bounded by snapshot age rather than log age.
How granular should events be, per unit or per adjustment?
Per state change with a quantity delta covers almost everything. Reach for per-unit events only where the physical world is serialized anyway, such as serial-numbered electronics or lot-tracked goods, because unit-level granularity everywhere multiplies volume without adding decisions.
Where does safety stock live in this model?
In the projection, never in the log. Buffers are policy applied when folding facts into a channel's ATP, so changing the buffer is a config change plus a rebuild, and the log keeps reporting what is physically true.
References
- Apache Kafka - Kafka Streams core concepts, duality of streams and tables
- Apache Kafka - introduction, ordering and retention guarantees
- Microsoft Azure Architecture Center - Event Sourcing pattern
- Microsoft Azure Architecture Center - CQRS pattern
- Shopify - inventory management apps and inventory states
- PostgreSQL - transaction isolation documentation
Related reading
More in ArchitectureInheriting a Codebase You Did Not Write: The First Two Weeks
Two weeks to observability, not to improvements. What to read first in an inherited codebase, what to baseline before touching anything, the 5 stabilization items to land, and an evidence-based test for rescue versus rewrite.
MACH-Aligned Without Being MACH-Certified: What Composable Actually Costs
MACH certification is a vendor membership programme, not a property of your architecture. Composable commerce is worth the money for the right retailer, and the difference is whether the team budgeted for the seams: optimistic concurrency, version conflicts, eventual consistency, and the operational surface you inherit.
Replatform, Modernize, or Rebuild: Telling the Three Apart
Replatform changes where the code runs, modernize changes what the code looks like, rebuild changes what the code believes about the business. Retail teams reach for the first when they need the second. The platform gets blamed because it has a vendor name and a renewal date, and the codebase has neither.