You split the monolith into services, gave each one its own database, and then discovered that the thing databases were quietly doing for you all along is now your problem. The saga pattern is the most common answer. It is also the most misunderstood, because the word "transaction" in "distributed transaction" is doing a lot of lying.
Here is the setup everyone runs into. You have an e-commerce store where customers have a credit limit. Placing an order has to create the order, reserve credit against the customer, hold inventory, and take payment. In a monolith those four writes sit inside one BEGIN TRANSACTION ... COMMIT, and the database guarantees they either all happen or none of them do. Split Orders, Customers, Inventory, and Payments into four services with four databases and that single guarantee is gone. There is no COMMIT that spans four Postgres instances owned by four teams.
The textbook fix used to be two-phase commit (2PC): a coordinator asks every participant to prepare, and if they all say yes, it tells them to commit. It works, and almost nobody uses it for microservices. 2PC holds locks across the network for the duration of the whole operation, so one slow participant stalls everyone. It needs every database and message broker to speak the same distributed-transaction protocol, which in practice they do not. And the coordinator is a single point of failure at the worst possible moment, the window between prepare and commit. So we reach for the saga instead. But a saga is not a smaller, cheaper 2PC. It is a different thing wearing the same coat.
What a saga actually is
A saga is a sequence of local transactions. Each step updates one service's database and then triggers the next step. If step four fails, the saga does not roll back the way a database would. It runs compensating transactions that semantically undo steps three, two, and one, in reverse order.
The idea is older than microservices by three decades. Hector Garcia-Molina and Kenneth Salem introduced it in a 1987 SIGMOD paper called "Sagas," and they were not thinking about services at all. They were thinking about long-lived transactions, the kind that hold database locks for minutes or hours and block everything shorter behind them. Their move was to break a long transaction T into a sequence T1, T2, ... Tn, where each Ti has a compensating transaction Ci that undoes its effects. The database then guarantees one of two outcomes: either T1 through Tn all complete, or some prefix runs and is followed by the compensations that unwind it. The locks are released after each small Ti instead of being held across the whole thing.
That last sentence is the whole story, and it is where the misunderstanding starts. Releasing the locks between steps is exactly what makes a saga cheap. It is also exactly what a saga takes away from you.
The part the tutorials skip: you lose isolation
ACID has four letters. A saga keeps three of them and quietly drops one. It is atomic in the "all or compensated" sense, it preserves consistency if you write your compensations correctly, and it is durable. What it does not give you is isolation, the I. Between step two and step three of a saga, the intermediate state is visible to the rest of the world. Other sagas, other queries, other users can see a half-finished operation and act on it.
Picture two orders from the same customer arriving at nearly the same time. Both check the credit limit, both see enough headroom because neither has committed the other's reservation yet, and both proceed. In a single ACID transaction the isolation level would have serialized them. In a saga, nothing did, and now the customer is over their limit. This is not a bug in your code. It is the property you gave up in exchange for not holding locks across the network.
Chris Richardson, who cataloged the saga for the microservices world, is blunt about this: the lack of isolation means concurrent sagas can produce anomalies, and the developer has to add "countermeasures," design techniques that put back just enough isolation to be safe. We will get to those. The point to internalize first is that a saga is ACD, not ACID, and the missing letter is your job now.
Two ways to coordinate: choreography and orchestration
There are two ways to string the steps together, and the choice shapes your whole system.
In choreography, there is no central brain. Each service does its local transaction and publishes an event; other services subscribe and react. The order saga looks like this:
OrderService: POST /orders -> save Order(PENDING) -> publish OrderCreated
CustomerService: on OrderCreated -> reserve credit -> publish CreditReserved
(or CreditLimitExceeded)
InventoryService: on CreditReserved -> hold stock -> publish StockReserved
PaymentService: on StockReserved -> charge card -> publish PaymentCompleted
OrderService: on PaymentCompleted -> Order.state = APPROVEDThe failure path is the same wiring in reverse. If PaymentService publishes PaymentFailed, InventoryService listens for it and releases the stock, CustomerService listens and releases the credit, and OrderService sets the order to REJECTED. Each service knows only about the events immediately around it.
Choreography is appealing for a small number of participants and it fits event-driven systems naturally. It gets ugly as the graph grows. With five or six services, the business process no longer lives in any single place. It is an emergent property of who subscribes to what, and the only way to answer "what happens when payment fails" is to read six codebases and hold the whole thing in your head. There is also a real risk of cyclic dependencies between the event handlers.
In orchestration, one component owns the workflow. An orchestrator sends a command to each participant and waits for the reply before deciding the next step:
class CreateOrderSaga:
steps = [
Step(action=reserve_credit, compensation=release_credit),
Step(action=reserve_inventory, compensation=release_inventory),
Step(action=process_payment, compensation=refund_payment),
]
def run(self, order):
completed = []
for step in self.steps:
try:
step.action(order)
completed.append(step)
except StepFailed:
for done in reversed(completed): # unwind, reverse order
done.compensation(order)
order.state = "REJECTED"
return
order.state = "APPROVED"The workflow is now one readable object. You can see the forward path, the compensations, and the order they run in. The cost is a component that has to know about every participant, which is a form of centralization, and if you write it as a naive in-memory loop like the sketch above, a crash halfway through leaves the saga stuck with no memory of where it was. Real orchestrators persist their state after every step, which is why people reach for engines like AWS Step Functions, Temporal, or Netflix's Conductor rather than hand-rolling the loop. Those engines keep a durable saga log, so a process that dies between "charge card" and "record result" resumes instead of vanishing.
A useful rule of thumb: choreography for two or three participants in an already event-driven system, orchestration the moment the workflow has branches, more than a handful of steps, or anyone who will ever ask "where is order 4821 stuck right now."
Compensations are new actions, not rollbacks
The most dangerous word in saga writing is "rollback," because it makes you picture the database quietly reverting rows. Compensations are nothing like that. A compensation is a brand new forward transaction that produces the opposite business effect. Garcia-Molina and Salem's own example is the cleanest one: a credit of x dollars is compensated by a debit of x dollars. The debit is a real transaction that runs later and is visible in the ledger. You did not erase history. You appended a correction.
This matters because some effects cannot be compensated at all. You can refund a payment, but if you already sent the "your order shipped" email you cannot unsend it. If you already handed a package to the courier you cannot un-ship it. The design response is to order your steps so that the un-undoable ones come last, after everything that might fail. Sagas are usually split into three kinds of steps for exactly this reason: compensatable steps that can be undone, a single pivot step that commits the saga (once it succeeds, the saga will run to completion), and retriable steps after the pivot that must eventually succeed and are never compensated. Send the email after the pivot, not before.
Two properties are non-negotiable for every action and compensation. They must be idempotent, because at-least-once messaging means the same command will sometimes arrive twice and running "refund payment" twice must not refund twice. And compensations must be retriable, because a compensation that itself fails cannot just give up and leave the system inconsistent. In practice this means every step carries a saga ID, and every service records which saga IDs it has already processed.
Putting isolation back: countermeasures
Because the saga dropped the I, you add it back selectively where anomalies actually hurt. The common countermeasures, from Richardson's catalog:
A semantic lock is an application-level flag. The order is created in a PENDING state, and anything reading orders knows PENDING means "not final, do not treat as real yet." The pivot step flips it to APPROVED. This is the single most used countermeasure, and you have probably written it without knowing it had a name. Commutative updates sidestep ordering entirely: if your operations commute, like credit reserve and release on a counter, the order they arrive in stops mattering. Reread value guards against lost updates by checking a record has not changed since you read it before you write, and re-running the step if it has, an optimistic-lock check. There is also the pessimistic view, reordering steps so the risky read happens after the risky write, and the version file, which records incoming operations so they can be reordered into the right sequence.
You do not apply all of these. You look at which concurrent interleavings actually cause a wrong answer for your business and add the cheapest countermeasure that closes that specific hole.
The atomicity gap nobody warns you about
There is one more trap, and it sits inside a single step. Each step must do two things: update its own database and publish the event or reply that moves the saga forward. If it writes to the database and then crashes before publishing, the saga stalls forever. If it publishes and then the database write rolls back, downstream services act on something that never happened. You cannot wrap "write row" and "publish to Kafka" in one atomic transaction, because they are two different systems, which is the very problem you were trying to escape.
The standard fix is the transactional outbox. Instead of publishing directly, the step writes the event as a row into an outbox table in the same local transaction as its business write. Now the two writes are atomic, they are in the same database. A separate process tails the outbox table, using change-data-capture or a simple poller, and publishes each row to the broker, marking it sent. The event goes out at least once, the business write and the intent to publish commit or fail together, and the atomicity gap closes. If you take one implementation detail away from this piece, make it this one, because it is the failure that will actually page you at 3am.
Where this runs in production
The saga is not an academic pattern. Order-payment-inventory flows on essentially every e-commerce platform are sagas, whether the teams call them that or not. AWS publishes the saga choreography and orchestration patterns in its Prescriptive Guidance and points teams at Step Functions to run the orchestrator with a durable execution history. Uber built Cadence, and its successor Temporal, largely to make long-running sagas survivable, so a workflow that spans minutes or days and crashes midway resumes from its last recorded step. Ride cancellations, travel bookings that reserve a flight then a hotel then a car, banking transfers between institutions, and telecom order provisioning are all sagas underneath. The common thread is a business operation that spans systems you do not own atomically and that has a sane, if sometimes awkward, way to undo each step after the fact.
When not to reach for it
If your operation fits inside one service and one database, use a plain local transaction and do not think about any of this. If you genuinely need strict isolation and cannot tolerate any window where intermediate state is visible, a saga is the wrong tool and you should question whether the operation should have been split across services at all. And if the steps have no meaningful compensation, if there is no debit for the credit, no refund for the charge, then the saga's core assumption does not hold and you need a different design, often a redesign of the service boundaries so the un-undoable work stays local.
The saga is a good pattern with an honest bill. It buys you availability and loose coupling across services, and it charges you the isolation you used to get for free, paid back in compensations, idempotency keys, semantic locks, and an outbox table. Go in knowing you are not getting a transaction. You are getting a workflow that is very good at cleaning up after itself, as long as you are the one who wrote the cleanup.

