Every legacy rewrite starts with the same sentence in the same meeting: "We know exactly what the old system does, so we will rebuild it on a modern stack and switch over in nine months." Eighteen months later the old system has shipped three urgent features, the new one is feature-complete against a spec that no longer exists, and nobody wants to be the person who flips the switch.
Martin Fowler has a sentence for this. Of the simple plan to replace a legacy system with one that behaves the same, he writes that we've seen it go down in flames most of the time. His reasons are mundane. Replacements take a long time. Users cannot wait for new features while you rebuild old ones. The existing behavior is hard to specify, and a lot of it is behavior nobody wants anyway.
The alternative has a botanical name. The strangler fig is a vine that starts in a crevice of a host tree, grows roots down to the ground, and slowly builds its own canopy. The host may die and leave the fig standing in its shape. Fowler saw these in the Queensland rain forests and used them as a metaphor for incremental modernization: build the new thing around the old thing, move behavior across in small pieces, and let the old thing wither once nothing depends on it.
My claim is that the pattern is simple to describe and easy to get wrong in three specific places: the routing layer, the shared data, and the proof that the new code behaves like the old. Most failed stranglings fail in one of those.
The shape: a facade that shifts traffic
The Azure Architecture Center version describes four phases. You introduce a facade (a proxy) between clients and the legacy system. You migrate functionality behind it in iterations, until the new system handles more requests than the old one. You decommission the legacy system once nothing depends on it. Then you remove the facade, or keep it as an adapter for clients that cannot change.
The important property is in the first sentence of the pattern: clients keep using the same interface and are not aware a migration is happening. That is what makes this a different animal from a rewrite. There is no cutover weekend. There is a routing table that gets one line longer every sprint.
In practice the facade is whatever already sits at your edge. An nginx config, an API gateway, a CDN rule, a load balancer with path-based routing. Here is the smallest useful version:
upstream legacy { server legacy.internal:8080; }
upstream billing { server billing-svc.internal:9000; }
server {
listen 443 ssl;
# Migrated: new service owns invoices
location /api/invoices/ {
proxy_pass http://billing;
}
# Everything else still goes to the monolith
location / {
proxy_pass http://legacy;
}
}
That is the entire pattern at day one. The default route points at the old system. Each migrated capability gets a more specific route pointing at the new one. Rolling back a bad release is deleting a location block and reloading.
If the strangler layer has to make a decision based on who the user is, not just the path, you want weighted or header-based routing so you can ramp. Percentages are better than switches, because the first 1% of traffic finds the bugs that the staging environment never will:
split_clients "${remote_addr}${http_user_agent}" $backend {
5% billing;
* legacy;
}
location /api/invoices/ {
proxy_pass http://$backend;
}
Hash on something stable, a user id or account id if you can get it, otherwise the same customer flips between systems on every request and you debug ghosts. The Azure guidance lists the facade's own risks: it must not become a single point of failure or a performance bottleneck, and you should treat it as transitional architecture, weighing its benefit against its temporary cost. A proxy you forgot about for four years is how a migration tool becomes a permanent dependency.
Where to cut: seams, not layers
The first real decision is what to move first. Fowler's advice is to find seams where the system can be split and to replace small components before large ones. A seam is a place where behavior can be intercepted without touching what is on either side.
Good first candidates tend to share traits. They sit at the edge of the system, so a URL or a message queue already marks the boundary. They have a narrow data footprint, so you are not dragging half the schema with you. They hurt, because there is a reason to move them: they are slow, they change weekly, or they block a launch. And they are not the core of the domain. Start with notifications, PDF generation, search, a reporting endpoint. Do not start with pricing.
The seam you cannot see in the code is often visible in the traffic. Pull a week of access logs, group by route, and sort by request volume and by error rate. The route that is both high volume and boring is your first target. The route that is low volume and weirdly shaped, with hand-written SQL and seven special cases for one enterprise customer, is your last.
Cutting by technical layer is the classic mistake. "Let's replace the data access layer first" produces a new layer that nothing calls yet and an old system that is now harder to change. Cut by capability, so that every step ends with something running in production that a user can touch.
The part everyone skips: the data
Routing is the easy half. The Azure pattern puts the hard half in its considerations: both systems need access to shared services and data stores at the same time. A new invoice service that calls the monolith's database directly is not decoupled, it is a second writer to someone else's schema.
There are three honest options, and all of them cost something.
The first is to let the new service read and write the legacy tables for a while. It is fast and it is a trap, because every schema change now needs two teams. Use it as a first step only, with a date on it.
The second is to give the new service its own database and replicate. The Azure guide describes this database variant explicitly: the new service starts on the monolith's domain tables, then you create an isolated domain database, migrate the tables with ETL, keep them synchronized with change data capture, validate consistency, and only then make the new database the system of record. Removal of the old tables comes last, because rolling back after removal is expensive.
The third is the anti-corruption layer. When the new code needs data in the legacy shape, or the legacy code needs to call the new service, a translation layer keeps legacy semantics from leaking into the new domain model. Without it the new service quietly inherits the old schema's names and its worst assumptions, and in two years you have a monolith with a different logo.
A concrete shape for the translation layer, in TypeScript:
// Legacy returns: { INV_NO, CUST_ID, AMT_CENTS, STAT_CD: "P" | "O" | "V" }
// New domain wants: Invoice { id, customerId, total: Money, status }
const STATUS = { P: "paid", O: "open", V: "void" } as const;
export function fromLegacy(row: LegacyInvoiceRow): Invoice {
return {
id: row.INV_NO,
customerId: row.CUST_ID,
total: { amount: row.AMT_CENTS, currency: "USD" },
status: STATUS[row.STAT_CD],
};
}
Twenty lines, one file, the only place in the new codebase that knows the word STAT_CD. When the legacy system finally dies, you delete the file and the old vocabulary goes with it.
Prove it before you route to it
The question that kills confidence in a strangler migration is never "does the new service work?" It is "does it give the same answer as the old one for the weird inputs we forgot we handle?" Tests written from the spec cannot answer that, because the spec is the thing nobody trusts.
The answer is to use production as the oracle. GitHub built a library called Scientist for this: you wrap the old behavior as the control and the new behavior as the candidate. The call always returns the control's value, so users see no change, but behind the scenes it runs both, randomizes the order, measures time, compares the results, records exceptions raised by the candidate, and hands everything to a publish hook you control. The README's example is a permissions check:
def allows?(user)
science "widget-permissions" do |experiment|
experiment.use { model.check_user(user).valid? } # old way
experiment.try { user.can?(:read, model) } # new way
end # returns the control value
end
You do not need the library to use the idea. In any language, the shape is the same:
async function getInvoice(id: string): Promise<Invoice> {
const control = await legacy.getInvoice(id); // source of truth
if (sampled(0.05)) {
queueMicrotask(async () => {
try {
const candidate = await billing.getInvoice(id);
if (!deepEqual(control, candidate)) {
metrics.increment("invoice.mismatch");
log.warn({ id, control, candidate }, "strangler mismatch");
}
} catch (err) {
metrics.increment("invoice.candidate_error");
}
});
}
return control; // callers never see the candidate
}
Run this in shadow mode for a week, then read the mismatch log. It is the most honest requirements document your system will ever produce. Some mismatches are bugs in the new code. Others are bugs in the old code that customers depend on, and now you have to decide, on purpose, whether to preserve them. Either way, the decision is made with evidence, not at 2 a.m. after a cutover.
A caution that is my own, not the library's: shadow-run reads, never writes. A candidate that writes will double-charge someone. For writes, use the ramp: 1%, 10%, 50%, 100%, with the legacy path still alive and the rollback still one config change away.
The yes, but
The strongest objection is that the strangler fig is slower and more expensive than a clean rewrite, because you pay for the facade, the dual-running period, the synchronization, and the translation layer, all of which get thrown away. That is true. Fowler says so himself: the approach does not make modernization easy, and the transitional architecture is a real cost. His argument is that reduced risk and earlier value outweigh it.
There are also cases where the pattern does not apply, and the Azure guide is blunt about them. If you cannot intercept requests to the back end, you have no place to put the facade. If you cannot modify the legacy source, you cannot disable migrated features or redirect internal calls. If the system is small and a wholesale replacement is simple, just replace it. And if the legacy system must be fully decommissioned quickly, a pattern designed for a long coexistence is the wrong tool.
Fowler adds a point that engineers like to skip because it is not a technical one. The organization that built the mess will build another one unless the way it works changes too. Conway's Law applies to the replacement. If the new services are sliced along the same team boundaries that produced the monolith's tangles, you will have recreated them over the network.
What to do on Monday
Open your edge config and write down the three routes with the highest traffic and the lowest business sensitivity. That is your candidate list. For the first one, put a proxy rule in front with the default pointing at the legacy system, so nothing changes, and ship that on its own. A strangler layer that is already in place and boring is worth more than a plan.
Then build the shadow comparison before you build the feature. If you cannot say how you will know the new service agrees with the old one, you are not ready to route a single request to it. And put an end date on the facade the day you create it, in the repo, as a ticket. Transitional architecture that nobody owns is just architecture.

