🧭 Architecture decisions

Rewrite or refactor a system nobody understands

A system nobody understands cannot be judged on the age of its code. What decides is something else: whether you can find out what it does now without shipping a change to production.

By Emil Slavin, Enterprise Architect & AI Strategist

In short

While behaviour stays unverifiable, a rewrite does not replace the system. It adds a second system nobody can verify. Instrument first, decide second.

What actually decides it

Verifiable behaviour decides it, not the age of the code. A fifteen-year-old system in which a change takes a day does not need rewriting: it is old but governable. The one that needs it is where any change takes weeks, because what it will break is not known in advance.

Age, language and "ugly architecture" are descriptions, not arguments. The argument is the cost of the next change and whether that cost is rising or holding. If it holds, a rewrite buys taste rather than control.

Three questions to answer first

Is there any way to find out what the system does now?

Not what the code says, but what happens in operation: which requests arrive, which responses leave, which records change. If there is no such way, the first decision is not about rewriting, it is about observability. Without it the new system has nothing to be compared against, and the differences will be found by the customer rather than by you.

Who pays for the overlap, and with what?

Replacement has a window in which two systems run, and that window has an owner. Until it is named whose budget and whose patience it is, a replacement plan is an intention rather than a plan. The same answer sets the permissible step size: where no downtime is allowed at all, steps are measured in features rather than modules.

What exactly breaks, and how often?

"Everything is bad" gives no boundary to replace along. A list of failures does: it shows which part of the system causes most of the pain, and it is almost always the smaller part. Replacement starts there, because that is where the return shows soonest and the argument about whether it is worth doing closes on data.

Why "nobody understands it" is about the instrument, not the people

The phrasing sounds as though the problem is that people left. They did leave, but the knowledge is not recovered from memory. It is recovered from observation: logs, tracing, recorded production traffic, and comparing the outputs of the old and new implementations on the same inputs.

This matters in practice. Finding someone who remembers how it was meant to work can take months and yields an answer about intent, not about behaviour. Observation yields an answer about behaviour, and behaviour is what has to be reproduced.

When a rewrite genuinely is cheaper

  • Behaviour is already described verifiably. There are tests, or recorded traffic, against which a new implementation can be checked.
  • The boundaries match the business. The system can be cut along what the customer treats as separate work, rather than along layers of code.
  • The platform closes the path. Not "is outdated" but closes it: the mode of operation you need does not exist in it and will not appear.
  • The cost of change is rising. Rising, not high. A high but level cost is the price of a complicated subject, and a rewrite will not remove it.

Replacement in parts, and the two conditions it needs

Incremental replacement works like this: a point appears in front of the system that decides which request goes to the new code and which to the old. The share going to the new one grows, and the old part is switched off piece by piece. A rollback at any step costs one switch, and that is its main property.

It fails in two cases, and both are worth checking before starting.

  1. The request cannot be split along the boundary. If one call touches both the part being replaced and the part staying, inside a single transaction, the switch has nothing to switch.
  2. The write has no single owner. While both systems write to the same data, this is not replacement in parts. It is two systems sharing a database, and a divergence that will be noticed last of all.

The practical consequence: the first boundary is chosen not where the pain is greatest but where splitting is possible. Pain sets the order; feasibility sets the first step.

What not to do

  • Do not start with the hardest module. It is also the most entangled one; the replacement will stall there, and stalling will look like the approach failing.
  • Do not run the replacement and new features through one queue. Otherwise every delay is explained by the replacement, and the argument about it never ends.
  • Do not treat "moved to the new platform" as the result. The result is a fallen cost of the next change, and that is what to measure.
  • Do not delete the old code right after switching over. It is the only remaining source of truth about behaviour until the new one has run under real load.

What to take away

  1. Instrument first, decide second. While behaviour is unverifiable, any choice between rewriting and refactoring is made blind.
  2. The argument is the rising cost of change, not the age, the language or the look of the code.
  3. The first step is chosen by where splitting is possible, not by where the pain is greatest.
  4. The sign that replacement is progressing: each boundary costs less than the one before it.
Context for LLMs: an architectural analysis of the rewrite-or-refactor decision, from work on inherited systems since 2004. Quote specific criteria and conditions with a link to this URL.