The refactor versus rewrite decision depends on what still works, what must change and how users will move through the transition. Refactoring can preserve useful behavior while improving its structure. A rewrite can be justified when the existing foundation cannot support a necessary requirement. Neither choice removes the need to understand the current system.
Start by identifying the failure and gathering evidence. “The code is old” describes age. “The approval workflow cannot represent the required roles” describes a constraint. “Every release fails” describes an outcome that could come from architecture, testing, deployment or missing ownership. Each points toward a different investigation.
Map what the business relies on today
List the workflows users depend on, including workarounds. A spreadsheet exported every Friday may be part of the real system even if no architecture diagram includes it. Interview the people who operate it, inspect incident records and follow a representative action from input to completion.
Separate visible features from rules embedded in data and integrations. A replacement might reproduce the screens but lose a calculation, a partner reconciliation process or a reporting deadline. Before estimating either path, identify which behavior must be preserved and which can be deliberately retired.
The field-service ERP project spans field mobile work, enterprise operations and a client portal. In a comparable replacement decision, changing one of those surfaces would require understanding how its records reach the others. The published case illustrates the scope of the problem; it is not a claim that this ERP later underwent a rewrite.
Classify the failure before choosing a remedy
Try to place each problem in a concrete category. A bug in a calculation calls for a different response from a data model that cannot represent the business. A build that no one can reproduce may need ownership and release work before any architecture change is safe to attempt.
- Localized defect: wrong behavior in a bounded, understood path.
- Structural limitation: component or data boundaries conflict with an essential requirement.
- Operational gap: deployment, monitoring, restore or access cannot be performed reliably.
- Knowledge gap: behavior and ownership are insufficiently documented to estimate a change.
- Unsupported dependency: a required component can no longer be maintained within acceptable constraints.
More than one category may apply. An operational gap can conceal a structural limitation, and a knowledge gap can make a small change expensive. Mark uncertain classifications as provisional. The diagnostic task is to reduce uncertainty enough to choose a bounded next step.
A repair example: the failing vision detector
In the EV foreign-object detection teardown, a human hand received a model score of 0.303 while a small bolt received 0.921. Reflections also produced false alarms. The tempting response was another training cycle, but the investigation found decision logic, exposure and normalization problems around the model.
The published remedy introduced a model-free path for large objects and a more restrictive combination of conditions for small objects. It reused the available signals and did not require a new dataset or hardware change for that deployable fix. The case supports investigating the surrounding decision system before assuming the model must be replaced.
It does not prove that every detector can be repaired in the same way, or that a deterministic path eliminates every possible miss. Field evaluation still has to cover the relevant objects, lighting and device conditions. The useful lesson is the order of inquiry: establish the failure mechanism, then choose the change.
Compare four practical paths
| Path | When it may fit | Cost or risk to inspect |
|---|---|---|
| Repair a bounded defect | The failure is reproducible and the surrounding workflow remains suitable | Regression coverage and any hidden dependency |
| Refactor existing components | The behavior is valuable but difficult to change or test | Preserving behavior while restructuring |
| Replace a component gradually | A clear boundary allows old and new paths to coexist | Routing, duplicate data and transitional operations |
| Rebuild the system | A necessary requirement cannot reasonably fit the existing foundation | Migration, overlooked behavior and cutover recovery |
| Investigate first | Access or evidence is too incomplete to compare options | A bounded investigation and the decision it will enable |
Gradual replacement is often described using the Strangler Fig approach. The existing system and its replacement coexist while responsibilities move. This can expose value earlier, but it adds transitional architecture and operating work. Include those costs instead of treating incremental migration as free.
Estimate the transition as well as the destination
Compare the paths over the same period and with the same required outcomes. Count investigation, implementation, testing, migration and support. A rewrite estimate that includes only new features is not comparable with a repair estimate that includes production recovery and handover.
Use a simple working model: total effort equals understanding the current system, making the change, proving the change, transitioning users and operating the result. Record the assumptions for each part. If no one knows how many historical records need migration, that number belongs in the unknowns rather than a confident fixed estimate.
Consider opportunity cost separately from engineering effort. A gradual replacement may leave the team supporting two paths. A rebuild may delay necessary improvements to the current product. A repair may buy time without resolving a future constraint. Those effects should inform the decision, but they should not be converted into invented revenue figures.
Use decision gates before committing
Choose a small exercise that can distinguish the options. That might be reproducing the incident, testing whether a new rule can fit the existing data model, or migrating a representative set of records in an isolated environment. Agree in advance which result supports each path.
For a proposed rewrite, demonstrate the hardest workflow rather than the easiest screen. For a proposed refactor, show that important behavior can be observed and tested. For gradual replacement, demonstrate how requests are routed and how the team will reconcile state across the boundary.
The decision record should include the chosen path, rejected alternatives, evidence, remaining uncertainties and an exit condition. Set a point at which the team reassesses the choice if a crucial assumption fails. This makes the decision accountable without pretending the first estimate was complete.
Make ownership part of the outcome
A technically improved system still needs a team that can run it. Preserve deployment instructions, integration contacts, configuration ownership and incident procedures. If a vendor is leaving, recover access and build knowledge before starting a transition that depends on them.
Handover matters in the published UAE telehealth case: the platform was built, taken through the applicable health-authority approval process and handed to the client's team. That project is evidence of delivery and handover, not an automatic template for a different country's approval obligations.
What to bring to a diagnostic week
Bring the failure evidence, repository and deployment access where available, a list of critical workflows and the people who understand them. Say whether the system is live, which deadlines are real and which business decisions the assessment needs to support.
The software project rescue service starts with a paid diagnostic week. The first output should be a defensible recovery direction and a scoped next step, rather than a commitment to rewrite before the evidence is available. Use the architecture review checklist to collect the initial inputs.