A rollback does not undo every consequence
Restoring a previous release can leave changed data, external effects, and interrupted work behind. A recovery plan starts with those consequences and identifies how the crew will resolve them.
A deployment introduces a defect that writes incorrect values to a shared record. The crew detects the problem and restores the previous application version. Requests begin succeeding again.
The deployment has been reversed. The incorrect values remain.
Rollback is valuable, but it addresses a particular part of recovery. A responsible plan considers what a change can affect, which effects can be reversed, and how the team will handle the rest. The difference becomes important whenever software does more than display information.
Start with the effects
Before a consequential change, trace what happens when it runs. It may write data, send messages, trigger downstream processing, change permissions, schedule work, or communicate with people. Some effects remain after the code that caused them has stopped.
Ask which effects are reversible in practice. A configuration can often be restored. A transformed record may require its previous value or enough information to reconstruct it. A message already acted upon may require a compensating operation. A notification already read cannot be unread.
This exercise should shape the release. If a failure would be difficult to repair at full exposure, use a smaller scope, a staged transition, an additional check, or a different design. A recovery plan is useful partly because writing it reveals when the bet is too large.
Detection belongs in the plan
A recovery path that starts with “if anything goes wrong” leaves the crew to discover what wrong looks like during the incident.
Name the signals that would trigger investigation or intervention. Include important failure modes that might look successful at the request level: duplicates, missing handoffs, incorrect permissions, inconsistent totals, or an accumulation of work that nobody can complete.
Identify who can see those signals and who can act. If only one developer knows where to look, the system remains dependent on that person’s availability. The necessary evidence and access should be usable by the people expected to operate the change.
The goal is a credible response under pressure, with enough context to avoid improvising the entire decision.
Separate containment from repair
The first action may need to prevent further harm. That could mean stopping new work, disabling a specific path, limiting access, or pausing a consumer. Restoring a release may be one option, provided the older code can still operate with the current data and dependencies.
Repair comes next. Determine which records or operations were affected, what their correct state should be, and how a correction will be verified. Preserve enough evidence to distinguish affected work from healthy work. An indiscriminate repair can create a second incident.
Some situations need reconciliation across systems. Agree which source establishes the intended state and how conflicting information will be handled. Include the people responsible for the business process when the software alone cannot determine the right answer.
Rehearse the uncertain step
Reading a recovery document can reveal missing instructions. Exercising the difficult part can reveal missing capability.
For a bounded change, the rehearsal might be restoring representative data in a safe environment, verifying that an older application can read the new format, or confirming that a stopped job can resume without repeating completed effects. Match the exercise to the risk; repeating a familiar deployment command may not test the uncertainty that matters.
Record prerequisites, access, and limitations. A backup is only one ingredient if the crew cannot identify the needed point in time or reconcile work performed after it. A procedure requiring unavailable permission is not yet an executable plan.
Know when recovery is complete
Successful requests are encouraging evidence, but they may not establish that interrupted work is resolved. Check the affected data, downstream obligations, remaining queues, and user-facing consequences. Explain any continuing limitation to the people who need to act around it.
Before your next risky change, finish three statements: we will detect failure through this evidence; we will contain it with this action; we will restore a dependable state through this process. If the third answer is only “roll back,” trace the effects once more. The ship may still be taking on water after the engine stops.