The previous release is ready to redeploy. The new release has already rewritten the customer records into a format the previous code cannot read. Pressing the rollback button now restores old code and creates a new outage. The team has a deployment reversal, but it does not have a complete recovery plan.
Rollback should be designed around the state and obligations a change creates. Returning to an earlier artifact, restoring data and compensating for an external action are different operations. A product should identify which recovery route applies, retain the evidence it needs and test the route before depending on it during a failure.
This chapter of The AI Software Factory follows a hypothetical reporting app through a faulty release and a mistaken customer update. Safe writes prevents mismatched commitments; durable workflows preserves progress. Here the question is what the operator can actually do after a change has produced an unwanted state.
Define the result recovery must restore
“Go back” can mean several things. The operator may want the service to answer requests again, the customer to see an earlier report or the database to recover missing records. Those outcomes can require different actions and may not occur at the same time.
Start with the customer obligation and tolerable loss. Which capability must return? Which accepted transactions must remain? Which data can be reconstructed? A blanket restore to yesterday can remove legitimate work performed after the backup.
For the hypothetical reporting app, restoring read access may be the immediate goal. Preserving reports accepted since the release remains another obligation. A recovery that makes the site available while silently deleting those reports is incomplete.
The recovery plan should therefore state its scope. A code reversal can restore an earlier behavior while leaving data reconciliation pending. The incident record should retain that distinction rather than use one green status to imply every consequence is resolved.
Identify the artifact before the incident
An operator needs to know which accepted version was running before the change and which version is now active. A moving branch name or vague release label can make that difficult.
Retain source revision, artifact identity and relevant configuration through the existing release system. The previous accepted artifact should remain available under the host’s normal retention rules. Rebuilding old source during an incident can introduce changed dependency or environment inputs.
Cloudflare’s current Workers rollback documentation describes redeploying a prior version and explicitly warns that connected resources are not rolled back with it. It also describes binding and lifecycle restrictions. Those narrow platform facts illustrate why deployment history is only one part of recovery.
Check the actual service and deployment model. A static site, a serverless worker and a database-backed application have different rollback mechanics. Use the repository’s established process instead of assuming one vendor command provides a general undo operation.
Stop creating new damage first
If the faulty release is writing incorrect records, continued traffic can expand the problem while the team investigates. A controlled pause or feature disablement may be the first useful action.
The stop should be scoped. Disable the affected import or write route while preserving safe read access where possible. A broad shutdown can create additional customer disruption, but leaving the harmful operation active can increase recovery cost. The choice depends on the observed failure and available isolation.
The stop control needs an authority path that survives the failure. If the only disablement interface depends on the broken app, the operator may be unable to use it. Retain a known operational route and test it with safe conditions.
Stopping new work does not resolve accepted work. Queued jobs, pending billing and customer promises remain. The incident state should identify them so they can be reconciled after the immediate harm is contained. App retirement rules address a permanent exit; a temporary rollback has its own narrower obligations.
Code and data need compatible transitions
A release may add a new field, change a representation or remove a structure. If the previous code cannot read the resulting records, code reversal is unsafe without a data plan.
A compatible migration often adds a new representation before retiring the old one. Readers and writers can transition deliberately while the old version remains usable. The exact technique depends on the database and application; it should be designed and tested against representative safe records.
For the hypothetical reporting app, a new cost field might initially coexist with the old field. A backfill can populate it, checks can compare results and a later release can remove the legacy representation. Combining all stages into one irreversible release makes recovery harder.
Compatibility also includes meaning. Keeping a field name while changing its units can break old code just as effectively as removing the field. Currency, timestamps and status labels deserve explicit migration rules. A schema that parses successfully can still describe the wrong customer state.
Backups become useful when restoration works
A backup is a retained copy or recovery mechanism. Its value depends on whether an authorized operator can restore the needed state within an acceptable time and interpret it correctly.
Test restoration with safe data in an isolated environment. Confirm that the app can read the restored records, that required relationships remain and that operation identities survive. A restored result table without its duplicate-prevention ledger can cause old events to create new effects.
Record the backup point and the likely gap between that point and the incident. Accepted work during the gap may need reconstruction from other authoritative records. A restore should not silently erase it and call the product repaired.
Backup access needs its own protection. A recovery copy can contain the same sensitive information as production. The privacy design chapter considers retention and access; recovery should preserve those boundaries rather than distribute broad copies during an emergency.
A customer-facing undo needs conflict rules
An app can offer undo for a recent change, but it should not overwrite a legitimate later action. The undo needs the original state, resulting state and a rule for what happens if the target changed again.
Suppose the app changes a report label from “draft” to “approved,” then another authorized user publishes it. Undoing the earlier approval cannot simply set the label back to draft without considering publication. The operation has crossed another boundary.
A conditional reversal can check whether the current state still matches the operation’s result. If it does, restore the prior state under valid authority. If it does not, stop and present a conflict or prepare a new corrective action.
The user should understand the scope. An undo button may restore a local field while leaving a sent message or external transaction intact. The interface should explain consequences beside the action, without overwhelming routine use with unrelated warnings.
External effects require compensation
Some consequences cannot be reversed in place. A customer received an email. A public report was seen. A remote service accepted an order. The application can take a new action that addresses the consequence, but it cannot make the original event never happen.
Compensation might mean sending a correction, issuing an authorized refund or withdrawing a published artifact with a retained explanation. Each is a new operation with its own permission, identity and verification. It should not be hidden inside a generic rollback command.
For the hypothetical reporting app, a report containing a wrong amount could be corrected with a new version and a clear notice to affected recipients. Deleting the old file alone may leave customers relying on a copy they already downloaded.
The compensation plan should identify affected operations from retained evidence. Guessing which customers saw the result can create more confusion. Where the audience cannot be fully identified, report that limit and choose a response appropriate to the uncertainty.
Recovery can move forward instead of backward
The previous version may contain another defect, lack the current data contract or no longer be available. A forward repair can be safer than forcing an old artifact onto new state.
The decision should compare the fastest reliable route to restoring the customer obligation. A narrowly scoped patch may correct the faulty parser while preserving compatible records. A code reversal may be preferable when the new release has a broad unknown defect and the prior state remains compatible.
Do not make “rollback” an ideological preference. Inspect the actual state, consequences and evidence. The incident owner should record why the chosen route addresses the failure and which risks remain.
A forward fix still needs review and verification appropriate to the urgency. Urgency can change the established release path, but it does not make an unexamined generated patch correct. Reuse the authorized emergency process and retain the evidence needed for subsequent review.
Preserve the incident’s evidence
Recovery can overwrite the state needed to understand what failed. Retain relevant logs, artifact identities, affected operation references and safe reproductions before making destructive repairs where feasible.
The evidence should be proportionate and protected. Copying all customer data into an unrestricted incident folder creates a new exposure. A sanitized fixture and scoped authoritative references can often preserve the failure without expanding access unnecessarily.
Distinguish observed facts from hypotheses. The outage began after the release, but a provider failure may also have occurred. The operator can use timing as an investigative lead while testing the cause. A confident incident summary written too early can direct the team toward the wrong repair.
After the service recovers, retain the accepted explanation and material uncertainty. A useful report tells future operators which condition triggered the failure and which evidence supports the repair. It should not fabricate a neat cause to close the incident.
Verification follows the recovery goal
Redeployment success is one observation. The product still needs to establish that the recovered customer path works and that the unwanted effect has stopped.
For the reporting app, verify that existing reports remain accessible, new safe inputs produce the intended result and the faulty write route no longer corrupts records. Inspect affected records or a representative reconciliation where the incident requires it. An application homepage returning successfully cannot establish data recovery.
Check related obligations too. Scheduled jobs may still run the faulty artifact, cached pages may show an old result or notifications may remain queued. The recovery scope should name which systems must agree before the incident can be considered resolved.
Behavior tests can supply controlled regression checks. Live verification examines the actual deployed state under authorized safe conditions. The report should distinguish those evidence types and avoid calling a local test a production result.
Rehearse recovery while the stakes are small
A written runbook can omit an unavailable credential, an obsolete command or an undocumented dependency. A rehearsal exposes those gaps before customers wait on them.
Use an isolated environment and safe data. Deploy a known change, introduce a controlled failure and execute the intended recovery route. Record how long each stage took, which decisions were unclear and whether the restored state met the independent success criteria.
Include a case where code reversal is incompatible with data. The expected result should be to detect the incompatibility and choose the defined alternative, not blindly complete the rollback. A rehearsal that always chooses the easiest reversible change gives little evidence about the hard boundary.
The result is local evidence for the tested configuration and operation. Revalidate after a new storage service, migration pattern or external effect changes the recovery assumptions. A past drill does not certify every future release.
Recovery authority should remain inspectable
An incident can motivate broad temporary access. The operator may need more capability than an ordinary coding task, but that capability should have a defined purpose and duration.
Record who can pause writes, deploy a recovery artifact, restore data and perform compensation. Separate technical access from legitimate authorization. An agent with a deployment key should not decide to refund customers because the product is failing.
Temporary authority should end when its purpose ends. Preserve the incident artifacts while narrowing credentials. This lets the team investigate and learn without leaving emergency powers active indefinitely.
The agent security guide develops least authority. Recovery design should implement it in a way that still permits timely action. A policy that forbids every repair without an unavailable owner can be as operationally fragile as a policy that lets anyone change everything.
Count recovery in the product economics
Recovery consumes support time, engineering effort and customer trust. A release process that produces fast changes but expensive incidents may have poor overall economics.
Track incident frequency, time to restore the named capability, affected customer work and reconciliation effort. Keep definitions stable. Time to redeploy is different from time to restore service, and both differ from time to resolve all data consequences.
A proposed pilot can compare recovery designs using controlled scenarios. Its timing results are not production availability statistics. Actual customer exposure requires observed operations over a defined period and careful reporting of what was measured.
The owner-hour metric makes that human burden visible. A small app that needs frequent emergency intervention can consume the operator’s capacity even if infrastructure costs are low.
Communicate the state customers need to act on
An incident message should tell affected customers which capability is unavailable, what work is preserved and which action they should take or avoid. It should reflect observed state. “Everything is fixed” is inappropriate while data reconciliation remains unresolved, even if the application is available again.
For the hypothetical reporting app, a useful message might distinguish restored access from reports still under review. Customers should not regenerate paid work simply because an old report is temporarily unavailable. The interface can preserve the operation and show its current recovery state.
Corrections should reach the people who may rely on the wrong result, within the product’s legitimate communication permissions. A new version placed quietly on the server may not help someone using a downloaded copy. The response needs to account for the actual distribution path.
Keep technical detail proportional to the customer decision. A customer usually needs the consequence and next step, while the engineering record retains artifact identities and migration details. Both descriptions should agree on the facts and uncertainty. Clear communication is part of recovery because it reduces the chance that people create new work or commitments while the system’s state remains unclear.
What Would We Do at Salars?
A proposed Forge app template would retain the previous accepted artifact, identify data compatibility and name the recovery owner. Each consequential feature would describe its reversal or compensation route beside its write design.
The initial merchant diagnostic pilot would rehearse restoring read access, retrying an interrupted safe job and correcting a published test report. It would include a deliberately incompatible data case and verify that the operator recognized the need for a different recovery path. These are proposed exercises, not completed Salars results.
Supplier Margin Guard would not gain automatic live write authority merely because its diagnostic could be redeployed. Price changes, notifications and paid operations would each need a scoped recovery model. The template could retain common mechanics while the app supplied its actual customer obligations.
A recovery feature is useful when the operator can state what it restores, perform it under valid authority and verify the resulting customer state. The previous version is one resource in that process; the obligation is the thing the process must recover.
Sources
- Cloudflare: Workers rollbacks, read October 7, 2026, for version redeployment, connected-resource limits and binding restrictions.
- Salars AI library, for related write, privacy and verification guides. Application designs, incident examples and rehearsal results in this chapter are hypothetical or proposed; no real outage or successful recovery is claimed.
Loading comments…