Written from operating and development evidence. Adapt the principles to the risk, authority and scale of your own system.

Back up the right state

Application files, database, environment, private data and relevant service configuration may all be needed. A backup should have a checksum and a tested restore path.

Define failure before launch

Health checks, important pages, forms, authentication, payment boundaries and service processes need explicit pass criteria. Without them, teams debate whether a release is broken while users wait.

Make rollback atomic

Versioned releases and controlled switching reduce partial deployments. Rollback should restore source, built output and private-state expectations together.

Capture evidence

Record version, hashes, route results, service health and invariant checks. Evidence helps diagnose subtle failures and gives the next release a trustworthy baseline.

Apply this to a real system.

Bring the context and we will identify the useful first decision.

Discuss the related system