Recovery from the Ground Up: Rollback, PITR, and Returning After Failure
- date
- category
- Reliability Engineering
- also in
- CI/CD · Databases · Incident Management
- reading
- 2 min / 362 words
The simplest recovery plan starts with one question:
which point do we need to return to?
Not with a tool.
Not with a cloud provider.
Not with a disaster recovery diagram that looks good in documentation and that nobody has practiced.
First, we need to name the loss.
how much data can we lose
how long can the system be unavailable
are we rolling back code, data, traffic, or everything at once
Without that, recovery becomes chaotic.
Someone says "rollback", someone else means restoring a backup, someone else wants to switch traffic, and the database has already accepted new records in a format the old application version does not understand.
This series is about simple, practical recovery.
Not a full strategy for a bank, active regions in two clouds, and automatic failover for everything.
The goal is to understand the basic moves before the word "restore" starts to mean anything.
What this series covers
We start with terms.
PITR means restoring data to a specific point in time.
RPO says how much data we can lose.
RTO says how long we can take to return to service. If TRO appears somewhere, it usually means RTO.
SLO describes a technical promise the system should meet.
SLA describes a commitment to a customer or organization.
Then we move to more concrete cases.
First, rolling back a release:
code
migrations
data compatibility
feature flags
queues and events
Because application rollback is often simple only when the database has not gone too far.
Then PITR for the whole server.
Not as a magic "restore" button, but as a sequence of decisions:
restore
check
freeze writes
switch traffic
confirm state
Then PITR for a single database on the same server, a smaller case that can be easier to explain when one tenant has one database.
At the end, we cover what happens after the system is back.
Recovery does not end when dashboards turn green.
We still need to reconstruct the timeline, find holes in the runbook, improve monitoring, add a restore test, and decide whether the accepted RPO and RTO were actually conscious decisions.
Minimal model
This series stays with a simple model:
application
database
full backup
WAL or binlog
routing to the active endpoint
runbook
That is enough to show the important mechanisms without hiding them behind a platform.
Kubernetes, managed databases, VMs, or bare metal change the implementation.
They do not change the basic questions:
what exactly did we lose
what are we returning to
what must not be overwritten anymore
who switches traffic
how do we know the restored state is correct
If these questions are clear, recovery is still stressful, but it stops being improvisation.
That is the point of this series.