Whole-Server PITR: Restore Is Only the Beginning
- date
- category
- Reliability Engineering
- reading
- 2 min / 449 words
Whole-server PITR sounds like a database operation.
In practice, it is a system operation.
The database is only the mechanism.
The goal is to return the application to a state that makes sense for users.
When it makes sense
We choose whole-server PITR when the problem affects shared state.
Examples:
a bad job changed many records
a migration damaged data
an operator deleted important tables
the application wrote wrong values for an hour
an attacker changed data and we need to return before that change
If the problem is only in code, release rollback may be enough.
If the problem affects one table or one database, full server PITR may be too large a hammer.
But if the whole production state is suspect, we should stop patching records live.
Then the question becomes:
which point do we return to?
Stop making the damage worse first
Before restoring, we need to limit writes.
That can mean:
enable maintenance mode
block write paths
stop workers
stop cron jobs
cut off integrations
take a snapshot of the current state
A snapshot of a broken state sounds odd, but it can be useful.
It may contain data we do not want to lose.
It may be evidence for analysis.
It may allow selected records to be moved later into the restored environment.
The worst case is restoring a database while old production still accepts writes and nobody knows which of those writes will be needed later.
Restore beside production, not blindly
The safe model is simple:
old production is frozen
a new server is restored beside it
the new state is validated
only then traffic is switched
Restoring in place can be tempting because it looks simpler.
But it removes the ability to compare.
If restore goes wrong, we have neither the old state nor the new one.
Restoring beside production gives time to check:
whether the database started
whether the application can connect
whether critical tables have the expected number of records
whether the bug we are escaping is gone
whether we did not go too far back
In PostgreSQL, details depend on the backup setup, but the mental model is the same:
base backup
WAL archive
target time
promote restored instance
validation
Switching traffic
After validation, the application must be switched to the restored database.
That can be a change to:
DNS
connection string
platform secret
managed database endpoint
Service or routing configuration
It is important to know where the source of truth really is.
If some workers still write to the old database, the system splits into two worlds.
Before the switch, we need a list of database clients:
API
workers
cron jobs
admin tools
integrations
read models
migration processes
After the switch, we need to check more than the web application.
We also need to check background processes.
They often write data long after the main traffic looks fine.
What about data after the restore point
PITR means consciously returning to the past.
If we restore the database to 10:42, everything after 10:42 disappears from the restored world.
That may be acceptable.
It may also require moving some data.
Examples:
orders created after the restore point
payments confirmed by the provider
support tickets submitted by customers
events published to other systems
There is no universal automation here.
We need to decide which data is the business truth and which data can be reconstructed from external systems.
Sometimes the best decision is:
we return to 10:42 and manually reconcile payments from the last hour
That may still be better than keeping a damaged state alive.
Minimal runbook
A whole-server PITR runbook should include:
who decides to start PITR
how we choose target time
how we freeze writes
how we snapshot the current state
where we restore the new instance
how we validate data
how we switch all clients
how we verify the old instance no longer accepts writes
how we document lost data
The most important sentence:
database restore is not the end of recovery
It is only the moment when we have a candidate for new production.
It becomes production only after validation and after all traffic has been switched.