Konrad Kowalski (rootsher)Principal Platform & Reliability Architect100110110000100010000011010000110110010010010101

Whole-Server PITR: Restore Is Only the Beginning

date
category
Reliability Engineering
also in
Databases · CI/CD
reading
2 min / 449 words

Whole-server PITR sounds like a database operation.

In practice, it is a system operation.

The database is only the mechanism.

The goal is to return the application to a state that makes sense for users.

When it makes sense

We choose whole-server PITR when the problem affects shared state.

Examples:

text
a bad job changed many records
a migration damaged data
an operator deleted important tables
the application wrote wrong values for an hour
an attacker changed data and we need to return before that change

If the problem is only in code, release rollback may be enough.

If the problem affects one table or one database, full server PITR may be too large a hammer.

But if the whole production state is suspect, we should stop patching records live.

Then the question becomes:

text
which point do we return to?

Stop making the damage worse first

Before restoring, we need to limit writes.

That can mean:

text
enable maintenance mode
block write paths
stop workers
stop cron jobs
cut off integrations
take a snapshot of the current state

A snapshot of a broken state sounds odd, but it can be useful.

It may contain data we do not want to lose.

It may be evidence for analysis.

It may allow selected records to be moved later into the restored environment.

The worst case is restoring a database while old production still accepts writes and nobody knows which of those writes will be needed later.

Restore beside production, not blindly

The safe model is simple:

text
old production is frozen
a new server is restored beside it
the new state is validated
only then traffic is switched

Restoring in place can be tempting because it looks simpler.

But it removes the ability to compare.

If restore goes wrong, we have neither the old state nor the new one.

Restoring beside production gives time to check:

text
whether the database started
whether the application can connect
whether critical tables have the expected number of records
whether the bug we are escaping is gone
whether we did not go too far back

In PostgreSQL, details depend on the backup setup, but the mental model is the same:

text
base backup
WAL archive
target time
promote restored instance
validation

Switching traffic

After validation, the application must be switched to the restored database.

That can be a change to:

text
DNS
connection string
platform secret
managed database endpoint
Service or routing configuration

It is important to know where the source of truth really is.

If some workers still write to the old database, the system splits into two worlds.

Before the switch, we need a list of database clients:

text
API
workers
cron jobs
admin tools
integrations
read models
migration processes

After the switch, we need to check more than the web application.

We also need to check background processes.

They often write data long after the main traffic looks fine.

What about data after the restore point

PITR means consciously returning to the past.

If we restore the database to 10:42, everything after 10:42 disappears from the restored world.

That may be acceptable.

It may also require moving some data.

Examples:

text
orders created after the restore point
payments confirmed by the provider
support tickets submitted by customers
events published to other systems

There is no universal automation here.

We need to decide which data is the business truth and which data can be reconstructed from external systems.

Sometimes the best decision is:

text
we return to 10:42 and manually reconcile payments from the last hour

That may still be better than keeping a damaged state alive.

Minimal runbook

A whole-server PITR runbook should include:

text
who decides to start PITR
how we choose target time
how we freeze writes
how we snapshot the current state
where we restore the new instance
how we validate data
how we switch all clients
how we verify the old instance no longer accepts writes
how we document lost data

The most important sentence:

text
database restore is not the end of recovery

It is only the moment when we have a candidate for new production.

It becomes production only after validation and after all traffic has been switched.