Konrad Kowalski (rootsher)Principal Platform & Reliability Architect101001001100010111110101010001101111000010110000

PITR, RPO, RTO, SLI, SLO, and SLA: A Recovery Vocabulary

date
category
Reliability Engineering
also in
Incident Management
reading
4 min / 735 words

During an incident, it is easy to confuse words that sound similar but lead to different decisions.

Someone says "backup".

Someone says "rollback".

Someone says "restore the state from an hour ago".

Someone asks whether we broke the SLA.

These are not language details.

Those words decide whether the team is trying to roll back code, restore data, switch traffic, or only reduce the impact.

PITR

PITR means point-in-time recovery.

The simplest definition:

text
restoring data to a specific point in time

Not "to the last backup".

Not "to the newest available state".

To a chosen moment.

Example:

text
restore the database to 2025-01-14 10:42:15

In PostgreSQL, this usually means combining a base backup with WAL records.

The base backup gives the starting point.

WAL lets the database replay changes up to the selected moment.

This matters because a bug is not always detected immediately.

If a bad job deleted data at 10:43 and monitoring fired at 11:20, the latest 10:00 backup may be too old, while the 11:20 state is already broken.

PITR gives a third option:

text
return just before the bad write

RPO

RPO means recovery point objective.

It answers this question:

text
how much data can we lose?

If RPO is 15 minutes, the organization accepts the risk of missing the last 15 minutes of writes after a failure.

That does not mean we always lose that much.

It means the system is designed around that limit.

RPO forces concrete decisions:

text
how often we take backups
whether we archive WAL or binlog
where we keep copies
how quickly we detect broken replication
whether we test restore, not only backup creation

A daily backup and a 5 minute RPO is wishful thinking.

It can be written in a document, but the system does not meet it.

RTO

RTO means recovery time objective.

It answers this question:

text
how long can we take to return to service?

If RTO is one hour, the recovery plan must fit within one hour from the decision or from the start of the incident, depending on how the organization defines it.

Sometimes TRO appears as a typo.

In practice, it almost always means RTO.

RTO covers more than restoring data.

We need to count:

text
detection time
decision time
procedure start time
restore time
validation time
traffic switch time
communication time

If the database restores in 20 minutes, but the team spends 40 minutes finding someone with permissions, the real RTO is not 20 minutes.

It is at least an hour.

SLI and SLO

Before SLO, we need to name SLI.

SLI means service level indicator.

It is a concrete measure of service behavior.

Example:

text
percentage of successful responses
p99 latency
percentage of successful writes
time from failure to detection

SLI must be countable.

It is not team mood or a general sentence like "the system is slow".

If we say during recovery that "checkout is back", SLI should show what that means.

text
whether it responds
whether it writes orders
whether it stays within latency
whether it does not return errors

SLO means service level objective.

It is a technical target based on an SLI.

In other words:

text
SLI measures
SLO says which result we consider good

Example:

text
99.9% successful responses in a month
p99 latency below 300 ms
99.95% availability for the checkout endpoint

SLO helps during an incident.

If the bug affects a small feature that does not touch the main SLO, rollback may be less urgent.

If the bug breaks the checkout path, every few minutes burn the error budget.

SLO does not say how to recover.

It says why recovery is urgent and what impact we are actually measuring.

Error budget

Error budget is the difference between perfect service and the accepted SLO.

If SLO says 99.9% successful responses, the system can have 0.1% errors in that window.

That is the error budget.

During an incident, the error budget shows how quickly we are consuming the margin.

It does not say how to restore a database.

It helps decide whether the problem needs to be cut off immediately or whether there is time to finish diagnosis calmly.

SLA

SLA means service level agreement.

It is not only a technical goal.

It is a commitment to a customer, business unit, or another side of an agreement.

SLA can have financial, formal, or reputational consequences.

The common distinction is:

text
SLO is what we target internally
SLA is what we promised externally

If SLO is stricter than SLA, the team has a buffer.

If SLA is stricter than the real SLO, the organization is selling a promise the system cannot keep.

Putting it together

These terms form a simple map.

text
PITR: which point we can restore data to
RPO: how much data we can lose
RTO: how long recovery can take
SLI: how we measure service behavior
SLO: which level of service we want to maintain
SLA: what we promised externally

Two more acronyms are worth knowing because they often appear after an incident.

MTTD means mean time to detect, the average time to detect a problem.

MTTR means mean time to recover or mean time to restore, depending on the organization. In practice, it is the average time to return to service.

They are not as precise as a specific RTO for a specific service, but they help show whether the team detects and closes incidents faster over time.

During an incident, we do not need an academic discussion.

Five questions are enough:

text
what is broken
since when is it broken
how much data can we lose
how much time do we have to return
did we break a promise to users

Only after that does it make sense to choose a mechanism.

Code rollback, backup restore, and PITR solve different problems.