PITR, RPO, RTO, SLI, SLO, and SLA: A Recovery Vocabulary
- date
- category
- Reliability Engineering
- also in
- Incident Management
- reading
- 4 min / 735 words
During an incident, it is easy to confuse words that sound similar but lead to different decisions.
Someone says "backup".
Someone says "rollback".
Someone says "restore the state from an hour ago".
Someone asks whether we broke the SLA.
These are not language details.
Those words decide whether the team is trying to roll back code, restore data, switch traffic, or only reduce the impact.
PITR
PITR means point-in-time recovery.
The simplest definition:
restoring data to a specific point in time
Not "to the last backup".
Not "to the newest available state".
To a chosen moment.
Example:
restore the database to 2025-01-14 10:42:15
In PostgreSQL, this usually means combining a base backup with WAL records.
The base backup gives the starting point.
WAL lets the database replay changes up to the selected moment.
This matters because a bug is not always detected immediately.
If a bad job deleted data at 10:43 and monitoring fired at 11:20, the latest 10:00 backup may be too old, while the 11:20 state is already broken.
PITR gives a third option:
return just before the bad write
RPO
RPO means recovery point objective.
It answers this question:
how much data can we lose?
If RPO is 15 minutes, the organization accepts the risk of missing the last 15 minutes of writes after a failure.
That does not mean we always lose that much.
It means the system is designed around that limit.
RPO forces concrete decisions:
how often we take backups
whether we archive WAL or binlog
where we keep copies
how quickly we detect broken replication
whether we test restore, not only backup creation
A daily backup and a 5 minute RPO is wishful thinking.
It can be written in a document, but the system does not meet it.
RTO
RTO means recovery time objective.
It answers this question:
how long can we take to return to service?
If RTO is one hour, the recovery plan must fit within one hour from the decision or from the start of the incident, depending on how the organization defines it.
Sometimes TRO appears as a typo.
In practice, it almost always means RTO.
RTO covers more than restoring data.
We need to count:
detection time
decision time
procedure start time
restore time
validation time
traffic switch time
communication time
If the database restores in 20 minutes, but the team spends 40 minutes finding someone with permissions, the real RTO is not 20 minutes.
It is at least an hour.
SLI and SLO
Before SLO, we need to name SLI.
SLI means service level indicator.
It is a concrete measure of service behavior.
Example:
percentage of successful responses
p99 latency
percentage of successful writes
time from failure to detection
SLI must be countable.
It is not team mood or a general sentence like "the system is slow".
If we say during recovery that "checkout is back", SLI should show what that means.
whether it responds
whether it writes orders
whether it stays within latency
whether it does not return errors
SLO means service level objective.
It is a technical target based on an SLI.
In other words:
SLI measures
SLO says which result we consider good
Example:
99.9% successful responses in a month
p99 latency below 300 ms
99.95% availability for the checkout endpoint
SLO helps during an incident.
If the bug affects a small feature that does not touch the main SLO, rollback may be less urgent.
If the bug breaks the checkout path, every few minutes burn the error budget.
SLO does not say how to recover.
It says why recovery is urgent and what impact we are actually measuring.
Error budget
Error budget is the difference between perfect service and the accepted SLO.
If SLO says 99.9% successful responses, the system can have 0.1% errors in that window.
That is the error budget.
During an incident, the error budget shows how quickly we are consuming the margin.
It does not say how to restore a database.
It helps decide whether the problem needs to be cut off immediately or whether there is time to finish diagnosis calmly.
SLA
SLA means service level agreement.
It is not only a technical goal.
It is a commitment to a customer, business unit, or another side of an agreement.
SLA can have financial, formal, or reputational consequences.
The common distinction is:
SLO is what we target internally
SLA is what we promised externally
If SLO is stricter than SLA, the team has a buffer.
If SLA is stricter than the real SLO, the organization is selling a promise the system cannot keep.
Putting it together
These terms form a simple map.
PITR: which point we can restore data to
RPO: how much data we can lose
RTO: how long recovery can take
SLI: how we measure service behavior
SLO: which level of service we want to maintain
SLA: what we promised externally
Two more acronyms are worth knowing because they often appear after an incident.
MTTD means mean time to detect, the average time to detect a problem.
MTTR means mean time to recover or mean time to restore, depending on the organization. In practice, it is the average time to return to service.
They are not as precise as a specific RTO for a specific service, but they help show whether the team detects and closes incidents faster over time.
During an incident, we do not need an academic discussion.
Five questions are enough:
what is broken
since when is it broken
how much data can we lose
how much time do we have to return
did we break a promise to users
Only after that does it make sense to choose a mechanism.
Code rollback, backup restore, and PITR solve different problems.