After Recovery: Green Dashboards Are Not the End
- date
- category
- Reliability Engineering
- also in
- Incident Management
- reading
- 3 min / 545 words
The easiest mistake is ending an incident too early.
Dashboards are green again.
The application responds.
Traffic goes to the right database.
Someone writes in the channel that "it works now".
That is an important moment.
But it is not the end of the work.
It is the end of the sharpest phase.
Confirm state first
After recovery, we need to check whether the system is really back.
Not only whether an endpoint returns 200.
We need to check the paths affected by the incident.
writes
reads
workers
queues
cron jobs
integrations
cache
reports
alerts
If we switched a database, we need to confirm that all clients use the new endpoint.
If we stopped workers, we need to start them deliberately.
If we froze writes, we need to know when they were unfrozen.
If we abandoned data after target time, we need to record exactly what was lost.
Without that, the team only has relief.
It does not have confirmed system state.
Reconstruct the timeline
A postmortem starts with a timeline.
Not with blame.
Not with one grand root cause.
First, we build chronology.
when the failure started
when the system detected it
when a human saw it
when the decision was made
when damage was stopped
when recovery started
when service was restored
when correct state was confirmed
The timeline reveals the truth about RTO.
If restore itself took 25 minutes, but the decision took an hour, the procedure does not fit in 25 minutes.
If the alert appeared after 40 minutes, technical recovery may have been good, but detection was weak.
If the team waited for access to secrets, the problem is operational readiness, not the database.
Check RPO
After recovery, we need to count the real data loss.
what target time was used
when writes were stopped
what data appeared between those moments
what was recovered
what needs manual reconstruction
what we consider lost
This should go into the incident note.
Not to dramatize.
To check whether the declared RPO was real.
If the organization talked about a 15 minute RPO, but actually lost two hours of data, that is not a small difference.
It is information about architecture, backups, detection, or decisions.
Improve the runbook
The runbook after an incident usually looks different from the runbook before the incident.
That is normal.
Recovery exposes things that were invisible in the document.
a command was missing
a link pointed to an old panel
only one person had permissions
validation was vague
there was no worker list
there was no decision owner for traffic switch
there was no plan for data after target time
Those fixes should be made quickly.
Not in a month.
Operational memory evaporates very quickly after an incident.
The best time to improve the runbook is while the team still remembers where it hurt.
Add a restore test
A backup that nobody restored is a hypothesis.
After an incident, it is worth adding the smallest restore test.
take the latest backup
restore to a chosen moment
start the application against it
check a few critical queries
measure time
save the result
It does not need to be a full disaster recovery drill immediately.
A small, repeatable test is better than a large document nobody runs.
More can be added over time:
sequence test
integrity test
login test
order write test
worker test
endpoint switch test
Recovery is a skill.
Skills do not appear just because instructions exist.
Postmortem without theater
A good postmortem should answer a few questions.
what happened
what was the impact
how we detected it
what we did
what worked
what did not work
what we are changing
who owns the changes
by when
The point is not to find the person who clicked the wrong button.
The point is to make the system less dependent on perfect people.
If one command could delete production data without a guardrail, the problem is not only a person.
The problem is the missing guardrail.
If the PITR decision took an hour, the problem is not only stress.
The problem is missing decision thresholds prepared earlier.
End of the series, start of practice
Simple recovery consists of a few boring elements:
known terms
compatible migrations
backup that can be restored
WAL or binlog
list of database clients
traffic switch plan
runbook
restore test
postmortem
None of these elements is spectacular.
Together, they make the difference between improvisation and a controlled return to service.
That is the whole magic.
Which means none.