Konrad Kowalski (rootsher)Principal Platform & Reliability Architect111011111001000001111101000100101001111100111001

After Recovery: Green Dashboards Are Not the End

date
category
Reliability Engineering
also in
Incident Management
reading
3 min / 545 words

The easiest mistake is ending an incident too early.

Dashboards are green again.

The application responds.

Traffic goes to the right database.

Someone writes in the channel that "it works now".

That is an important moment.

But it is not the end of the work.

It is the end of the sharpest phase.

Confirm state first

After recovery, we need to check whether the system is really back.

Not only whether an endpoint returns 200.

We need to check the paths affected by the incident.

text
writes
reads
workers
queues
cron jobs
integrations
cache
reports
alerts

If we switched a database, we need to confirm that all clients use the new endpoint.

If we stopped workers, we need to start them deliberately.

If we froze writes, we need to know when they were unfrozen.

If we abandoned data after target time, we need to record exactly what was lost.

Without that, the team only has relief.

It does not have confirmed system state.

Reconstruct the timeline

A postmortem starts with a timeline.

Not with blame.

Not with one grand root cause.

First, we build chronology.

text
when the failure started
when the system detected it
when a human saw it
when the decision was made
when damage was stopped
when recovery started
when service was restored
when correct state was confirmed

The timeline reveals the truth about RTO.

If restore itself took 25 minutes, but the decision took an hour, the procedure does not fit in 25 minutes.

If the alert appeared after 40 minutes, technical recovery may have been good, but detection was weak.

If the team waited for access to secrets, the problem is operational readiness, not the database.

Check RPO

After recovery, we need to count the real data loss.

text
what target time was used
when writes were stopped
what data appeared between those moments
what was recovered
what needs manual reconstruction
what we consider lost

This should go into the incident note.

Not to dramatize.

To check whether the declared RPO was real.

If the organization talked about a 15 minute RPO, but actually lost two hours of data, that is not a small difference.

It is information about architecture, backups, detection, or decisions.

Improve the runbook

The runbook after an incident usually looks different from the runbook before the incident.

That is normal.

Recovery exposes things that were invisible in the document.

text
a command was missing
a link pointed to an old panel
only one person had permissions
validation was vague
there was no worker list
there was no decision owner for traffic switch
there was no plan for data after target time

Those fixes should be made quickly.

Not in a month.

Operational memory evaporates very quickly after an incident.

The best time to improve the runbook is while the team still remembers where it hurt.

Add a restore test

A backup that nobody restored is a hypothesis.

After an incident, it is worth adding the smallest restore test.

text
take the latest backup
restore to a chosen moment
start the application against it
check a few critical queries
measure time
save the result

It does not need to be a full disaster recovery drill immediately.

A small, repeatable test is better than a large document nobody runs.

More can be added over time:

text
sequence test
integrity test
login test
order write test
worker test
endpoint switch test

Recovery is a skill.

Skills do not appear just because instructions exist.

Postmortem without theater

A good postmortem should answer a few questions.

text
what happened
what was the impact
how we detected it
what we did
what worked
what did not work
what we are changing
who owns the changes
by when

The point is not to find the person who clicked the wrong button.

The point is to make the system less dependent on perfect people.

If one command could delete production data without a guardrail, the problem is not only a person.

The problem is the missing guardrail.

If the PITR decision took an hour, the problem is not only stress.

The problem is missing decision thresholds prepared earlier.

End of the series, start of practice

Simple recovery consists of a few boring elements:

text
known terms
compatible migrations
backup that can be restored
WAL or binlog
list of database clients
traffic switch plan
runbook
restore test
postmortem

None of these elements is spectacular.

Together, they make the difference between improvisation and a controlled return to service.

That is the whole magic.

Which means none.