SRE
posts (6)
- Recovery from the Ground Up: Rollback, PITR, and Returning After Failure5/5
After Recovery: Green Dashboards Are Not the End
Recovery ends the impact of an incident, but it does not end the work. We still need to reconstruct the timeline, improve the runbook, and check whether the…
- Recovery from the Ground Up: Rollback, PITR, and Returning After Failure4/5
PITR for One Tenant: Restoring One Database
A simple PITR example in a multi-tenant system: one tenant has one database, so we restore only that database instead of rolling back the whole production…
- Recovery from the Ground Up: Rollback, PITR, and Returning After Failure3/5
Whole-Server PITR: Restore Is Only the Beginning
Whole-server PITR does not end when the database is restored. We still need to freeze writes, validate state, switch traffic, and consciously abandon or move…
- Recovery from the Ground Up: Rollback, PITR, and Returning After Failure2/5
Rolling Back a Release: Code Is the Easy Part
Release rollback often looks like returning to the previous application image. In practice, the hard parts are data, migrations, and side effects that already…
- Recovery from the Ground Up: Rollback, PITR, and Returning After Failure1/5
PITR, RPO, RTO, SLI, SLO, and SLA: A Recovery Vocabulary
Recovery has its own vocabulary, but the important terms describe simple questions: how far back do we go, how much do we lose, how long does it take, and what…
- series · 5 parts
Recovery from the Ground Up: Rollback, PITR, and Returning After Failure
Recovery does not start with a large disaster recovery plan. It starts with a simple question: what exactly do we need to return to, and how much can we lose?