Rolling Back a Release: Code Is the Easy Part
- date
- category
- Reliability Engineering
- reading
- 3 min / 516 words
Rolling back a release sounds simple.
return to the previous version
For code alone, it often really is simple.
You can point to the previous container image.
You can roll back the release in the platform.
You can switch traffic from green to blue.
You can turn off a flag.
The problem starts when the new release has already changed the world outside its own process.
Code is reversible
The application artifact is usually immutable.
api:1.41.0
api:1.42.0
If 1.42.0 has a bug, we can run 1.41.0.
That is the clean rollback case.
There is one condition:
the old version must still be able to run in the current environment
That environment includes configuration, database schema, message formats, external APIs, and data that users have already saved.
If the new version only changed discount calculation logic, rolling back code may be enough.
If the new version changed the data shape, rolling back code can start an old process in a new world.
That is where the real problem begins.
Migrations do not roll back by themselves
The most dangerous rollbacks involve migrations.
Example:
ALTER TABLE orders DROP COLUMN legacy_status;
The new application version no longer uses legacy_status.
The old version still does.
If we only roll back code after this migration, the old application may fail immediately.
That is why safe schema changes are usually done in steps.
add the new field
write to the old and new field
switch reads to the new field
stop writing to the old field
remove the old field only after some time
This is more boring than one migration.
But it lets us roll back code in the middle of the process.
Rollback does not mean "undo the migration".
Rollback means:
return to the previous application version, because the schema still supports it
Data is more often one-way than code
Sometimes a migration does not remove a column.
Sometimes it changes the meaning of data.
Example:
status = paid
becomes:
payment_state = captured
fulfillment_state = ready
We can write a forward migration.
We cannot always write a backward migration without losing information.
It gets worse if users started creating new records in the new model after deployment.
The old application version may not know what to do with them.
So the question before release is not only:
do we have rollback?
A better question is:
does the previous version understand the state that can appear after deploy?
Side effects leave the database
A release can also send events, emails, webhooks, or queue jobs.
Rolling back code will not undo:
sent message
published event
executed payment
job waiting in a queue
cache with a new format
search index
If the new version published an event in a new format, the old consumer may not understand it.
If the new version wrote cache in a new shape, the old version may read wrong data.
If the new version sent a webhook to a partner, we cannot pretend it did not happen.
Rollback limits future damage.
It does not automatically erase effects that already happened.
A feature flag can be a better rollback
If the problem affects one path, the best rollback may be turning off a flag.
new code stays
new behavior disappears
That is faster than rebuilding an image or doing another rollout.
But it works only if the flag was designed as a real safety switch.
A flag that hides a button in the UI will not stop a cron job.
A flag that disables an endpoint will not undo a migration.
A flag that works only on reads will not protect writes.
Minimal rollback runbook
A good release rollback runbook should include a few points.
which symptom triggers rollback
who makes the decision
which artifact is the previous good version
whether migrations are backward compatible
whether flags need to be disabled
whether queues or jobs need to stop
how we check that rollback helped
The most important point is simple:
code rollback is not system rollback
It is only one tool.
If a bug has already damaged data, we need data recovery.