Konrad Kowalski (rootsher)Principal Platform & Reliability Architect110011110111101110011011101101101000010000100111

Rolling Back a Release: Code Is the Easy Part

date
category
Reliability Engineering
also in
CI/CD · Databases
reading
3 min / 516 words

Rolling back a release sounds simple.

text
return to the previous version

For code alone, it often really is simple.

You can point to the previous container image.

You can roll back the release in the platform.

You can switch traffic from green to blue.

You can turn off a flag.

The problem starts when the new release has already changed the world outside its own process.

Code is reversible

The application artifact is usually immutable.

text
api:1.41.0
api:1.42.0

If 1.42.0 has a bug, we can run 1.41.0.

That is the clean rollback case.

There is one condition:

text
the old version must still be able to run in the current environment

That environment includes configuration, database schema, message formats, external APIs, and data that users have already saved.

If the new version only changed discount calculation logic, rolling back code may be enough.

If the new version changed the data shape, rolling back code can start an old process in a new world.

That is where the real problem begins.

Migrations do not roll back by themselves

The most dangerous rollbacks involve migrations.

Example:

sql
ALTER TABLE orders DROP COLUMN legacy_status;

The new application version no longer uses legacy_status.

The old version still does.

If we only roll back code after this migration, the old application may fail immediately.

That is why safe schema changes are usually done in steps.

text
add the new field
write to the old and new field
switch reads to the new field
stop writing to the old field
remove the old field only after some time

This is more boring than one migration.

But it lets us roll back code in the middle of the process.

Rollback does not mean "undo the migration".

Rollback means:

text
return to the previous application version, because the schema still supports it

Data is more often one-way than code

Sometimes a migration does not remove a column.

Sometimes it changes the meaning of data.

Example:

text
status = paid

becomes:

text
payment_state = captured
fulfillment_state = ready

We can write a forward migration.

We cannot always write a backward migration without losing information.

It gets worse if users started creating new records in the new model after deployment.

The old application version may not know what to do with them.

So the question before release is not only:

text
do we have rollback?

A better question is:

text
does the previous version understand the state that can appear after deploy?

Side effects leave the database

A release can also send events, emails, webhooks, or queue jobs.

Rolling back code will not undo:

text
sent message
published event
executed payment
job waiting in a queue
cache with a new format
search index

If the new version published an event in a new format, the old consumer may not understand it.

If the new version wrote cache in a new shape, the old version may read wrong data.

If the new version sent a webhook to a partner, we cannot pretend it did not happen.

Rollback limits future damage.

It does not automatically erase effects that already happened.

A feature flag can be a better rollback

If the problem affects one path, the best rollback may be turning off a flag.

text
new code stays
new behavior disappears

That is faster than rebuilding an image or doing another rollout.

But it works only if the flag was designed as a real safety switch.

A flag that hides a button in the UI will not stop a cron job.

A flag that disables an endpoint will not undo a migration.

A flag that works only on reads will not protect writes.

Minimal rollback runbook

A good release rollback runbook should include a few points.

text
which symptom triggers rollback
who makes the decision
which artifact is the previous good version
whether migrations are backward compatible
whether flags need to be disabled
whether queues or jobs need to stop
how we check that rollback helped

The most important point is simple:

text
code rollback is not system rollback

It is only one tool.

If a bug has already damaged data, we need data recovery.