Konrad Kowalski (rootsher)Principal Platform & Reliability Architect111111010110011101100100100000101010011001110010

Rolling Update: Replacing Pieces Gradually

date
category
CI/CD
also in
Reliability Engineering
reading
1 min / 246 words

At first, deployment meant an interruption.

text
stop old
start new

When applications started running as multiple instances behind a load balancer, a simpler option appeared:

text
remove one old instance
add one new instance
repeat

That is rolling update.

The whole system is not switched at once.

It is replaced piece by piece.

Minimal example

In Kubernetes, Deployment uses RollingUpdate by default.

It is still worth writing the parameters explicitly:

yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api
spec:
  replicas: 4
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  selector:
    matchLabels:
      app: api
  template:
    metadata:
      labels:
        app: api
    spec:
      containers:
        - name: api
          image: ghcr.io/example/api:2.0.0
          ports:
            - containerPort: 3000
          readinessProbe:
            httpGet:
              path: /ready
              port: 3000

With replicas: 4, maxSurge: 1, and maxUnavailable: 0, Kubernetes can temporarily run a fifth instance, but should not go below four ready Pods.

readinessProbe is the traffic gate here.

A new Pod should not receive requests just because the process started.

It should receive traffic only after it can handle it.

What it is really for

Rolling update is good when the new version is compatible with the old one.

For a while, both versions live in the system:

text
api:1.9.0
api:2.0.0

That must be safe for:

text
HTTP API
events
database schema
cache
background workers

If the new version writes data that the old version cannot read, rolling update becomes risky.

Not because Kubernetes deploys badly.

Because the application does not tolerate mixed versions.

Where the name lies

Rolling update sounds like a safety strategy.

Often, it is only an availability strategy.

It protects against simple downtime:

text
all instances disappeared at once

It does not protect against a logic bug gradually reaching more Pods.

Rollback does not reverse the world either.

You can return to the previous container image, but you do not automatically undo data migrations, sent emails, or published events.

Rolling update is a good default mechanism.

It is not permission to ignore compatibility.