Konrad Kowalski (rootsher)Principal Platform & Reliability Architect111100101110000000101111100001010000001101010001

Introduction to Service Mesh: Why Istio and Linkerd Exist

date
category
Networking
also in
Containers · Infrastructure & Cloud Security
reading
3 min / 594 words

The problem does not start with the mesh.

It starts when one application stops being one application.

text
frontend
  -> api
     -> payments
     -> inventory
     -> notifications

At first, an HTTP client in code is enough.

Then the questions appear:

text
is this request encrypted?
who is really calling this service?
how long does the call to payments take?
will retry make the outage worse?
is the timeout the same in every service?
where can we see the dependency graph?

You can answer with a library in every application.

But then every application has to know the same things.

What existed before service mesh

In a simpler system, a lot of logic lived at the edge.

text
client
  |
  v
load balancer / ingress / API gateway
  |
  v
application

The edge could terminate TLS, route requests, and collect some metrics.

That works while the most important traffic enters from outside.

In microservices, a large part of the problem moves inside the cluster:

text
service A -> service B
service B -> service C
service C -> service D

Ingress does not see that whole conversation.

An API gateway usually does not see every service-to-service connection either.

If every service implements retries, timeouts, TLS, metrics, and access policy by itself, after a few years the platform has as many communication models as teams.

Service mesh appeared as an attempt to move this repeated logic out of applications.

What service mesh is

Service mesh is an infrastructure layer for traffic between services.

It usually adds:

text
mTLS between services
workload identity
request metrics
tracing
retries and timeouts
routing between versions
access policies

The important part is where this logic runs.

The application still sends a normal request:

js
await fetch("http://payments/charge");

A proxy runs next to the application or in front of it.

text
app container
  |
  v
local proxy
  |
  v
network
  |
  v
local proxy
  |
  v
other app container

The proxy can add TLS, collect metrics, apply policy, or change routing without putting that code into the service.

Simplest Kubernetes model

In the classic sidecar model, the mesh adds a proxy to the Pod.

Before the mesh, the Pod looks like this:

text
Pod
  api

After sidecar injection:

text
Pod
  api
  proxy

The application does not need to know that a proxy runs next to it.

The platform routes traffic through that proxy.

The simplest operational example is enabling the mesh for a namespace:

bash
kubectl label namespace apps istio-injection=enabled
kubectl rollout restart deployment/api -n apps

or in Linkerd:

bash
kubectl annotate namespace apps linkerd.io/inject=enabled
kubectl rollout restart deployment/api -n apps

This is not the whole production configuration.

It is only the moment when a workload becomes part of the mesh.

Why Istio

Istio is the heavier and more extensible mesh.

It makes sense when we need a lot of control over traffic, security, and policy.

Typical reasons:

text
mTLS as the default communication model
policy per workload
advanced routing
canary and traffic shifting
service-to-service traffic telemetry
Envoy integration
multi-cluster or hybrid environments

Istio has historically been associated with Envoy-based sidecars, but today it also has ambient mode, which tries to reduce the operational cost of sidecars.

It is still a tool for situations where the platform wants to program the application network strongly.

Why Linkerd

Linkerd goes in a different direction.

Its main argument is a simpler, lighter mesh for Kubernetes.

Typical reasons:

text
automatic mTLS
metrics per service
retries and timeouts
load balancing
simpler operational model
smaller proxy

If a team mostly wants encryption, basic observability, and safer service-to-service communication without building a large policy platform, Linkerd can be the more natural choice.

That does not mean Linkerd is "only simple".

Rather, it chooses a smaller operational surface.

What service mesh does not solve

Service mesh does not fix bad boundaries between services.

If services call each other too often, the mesh will make that problem more visible, but it will not remove the dependency.

Service mesh does not replace a good API.

It also does not decide which timeouts, retries, and policies are correct for the business.

It can enforce mTLS, but it will not decide by itself which service should access which operation.

It can give excellent metrics, but someone still needs to know which ones indicate a problem.

When it makes sense

Service mesh makes sense when the service-to-service problem is already real:

text
many services
many teams
inconsistent timeouts
no shared mTLS
hard dependency debugging
need for policies between workloads
rollouts that depend on internal cluster routing

It does not make sense to install a mesh only because the cluster runs Kubernetes.

If the system has a few services and simple dependencies, a mesh can add more layers than problems it solves.

The simplest question is:

text
is the problem in the application, or in repeated traffic between applications?

If the problem repeats in every service, service mesh starts being a platform tool, not architecture decoration.