Konrad Kowalski (rootsher)Principal Platform & Reliability Architect111000011010110010010001100110100111110001100110

The OSI Model in the Cluster and the Cloud: Who Sees What When You Open a Page

date
category
Networking
also in
Containers · Cloud
reading
9 min / 1829 words

Someone in a meeting says:

"It runs on L4, but we do routing on L7."

A few people nod.

If you are not quite sure what those numbers mean, let's follow a single request from start to finish.

You open an online shop, click "Pay", and the browser sends a request that travels across the internet, the cloud and a Kubernetes cluster until it reaches the payments application.

It is the same request the whole way. The only thing that changes is how much each piece of infrastructure can see.

A parcel helps as a metaphor:

  • L1: the road,
  • L2: local delivery inside one building,
  • L3: the address on the envelope,
  • L4: the flat number and the delivery method,
  • L7: what the letter says.

The higher we go, the more we know about what the user actually wants.

Why layers exist at all

Networks work because no single component has to understand everything.

A network card does not need to know HTTP. A router does not need to know that a customer is paying for an order. A load balancer does not always need to read request headers.

Each layer solves its own problem and hands the data on.

The classic OSI model has seven layers. In practice the internet runs on the simpler TCP/IP model, but the OSI numbers stuck in the language of infrastructure.

That is why you still hear:

  • "L4 load balancer",
  • "L7 routing",
  • "L3/L4 firewall".

You can read those numbers as shorthand for one question:

how deeply does this component look into the traffic?

Let's see it on our payment.

L1: physics

You click "Pay".

At the very bottom, everything is physics: an electrical signal, light in a fibre, or a Wi-Fi radio wave.

That is L1.

This is where cables, ports, network cards and physical switches live.

In the cloud you usually do not see any of it. The provider hides the hardware behind an API, and you mostly feel it as a bandwidth limit on an instance or a network interface.

L1 knows nothing about IP, ports or HTTP.

It carries bits.

L2: who is my neighbour

The request now has to be handed to the next device on the local network.

This is L2.

The classic address at this layer is the MAC address. With IPv4 there is also ARP, which helps work out which MAC address belongs to a given IP on the local network.

If L3 says:

"I want to reach this IP address"

L2 replies:

"Fine, but which device right here should I hand the frame to?"

In Kubernetes the network is virtual, but the principle is much the same.

The payments pod has its own network interface, often connected to the node's network with a veth pair.

Simplified:

text
pod payments <---- veth ----> node network

On the node side, traffic may go through a bridge, kernel routing, eBPF or whatever else the particular CNI plugin uses.

You do not need those details to get the main point:

a pod is plugged into the node's virtual network roughly the way a device is plugged into a switch.

L2 still does not know a payment is happening.

What it knows is who to hand the data to locally.

L3: the IP address and the route

Now the request has to find its way to the right network.

We are at L3.

What matters here is the IP address.

Your browser talks to the shop's public address. From there, traffic may pass through a load balancer, the cloud network and the cluster nodes before it finally reaches the right pod.

In Kubernetes every pod has its own IP address.

Say payments has:

text
10.42.3.17

The customer's browser obviously does not know that address. The infrastructure along the way steers traffic where it needs to go.

In the cloud, a similar role is played by a VPC or VNet, subnets and route tables.

Logically, routing can look like this:

text
10.42.0.0/16 -> cluster network
10.50.0.0/16 -> another network
0.0.0.0/0    -> internet

A router looks at the destination IP and answers one question:

"Which way should this packet go?"

It does not care whether the user is paying, logging in or downloading a product photo.

This is also where security rules show up.

In Kubernetes that can be a NetworkPolicy; in the cloud, a security group or a firewall.

Simplified:

text
frontend -> payments: allowed
internet -> database: denied

It is a decision of the kind:

who is allowed to talk to whom?

We do not yet care whether the request is:

http
POST /payments

or:

http
GET /health

For that we need a higher layer.

L4: the connection and the port

An IP address alone is not enough.

One machine can run many services, so we also need a port number.

For example:

text
10.42.3.17:8080

The IP says where.

The port says which service.

A bit like a building address and a flat number.

L4 is mostly TCP and UDP.

TCP

TCP sets up a connection and takes care of, among other things, data ordering and retransmitting lost pieces.

UDP

UDP is simpler. It sends datagrams but does not offer the same reliability guarantees.

For our story, something else matters more.

The payments pod can disappear and be replaced by a new one. We may also run several copies of it:

text
payments-1
payments-2
payments-3

That is why a client inside the cluster usually does not connect to a specific pod directly.

In Kubernetes we use a Service, which gives a stable access point and sends traffic to one of the backends.

From the L4 point of view, the infrastructure may see:

text
protocol: TCP
port: 8080

and decide where to send the connection based on that.

A network load balancer in the cloud, such as an NLB, works in a similar way.

It may see:

text
source IP
destination IP
TCP
port 443

and spread connections across backends.

It does not need to know there is HTTP inside.

It does not need to read the URL.

It does not need to know the headers.

That is what L4 load balancing means.

The balancer sees the connection.

It does not yet understand the conversation.

L5 and L6: the layers almost nobody counts

Classic OSI also has a session layer and a presentation layer.

In everyday infrastructure conversations you will almost never hear:

"We have a problem on L6."

Those functions did not disappear. We just stopped calling them by number.

In our payment, TLS is one example.

When you open:

text
https://shop.example.com

the data is encrypted.

The request also has a particular representation. It may carry JSON:

json
{
  "orderId": "12345",
  "amount": 19900,
  "currency": "PLN"
}

or protobuf. It may also be compressed.

Modern infrastructure simply says:

  • TLS,
  • JSON,
  • protobuf,
  • gzip.

That is why conversations usually jump from:

text
L4

straight to:

text
L7

L7: the infrastructure understands the request

We reach the application layer.

Here the infrastructure can see that the data is an HTTP request.

It no longer sees only:

text
10.20.1.8 -> 10.42.3.17:8080

It can see:

http
POST /api/payments
Host: shop.example.com
Content-Type: application/json
X-Preview: true

And that changes everything.

At L3 we asked:

Which IP is the traffic coming from, and where is it going?

At L4:

Which port?

At L7 we can ask:

What is this user actually trying to do?

That is why a gateway or an L7 proxy can route traffic by path:

text
/api/catalog  -> catalog-service
/api/orders   -> orders-service
/api/payments -> payments-service

All of it can arrive at the same IP address and port 443.

Only the HTTP proxy reads the request and picks the right backend.

Canary

We have two versions:

text
payments-v1
payments-v2

We want to try the new one.

An L7 proxy can do:

text
90% -> payments-v1
10% -> payments-v2

Header-based routing

Testers send:

http
X-Preview: true

The gateway can send just them to payments-v2.

Retry and timeout

The proxy can also say:

if the backend briefly fails to respond, try again.

Or:

if it does not respond within two seconds, end the request.

You will find these mechanisms in gateways such as Envoy Gateway, in cloud application load balancers such as the ALB, and in service meshes.

The common denominator is simple:

to make a decision like this, the infrastructure has to understand the application protocol.

That is L7.

The whole journey in one go

The user clicks "Pay".

The browser builds:

http
POST /api/payments

At L7, the gateway sees the path /api/payments and decides the request goes to payments.

TLS encrypts the communication, and the payload may be JSON.

At L4, traffic uses TCP and port 443, and maybe 8080 later inside the cluster. A load balancer or a Kubernetes Service picks a backend.

At L3, packets carry IP addresses. Routing in the cloud and in the cluster decides how they get to the right node and pod.

At L2, data is passed between local interfaces and network neighbours.

At L1, everything ends up as a signal travelling through the physical infrastructure of a data centre.

The request reaches payments.

And the response comes back along a similar path.

The higher you go, the more you see

This is the most important rule.

At L3 the infrastructure sees roughly:

text
IP -> IP

At L4:

text
IP:port -> IP:port

At L7:

http
POST /api/payments

and even:

http
X-Preview: true

The higher you go, the more you can do.

But that knowledge has a price.

An L4 proxy can be relatively simple. It does not have to parse HTTP.

An L7 proxy has to understand the protocol, parse requests, keep more state, and often take part in TLS termination as well.

So the thing to remember is:

L4 is faster, simpler and blinder.

L7 is smarter, but costs more resources and complexity.

The trap: a policy only protects traffic that passes through it

Picture a security guard standing at one entrance to an office.

They check the badge of everyone who comes through the door.

The rule is:

only people from the finance team may enter.

That works perfectly, as long as everyone uses that door.

If there is a side entrance that lets you get around the guard, the rule stops being full protection.

L7 policies work the same way.

If a security rule is enforced by an L7 proxy, it protects the traffic that actually goes through that proxy.

If some of the traffic can bypass it, that particular control will not apply.

This matters in modern service meshes too, where traffic can be intercepted and controlled in several different places.

So instead of asking only:

"Do we have a policy?"

it is better to ask:

"Where is it enforced, and does all the traffic we care about go through that point?"

Cheat sheet

LayerWhat it seesIn the clusterIn the cloud
L1signal, bitsphysical network under the nodescables, NICs, bandwidth
L2local neighbours, MACveth, node networkvirtualised Ethernet
L3IP addresses, routespod IPs, routing, NetworkPolicyVPC/VNet, routing, security groups
L4TCP/UDP, portsService, forwarding to podsNLB and transport load balancers
L5/L6session, encryption, formatTLS, JSON, protobufTLS, certificates, encoding
L7HTTP, URL, method, headersGateway, ingress, Envoy, service meshALB, application proxies

If you only remember three lines:

text
L3: where to?
L4: which port?
L7: what does the client want to do?

Or the parcel version:

text
L1: how it physically travels
L2: who to hand it to locally
L3: which address
L4: which flat
L7: what the letter says

So the sentence:

"The NLB does L4, but path-based routing needs L7"

simply means:

L4 sees the connection. L7 understands the request.

top