Konrad Kowalski (rootsher)Principal Platform & Reliability Architect100110011100100110000101001101010001100110010001

Cluster Policies: Why Kyverno and CEL Exist

date
category
Infrastructure & Cloud Security
also in
Containers
reading
4 min / 864 words

The problem often looks innocent. Someone is allowed to create a Deployment, RBAC says "yes", and the request moves on. The problem only becomes visible inside the manifest:

yaml
securityContext:
  privileged: true

or:

yaml
containers:
  - image: nginx:latest

RBAC answers:

text
can this person or pipeline create a Deployment?

It does not answer this question well:

text
does this specific Deployment follow the cluster rules?

That is why cluster policies exist.

What existed before

At first, control often lived outside the cluster:

text
pull request review
CI check
YAML validation script
rules written in documentation

That helps, but it has gaps. Not everything goes through the same pipeline, not every manifest has the same review process, an operator can create a resource manually, and a controller can generate an object no one reviewed in a pull request.

In Kubernetes, the real decision point is the API server. Every cluster state change goes through an API request, so control had to move closer to that place.

Admission control

Admission control runs after authentication and authorization, but before the object is stored in the cluster.

Simplified flow:

text
request
  |
  v
authentication
  |
  v
authorization / RBAC
  |
  v
admission
  |
  v
etcd

Admission answers a different question than RBAC:

text
RBAC:
  are you allowed to create a Pod?

admission:
may this Pod be created in this shape?

That is the difference between permission and policy.

Why Kyverno

Kyverno appeared as a policy engine designed for Kubernetes. Instead of writing a custom admission webhook in Go, you can write a policy as a Kubernetes resource and keep it with the rest of the platform configuration.

Example goals:

text
do not allow images with the latest tag
require an owner label
require requests and limits
block privileged containers
add default labels
verify image signatures
generate NetworkPolicy for a namespace

Kyverno runs as an admission controller: the API server sends it the request, Kyverno evaluates policies, and returns a decision. The important part is that a policy does not have to be only a ban. Kyverno can validate, mutate, generate resources, clean up resources, and verify images:

text
validate
mutate
generate
cleanup
verify images

Sometimes we want to reject a resource. Sometimes we want to complete it. Sometimes we want to generate a companion resource, for example a default NetworkPolicy for a new namespace.

Kyverno and CEL together

CEL is not Kyverno's competitor. CEL, Common Expression Language, is an expression language. It can be used directly by Kubernetes in ValidatingAdmissionPolicy, but it can also be used in Kyverno, because newer Kyverno policy types are based on CEL.

A minimal Kyverno policy can require the owner label on every Pod:

yaml
apiVersion: policies.kyverno.io/v1
kind: ValidatingPolicy
metadata:
  name: require-owner-label
spec:
  validationActions:
    - Deny
  matchConstraints:
    resourceRules:
      - apiGroups: [""]
        apiVersions: ["v1"]
        operations: ["CREATE", "UPDATE"]
        resources: ["pods"]
  validations:
    - expression: "has(object.metadata.labels.owner)"
      message: "Pod must have an owner label"

This is not application security. It protects the process of changing cluster state: if someone tries to create a Pod without the required label, the request is rejected before the object is stored.

CEL built into Kubernetes

Kubernetes added CEL-based mechanisms so some validation can run without an external webhook and a separate controller. That makes sense for simple rules that are local to the request:

text
this field must exist
this value must not be true
this image must not use the latest tag

The same condition can be written as a ValidatingAdmissionPolicy:

yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: require-owner-label
spec:
  matchConstraints:
    resourceRules:
      - apiGroups: [""]
        apiVersions: ["v1"]
        operations: ["CREATE", "UPDATE"]
        resources: ["pods"]
  validations:
    - expression: "has(object.metadata.labels.owner)"
      message: "Pod must have an owner label"

The difference is not "Kyverno or CEL". The difference is where we use CEL and how much platform machinery we need around the expression.

When Kubernetes CEL is enough

Built-in CEL admission policies are a good fit for simple validation:

text
required label
forbid privileged
forbid hostNetwork
limit a field value
simple image condition

If the policy only looks at the object from the request and returns "allow or reject", the built-in Kubernetes mechanism may be enough. It has fewer moving parts because the decision is evaluated in the API server.

When Kyverno gives more

Kyverno starts to matter when the policy does not end at a simple CEL condition:

text
resource mutation
resource generation
policy reports
image verification
exceptions
policy testing
larger policy as code workflow

In practice, both levels can live together. Simple rules can sit in native Kubernetes admission policies, while a fuller policy workflow can live in Kyverno. Or you can write CEL in Kyverno and use reports, exceptions, tests, and other platform features that ValidatingAdmissionPolicy does not try to replace.

Example: image signing with Sigstore

A good example of the difference is image signature verification. In YAML, we only see:

yaml
image: ghcr.io/example/api:1.4.2

That tells Kubernetes which image should run. It does not say who built it, whether it passed the right pipeline, or whether someone replaced the artifact in the registry.

In this model, the pipeline builds an image and signs it with Sigstore/Cosign. The cluster does not trust the string ghcr.io/example/api:1.4.2 by itself. During admission, Kyverno checks whether the image has a valid signature from a trusted attestor.

Minimal policy sketch:

yaml
apiVersion: policies.kyverno.io/v1
kind: ImageValidatingPolicy
metadata:
  name: require-signed-images
spec:
  validationActions:
    - Deny
  matchConstraints:
    resourceRules:
      - apiGroups: [""]
        apiVersions: ["v1"]
        operations: ["CREATE", "UPDATE"]
        resources: ["pods"]
  matchImageReferences:
    - glob: "ghcr.io/example/*"
  attestors:
    - name: githubActions
      cosign:
        keyless:
          identities:
            - issuer: "https://token.actions.githubusercontent.com"
              subject: "https://github.com/example/api/.github/workflows/release.yml@refs/heads/main"
        ctlog:
          url: "https://rekor.sigstore.dev"
  validations:
    - expression: >-
        images.containers.map(image,
          verifyImageSignatures(image, [attestors.githubActions])
        ).all(result, result > 0)
      message: "Container image must be signed by the release workflow"

This is a different kind of control than has(object.metadata.labels.owner). The policy has to check the image, the signature, the issuer identity, and the transparency log. CEL still appears in the validation expression, but Kyverno provides the integration with Cosign, attestors, and container images.

That is why Kyverno is used for cases like this: not because CEL is weak as expression syntax, but because expression syntax alone is not the whole supply chain workflow.

What policies do not solve

Policies do not fix bad ownership. If nobody knows who owns a namespace, a required owner label becomes just another field to bypass.

Policies do not replace team education. If the error message only says:

text
denied by policy

the platform produces frustration, not security. Policies also should not replace good defaults. If every team must remember ten fields manually, it is better to generate some of them or provide them in a template.

What this is really for

Cluster policies exist so platform rules are enforced at a place that is hard to bypass: not only in documentation, not only in CI, not only in review, but in the Kubernetes API itself.

Good policies do three things:

text
block obviously unsafe configuration
give a clear message about what to fix
automate boring repeated requirements

Bad policies do one thing:

text
turn the cluster into a black box that sometimes says no

The difference is rarely in the tool itself. It is in whether the policy describes a real platform rule, or just fear written in YAML.