Konrad Kowalski (rootsher)Principal Platform & Reliability Architect101110011111000110101110000011000010111101001101

Evals and quality gates

date
category
AI Agents
also in
AI Engineering · Engineering Practices
reading
4 min / 873 words

With observability in place we already know a lot.

For a single task we see:

text
9 model calls
4 tool calls
1 RAG query
37 seconds
31k tokens
PR created

That still does not answer the question:

Did the agent do the task well?

For that we need evals.

Where we define evals

Simplest: next to the agent.

text
engineering-agents/
└── implementer/
    ├── agent.py
    ├── workflow.py
    └── evals/
        ├── cases.yaml
        ├── rubrics.yaml
        └── datasets/

Just as code has tests/, an agent can have evals/.

When we run evals

Not only when the model version changes.

It is worth treating evals as regression tests for the agent's behavior.

1. When the model changes

text
GPT X -> GPT Y
Claude X -> Claude Y

The same prompt and the same tools can start producing different behavior.

This is the most obvious case.

2. When the system prompt or agent instructions change

Changing a few sentences can change:

  • the order of actions,
  • tool choice,
  • how eagerly it uses RAG,
  • how it interprets the ticket.

That is why the prompt is part of the agent's versioned configuration.

3. When a tool or its description changes

If we change:

text
github.create_pr

or even just its description, the model may start choosing it in different situations.

That needs a regression run too.

4. When RAG changes

For example:

  • a new index,
  • a new embedding model,
  • different chunking,
  • a different top-k,
  • a new document source,
  • changed permissions.

The agent code did not change, but the context it gets did.

5. When the agent workflow changes

For example:

text
before:
implement -> test -> PR

now:
plan -> implement -> test -> self-review -> PR

An obvious candidate for the full eval suite.

6. Before releasing a new agent version

The simplest organizational rule:

text
agent PR / release
|
v
eval suite
|
v
compare with baseline
|
v
quality gate
|
v
deploy

Not every change has to run the whole expensive dataset.

You can have:

text
small eval suite
-> every PR

full regression
-> release / bigger change

7. Periodically, even when nothing changed

This matters because something outside the agent repo can change:

  • the model provider can update its backend,
  • RAG can get new documents,
  • the Jira / GitHub tool can return slightly different data,
  • the nature of real tasks can shift.

That is why it is worth running the regression suite periodically, e.g.:

text
daily / weekly

depending on criticality and cost.

There is no single correct frequency.

For a coding agent a sensible start is:

text
small set -> every PR
full set -> before release
production sample -> daily or weekly

8. On real production tasks

Observability gives traces of real runs.

We can take a sample, e.g.:

text
5% of completed tasks

and score automatically:

  • task completion,
  • adherence,
  • tool use,
  • groundedness.

This catches regressions that were not in the test dataset.

An example eval case

yaml
name: retry-policy

task:
  issue: PAY-123
  repository: test/payments-service

expected:
  tests_pass: true
  pull_request_created: true
  public_api_changed: false

rubric:
  - uses_company_retry_standard
  - adds_tests
  - uses_correct_rag_source

What we measure deterministically

With ordinary code:

text
did the tests pass?
was a PR created?
was a forbidden file changed?
did the agent use the right repo?
did it call a disallowed tool?

What an AI evaluator can judge

Things that are harder to write as an assert:

  • whether the task was really done,
  • whether the agent followed the instructions,
  • whether it used the right tools,
  • whether it used RAG correctly,
  • whether it took unnecessary steps,
  • whether the final output matches expectations.

Example metrics

QuestionMetric
Did it finish the task?task completion
Did it follow the rules?task adherence
Did it use tools correctly?tool accuracy
Was the output based on the context?groundedness
Was the flow sensible?trajectory / navigation quality
How much did the task cost?cost / tokens / latency

How often, really?

There is no point scoring 100% of everything all the time.

A practical model:

MomentScope
every PR to the agent reposmall, fast regression suite
model / prompt / RAG / tools changeextended suite
new version releasefull suite + baseline comparison
productionsampling of real traces
periodicallyfull regression even without changes

Evals in CI

The flow may look familiar:

text
PR to the agent repo
|
v
CI
|
v
run eval dataset
|
v
compare with baseline
|
v
quality gate

Example:

Metricv1v2
Task completion89%95%
Tool accuracy96%97%
Groundedness4.34.6
Avg. cost / task$0.20$0.34
Avg. duration32s41s

The new version is better in quality but more expensive and slower.

That makes it possible to make a deliberate decision.

Tools on the market

ToolCloud / typeRole
Microsoft Foundry EvaluationsAzuretask/adherence/tool/output evaluation, including trace evaluation
Amazon Bedrock / AgentCore EvaluationsAWSagent and model evaluation, built-in and custom evaluators
Vertex AI Gen AI EvaluationGCPfinal response, trajectory and tool evaluation
LangSmithindependentdatasets, experiments, traces, evals
Langfuseopen source / managedtracing + evals + prompt management
Arize Phoenixopen source / managedtracing, RAG and agent evals

What the organization standardizes

Not every agent needs the same dataset.

What can be shared:

  • the minimum required metrics,
  • how evals are run,
  • a standard small and full suite,
  • quality thresholds,
  • the evaluator/judge,
  • sampling of production traces,
  • a dashboard,
  • the requirement to pass the quality gate before rollout.

The team still defines the cases specific to its role.

Materials

top