Evals and quality gates
- date
- category
- AI Agents
- also in
- AI Engineering · Engineering Practices
- reading
- 4 min / 873 words
With observability in place we already know a lot.
For a single task we see:
9 model calls
4 tool calls
1 RAG query
37 seconds
31k tokens
PR created
That still does not answer the question:
Did the agent do the task well?
For that we need evals.
Where we define evals
Simplest: next to the agent.
engineering-agents/
└── implementer/
├── agent.py
├── workflow.py
└── evals/
├── cases.yaml
├── rubrics.yaml
└── datasets/
Just as code has tests/, an agent can have evals/.
When we run evals
Not only when the model version changes.
It is worth treating evals as regression tests for the agent's behavior.
1. When the model changes
GPT X -> GPT Y
Claude X -> Claude Y
The same prompt and the same tools can start producing different behavior.
This is the most obvious case.
2. When the system prompt or agent instructions change
Changing a few sentences can change:
- the order of actions,
- tool choice,
- how eagerly it uses RAG,
- how it interprets the ticket.
That is why the prompt is part of the agent's versioned configuration.
3. When a tool or its description changes
If we change:
github.create_pr
or even just its description, the model may start choosing it in different situations.
That needs a regression run too.
4. When RAG changes
For example:
- a new index,
- a new embedding model,
- different chunking,
- a different top-k,
- a new document source,
- changed permissions.
The agent code did not change, but the context it gets did.
5. When the agent workflow changes
For example:
before:
implement -> test -> PR
now:
plan -> implement -> test -> self-review -> PR
An obvious candidate for the full eval suite.
6. Before releasing a new agent version
The simplest organizational rule:
agent PR / release
|
v
eval suite
|
v
compare with baseline
|
v
quality gate
|
v
deploy
Not every change has to run the whole expensive dataset.
You can have:
small eval suite
-> every PR
full regression
-> release / bigger change
7. Periodically, even when nothing changed
This matters because something outside the agent repo can change:
- the model provider can update its backend,
- RAG can get new documents,
- the Jira / GitHub tool can return slightly different data,
- the nature of real tasks can shift.
That is why it is worth running the regression suite periodically, e.g.:
daily / weekly
depending on criticality and cost.
There is no single correct frequency.
For a coding agent a sensible start is:
small set -> every PR
full set -> before release
production sample -> daily or weekly
8. On real production tasks
Observability gives traces of real runs.
We can take a sample, e.g.:
5% of completed tasks
and score automatically:
- task completion,
- adherence,
- tool use,
- groundedness.
This catches regressions that were not in the test dataset.
An example eval case
name: retry-policy
task:
issue: PAY-123
repository: test/payments-service
expected:
tests_pass: true
pull_request_created: true
public_api_changed: false
rubric:
- uses_company_retry_standard
- adds_tests
- uses_correct_rag_source
What we measure deterministically
With ordinary code:
did the tests pass?
was a PR created?
was a forbidden file changed?
did the agent use the right repo?
did it call a disallowed tool?
What an AI evaluator can judge
Things that are harder to write as an assert:
- whether the task was really done,
- whether the agent followed the instructions,
- whether it used the right tools,
- whether it used RAG correctly,
- whether it took unnecessary steps,
- whether the final output matches expectations.
Example metrics
| Question | Metric |
|---|---|
| Did it finish the task? | task completion |
| Did it follow the rules? | task adherence |
| Did it use tools correctly? | tool accuracy |
| Was the output based on the context? | groundedness |
| Was the flow sensible? | trajectory / navigation quality |
| How much did the task cost? | cost / tokens / latency |
How often, really?
There is no point scoring 100% of everything all the time.
A practical model:
| Moment | Scope |
|---|---|
| every PR to the agent repo | small, fast regression suite |
| model / prompt / RAG / tools change | extended suite |
| new version release | full suite + baseline comparison |
| production | sampling of real traces |
| periodically | full regression even without changes |
Evals in CI
The flow may look familiar:
PR to the agent repo
|
v
CI
|
v
run eval dataset
|
v
compare with baseline
|
v
quality gate
Example:
| Metric | v1 | v2 |
|---|---|---|
| Task completion | 89% | 95% |
| Tool accuracy | 96% | 97% |
| Groundedness | 4.3 | 4.6 |
| Avg. cost / task | $0.20 | $0.34 |
| Avg. duration | 32s | 41s |
The new version is better in quality but more expensive and slower.
That makes it possible to make a deliberate decision.
Tools on the market
| Tool | Cloud / type | Role |
|---|---|---|
| Microsoft Foundry Evaluations | Azure | task/adherence/tool/output evaluation, including trace evaluation |
| Amazon Bedrock / AgentCore Evaluations | AWS | agent and model evaluation, built-in and custom evaluators |
| Vertex AI Gen AI Evaluation | GCP | final response, trajectory and tool evaluation |
| LangSmith | independent | datasets, experiments, traces, evals |
| Langfuse | open source / managed | tracing + evals + prompt management |
| Arize Phoenix | open source / managed | tracing, RAG and agent evals |
What the organization standardizes
Not every agent needs the same dataset.
What can be shared:
- the minimum required metrics,
- how evals are run,
- a standard small and full suite,
- quality thresholds,
- the evaluator/judge,
- sampling of production traces,
- a dashboard,
- the requirement to pass the quality gate before rollout.
The team still defines the cases specific to its role.