Konrad Kowalski (rootsher)Principal Platform & Reliability Architect011001100101100110100011110010010110100100100111

Human-in-the-loop: when the agent cannot continue on its own

date
category
AI Agents
also in
Automation · Engineering Practices
reading
4 min / 733 words

So far we assumed a simple scenario:

text
Jira task
|
v
Implementer Agent
|
v
repo
|
v
implementation
|
v
tests
|
v
PR

In practice the agent will not always be able to get through the whole flow on its own.

Not because it "does not know what to do" technically.

Sometimes it reaches a point where the next decision needs information that is not in the repo, RAG or the ticket, or a decision the organization deliberately does not want to delegate to a model.

An example.

The ticket says:

text
PAY-123

Add automatic retry for failed payment provider calls.

Acceptance criteria:
- retry transient failures
- maximum 3 attempts
- add metrics

The agent analyzes the code and documentation and notices that the current system treats a timeout after sending a payment request as an ambiguous state.

Retrying the request could create a second payment.

In that case the agent should not decide on its own:

text
"retry will probably be fine"

It should stop and ask:

On timeout the provider may return an unknown operation result, and retrying the request without an idempotency key can create a second payment. Should this task add idempotency key support first, or do we leave timeouts out of retry?

That is human-in-the-loop.

The agent goes into a waiting state

The flow changes from:

text
RUNNING
|
v
DONE

to:

text
RUNNING
|
v
NEEDS_INPUT
|
v
WAITING_FOR_HUMAN
|
v
RUNNING
|
v
DONE

Before stopping, the agent saves a checkpoint.

For example:

json
{
  "issue": "PAY-123",
  "repository": "payments-service",
  "branch": "agent/PAY-123",
  "step": "implementation",
  "status": "waiting_for_input",
  "question": "On timeout the provider may return an unknown result, and retry without an idempotency key can create a second payment. Do we add an idempotency key first, or leave timeouts out of retry?"
}

That way, after a human answers, it does not have to start the task from scratch.

Where the question goes

The simplest option in our flow is Jira.

The agent can add a comment:

text
Implementer Agent needs input

Provider may return an unknown result after timeout.
Retrying without an idempotency key can create a duplicate payment.

Decision required:
1. Add idempotency support in this task.
2. Exclude timeout from retry scope.

At the same time the ticket can move to the status:

text
Waiting for input

Other options:

  • Teams,
  • Slack,
  • a pull request comment,
  • a dedicated agent operations panel,
  • an approval task in a pipeline.

The mechanism is secondary.

What matters is that the agent emits an explicit event:

text
human input required

How the agent gets back to work

A human answers:

text
We do not retry timeouts in this task.
Retry only on explicit 5xx responses.
Idempotency will be a separate task.

The answer produces another event:

text
Jira comment / status change
|
v
resume trigger
|
v
agent runtime
|
v
load checkpoint
|
v
continue

The agent recovers:

text
ticket
repo
branch
current step
previous decisions
human answer

and continues:

text
exclude timeout retries
|
v
implement retry for 5xx
|
v
tests
|
v
PR

Why state is needed here

This is one of the most practical reasons to have state.

Without a checkpoint, after the human answers, the agent would have to rebuild everything:

text
fetch Jira
|
v
clone repo
|
v
analyze the code
|
v
reconstruct the previous line of work
|
v
try to understand what the answer was about

With state:

text
load checkpoint
|
v
apply human decision
|
v
continue

So state is not just conversation history.

It is the state of a running process.

When the agent should ask

Not at every uncertainty.

If the agent escalates every small decision, the automation quickly stops making sense.

Good candidates for human-in-the-loop are situations where at least one of these holds:

A business decision is missing

Example:

text
Should a timeout trigger a retry
if a duplicate payment can happen?

Requirements contradict each other

For example:

text
Acceptance criteria:
- API must be backward compatible

Architecture note:
- remove the old API

The agent should not pick on its own which statement wins.

The operation has a large blast radius

For example:

text
terraform destroy
production deployment
database migration
permission escalation

Extra permissions are needed

The agent has:

text
read repository
create branch
create PR

but suddenly needs:

text
production write access

That should require an explicit decision.

Information is missing and cannot be recovered from the available sources

The agent has checked:

text
ticket
repo
RAG
tools

and still has no answer.

Only then does it escalate.

Clarification versus approval

It is worth separating two cases.

Clarification

The agent needs information:

text
Is timeout within the retry scope?

After the answer it can continue.

Approval

The agent knows what to do, but the organization requires sign-off:

text
Plan includes a production database migration.

Approve execution?

Technically the flow looks similar:

text
pause
|
v
human action
|
v
resume

but semantically it is something else.

In a larger organization it is worth distinguishing these two event types.

Timeout and escalation

What if nobody answers?

The agent should not hang forever.

You can define a policy:

text
WAITING_FOR_HUMAN
|
v
24h
|
v
reminder
|
v
48h
|
v
escalation
|
v
7 days
|
v
cancel task

This too can be part of the shared platform.

The agent repo defines:

text
when I need a human

and the platform can provide:

text
pause
resume
timeout
notification
escalation

What is configured where

ElementWhere
when the agent should escalateagent repo / workflow
checkpointruntime / state layer
content of the questionagent
communication channelplatform / integration layer
who may answeridentity / policy
resume triggerevent layer
timeoutworkflow / orchestration
audit of answersobservability / state

The final flow

Our Implementer Agent no longer looks only like this:

text
ticket
|
v
implement
|
v
PR

The real flow is closer to this:

text
Jira ticket
|
v
Implementer Agent
|
v
repo + RAG + tools
|
v
implementation
|
v
decision missing?
|-- no
|   |
|   v
|   continue
|
`-- yes
    |
    v
    checkpoint
    |
    v
    WAITING_FOR_HUMAN
    |
    v
    Jira / Teams / Slack
    |
    v
    human response
    |
    v
    resume event
    |
    v
    load checkpoint
    |
    v
    continue
    |
    v
    tests
    |
    v
    PR

So human-in-the-loop does not mean the agent stops being autonomous.

It means the organization draws a clear line:

Which decisions the agent may make on its own, and at which ones it has to stop and ask a human to decide.

Materials

top