Human-in-the-loop: when the agent cannot continue on its own
- date
- category
- AI Agents
- also in
- Automation · Engineering Practices
- reading
- 4 min / 733 words
So far we assumed a simple scenario:
Jira task
|
v
Implementer Agent
|
v
repo
|
v
implementation
|
v
tests
|
v
PR
In practice the agent will not always be able to get through the whole flow on its own.
Not because it "does not know what to do" technically.
Sometimes it reaches a point where the next decision needs information that is not in the repo, RAG or the ticket, or a decision the organization deliberately does not want to delegate to a model.
An example.
The ticket says:
PAY-123
Add automatic retry for failed payment provider calls.
Acceptance criteria:
- retry transient failures
- maximum 3 attempts
- add metrics
The agent analyzes the code and documentation and notices that the current system treats a timeout after sending a payment request as an ambiguous state.
Retrying the request could create a second payment.
In that case the agent should not decide on its own:
"retry will probably be fine"
It should stop and ask:
On timeout the provider may return an unknown operation result, and retrying the request without an idempotency key can create a second payment. Should this task add idempotency key support first, or do we leave timeouts out of retry?
That is human-in-the-loop.
The agent goes into a waiting state
The flow changes from:
RUNNING
|
v
DONE
to:
RUNNING
|
v
NEEDS_INPUT
|
v
WAITING_FOR_HUMAN
|
v
RUNNING
|
v
DONE
Before stopping, the agent saves a checkpoint.
For example:
{
"issue": "PAY-123",
"repository": "payments-service",
"branch": "agent/PAY-123",
"step": "implementation",
"status": "waiting_for_input",
"question": "On timeout the provider may return an unknown result, and retry without an idempotency key can create a second payment. Do we add an idempotency key first, or leave timeouts out of retry?"
}
That way, after a human answers, it does not have to start the task from scratch.
Where the question goes
The simplest option in our flow is Jira.
The agent can add a comment:
Implementer Agent needs input
Provider may return an unknown result after timeout.
Retrying without an idempotency key can create a duplicate payment.
Decision required:
1. Add idempotency support in this task.
2. Exclude timeout from retry scope.
At the same time the ticket can move to the status:
Waiting for input
Other options:
- Teams,
- Slack,
- a pull request comment,
- a dedicated agent operations panel,
- an approval task in a pipeline.
The mechanism is secondary.
What matters is that the agent emits an explicit event:
human input required
How the agent gets back to work
A human answers:
We do not retry timeouts in this task.
Retry only on explicit 5xx responses.
Idempotency will be a separate task.
The answer produces another event:
Jira comment / status change
|
v
resume trigger
|
v
agent runtime
|
v
load checkpoint
|
v
continue
The agent recovers:
ticket
repo
branch
current step
previous decisions
human answer
and continues:
exclude timeout retries
|
v
implement retry for 5xx
|
v
tests
|
v
PR
Why state is needed here
This is one of the most practical reasons to have state.
Without a checkpoint, after the human answers, the agent would have to rebuild everything:
fetch Jira
|
v
clone repo
|
v
analyze the code
|
v
reconstruct the previous line of work
|
v
try to understand what the answer was about
With state:
load checkpoint
|
v
apply human decision
|
v
continue
So state is not just conversation history.
It is the state of a running process.
When the agent should ask
Not at every uncertainty.
If the agent escalates every small decision, the automation quickly stops making sense.
Good candidates for human-in-the-loop are situations where at least one of these holds:
A business decision is missing
Example:
Should a timeout trigger a retry
if a duplicate payment can happen?
Requirements contradict each other
For example:
Acceptance criteria:
- API must be backward compatible
Architecture note:
- remove the old API
The agent should not pick on its own which statement wins.
The operation has a large blast radius
For example:
terraform destroy
production deployment
database migration
permission escalation
Extra permissions are needed
The agent has:
read repository
create branch
create PR
but suddenly needs:
production write access
That should require an explicit decision.
Information is missing and cannot be recovered from the available sources
The agent has checked:
ticket
repo
RAG
tools
and still has no answer.
Only then does it escalate.
Clarification versus approval
It is worth separating two cases.
Clarification
The agent needs information:
Is timeout within the retry scope?
After the answer it can continue.
Approval
The agent knows what to do, but the organization requires sign-off:
Plan includes a production database migration.
Approve execution?
Technically the flow looks similar:
pause
|
v
human action
|
v
resume
but semantically it is something else.
In a larger organization it is worth distinguishing these two event types.
Timeout and escalation
What if nobody answers?
The agent should not hang forever.
You can define a policy:
WAITING_FOR_HUMAN
|
v
24h
|
v
reminder
|
v
48h
|
v
escalation
|
v
7 days
|
v
cancel task
This too can be part of the shared platform.
The agent repo defines:
when I need a human
and the platform can provide:
pause
resume
timeout
notification
escalation
What is configured where
| Element | Where |
|---|---|
| when the agent should escalate | agent repo / workflow |
| checkpoint | runtime / state layer |
| content of the question | agent |
| communication channel | platform / integration layer |
| who may answer | identity / policy |
| resume trigger | event layer |
| timeout | workflow / orchestration |
| audit of answers | observability / state |
The final flow
Our Implementer Agent no longer looks only like this:
ticket
|
v
implement
|
v
PR
The real flow is closer to this:
Jira ticket
|
v
Implementer Agent
|
v
repo + RAG + tools
|
v
implementation
|
v
decision missing?
|-- no
| |
| v
| continue
|
`-- yes
|
v
checkpoint
|
v
WAITING_FOR_HUMAN
|
v
Jira / Teams / Slack
|
v
human response
|
v
resume event
|
v
load checkpoint
|
v
continue
|
v
tests
|
v
PR
So human-in-the-loop does not mean the agent stops being autonomous.
It means the organization draws a clear line:
Which decisions the agent may make on its own, and at which ones it has to stop and ask a human to decide.
Materials
- Beyond Dynamic Workflows, on long waits and approvals outside the LLM
- Anthropic: Building effective agents
- LangGraph: human-in-the-loop and interrupts
- Temporal: signals, resuming a workflow from outside
- AWS Step Functions: wait for a callback with a task token
- Azure Durable Functions: the human interaction pattern
- GitHub Actions: environments with required reviewers