Konrad Kowalski (rootsher)Principal Platform & Reliability Architect011010000000000101010011000001101011010100101110

Beyond Dynamic Workflows: when the workflow leaves Claude

date
category
AI Agents
also in
Automation · AI Engineering · Engineering Practices
reading
2 min / 461 words

Dynamic Workflows solve a very specific problem:

text
analyze
-> implement
-> review
-> fix
-> verify

Instead of assuming Claude will remember the whole order of steps, you write it as an executable workflow.

This works well as long as most of the work happens inside Claude's world.

The problem starts when the workflow grows beyond the agent.

For example:

text
GitHub event
-> Claude analyzes the problem
-> run backend job
-> wait for result
-> call API
-> wait for approval
-> store state
-> run another agent
-> create PR

Claude can still perform intelligent steps.

But it should not be responsible for supervising the whole system.

Claude does not need to own the whole workflow

This is an important change.

So far we had:

text
Dynamic Workflow
-> Claude tasks
-> subagents
-> verification

Now we can have:

text
external orchestrator
-> workflow
   +-- Claude task
   +-- backend job
   +-- API call
   +-- wait
   +-- human approval
   +-- another Claude task

Claude becomes one of the executors.

Not the orchestrator of the whole infrastructure.

Why an external orchestrator?

Imagine a workflow:

text
PR opened
-> Claude review
-> if blocking issue: comment
-> if everything OK: wait for CI
-> CI finished
-> if fail: run investigation
-> if pass: wait for approval
-> after approval: next step

This introduces problems you do not want to keep in a context window or a single agent session:

  • retries,
  • timeouts,
  • long waits,
  • webhooks,
  • queues,
  • concurrency,
  • workflow state,
  • human approval,
  • recovery after failure.

Tools for this include:

  • Trigger.dev,
  • Hatchet,
  • Temporal,
  • Inngest,
  • Restate.

Trigger.dev offers durable tasks, queues, retries, checkpointing, long waits and human-in-the-loop.

Hatchet has a similar model based on durable tasks, workers and workflows with retries and checkpointing.

Temporal goes even further toward Durable Execution: the workflow keeps state and can resume after a crash, outage or even days of waiting.

Inngest describes durable workflows started by an event, schedule, webhook or another function, with steps whose results can be saved and reused after retry.

Dynamic Workflow vs external orchestration

The simplest mental model:

text
Dynamic Workflow
= how Claude organizes agent work

External orchestrator
= how the whole system organizes Claude and everything around it

Example:

text
Trigger.dev
-> task: fetch PR
-> task: start Claude review
-> task: wait for CI
-> task: request approval
-> task: call deployment API

Claude can still perform:

text
review
investigation
reasoning
fix generation

But retrying an API call or waiting three hours for approval should not be the LLM's problem.

When to stay with Dynamic Workflows?

If you have:

text
research
-> implementation
-> review
-> fix
-> verify

and everything happens inside one agent's work, do not complicate the system.

Dynamic Workflow is enough.

External orchestration starts making sense when you notice:

More and more steps of my workflow are not Claude work.

For example:

text
wait
webhook
database update
queue
approval
external job
retry

Then adding more instructions to the agent is the wrong direction.

Example: failing CI

In the Claude-first version:

text
Dynamic Workflow
-> investigate
-> fix
-> test
-> review

In a larger system:

text
CI event
-> Trigger.dev / Hatchet / Temporal
-> start Claude investigation
-> result
-> run deterministic CI job
-> pass?
   no  -> Claude fix
   yes -> wait for approval
          -> create PR

Notice the difference.

Claude makes decisions where reasoning is needed.

The orchestrator supervises the rest.

Implement this today

Do not install an orchestrator just because it exists.

Take the workflow from the Dynamic Workflows article and mark each step as:

text
CLAUDE

or:

text
SYSTEM

For example:

text
analyze failure       -> CLAUDE
find root cause       -> CLAUDE
run CI job            -> SYSTEM
wait for webhook      -> SYSTEM
review implementation -> CLAUDE
wait for approval     -> SYSTEM
create PR             -> SYSTEM

If almost everything is CLAUDE, stay with Dynamic Workflows.

If a large part becomes SYSTEM, look at tools such as Trigger.dev, Hatchet, Temporal, Inngest or Restate.

Level complete

The level is complete if you stop treating Claude as the owner of the whole orchestration and can separate:

text
reasoning

from:

text
durable orchestration

Claude should do what it is good at.

The system should supervise things that should be deterministic.

Materials