Skip to main content

What are you looking for?

Explore our services and discover how we can help you achieve your goals

Agentic AI architecture with human-in-the-loop controls: where to put the checkpoint

Place a human checkpoint by how reversible the action is, the latency a business accepts, and the agent's measured confidence: synchronous blocking for irreversible or low-confidence steps, asynchronous queued for the rest, optimistic execution with a rollback only where a compensating action exists. Persist enough state at the pause to resume without replaying a side effect.

Book a solution review See the related solution

By Netbase AI and Automation Engineering Desk · Reviewed by David (CEO) · Updated 6 Oct 2026 · 9 min read

star

This guide is for CTOs and engineering leads designing or reviewing an agent architecture, and for technical product owners deciding how much of a workflow an agent should run alone. Our AI agent vs workflow automation guide decides whether a task needs an agent at all; our AI agent security and human approval controls guide secures one approval gate against injection and non-compliant automated decisions. This page sits between them: once you have an agent and know a gate is needed, it designs where the gate sits in the control loop and how the system survives the wait.

In this guide

The control loop has three places for a person, not one

An agent runs a perceive-decide-act loop: it reads context, proposes a tool call, and (where a gate applies) waits for a decision before it acts. Most teams default to one checkpoint shape — a chat message that blocks until someone answers — because it is the easiest to build first. It is also the slowest at scale and, as our security and approval guide notes, the shape that trains approvers to stop reading. Three shapes exist, and an architecture should choose deliberately among them, not by default:

  • Synchronous blocking. The run halts at the gated step; nothing downstream in that run proceeds until a person answers.
  • Asynchronous queued. The gated action is written to a queue or a durable promise; the run continues other independent work, and the agent (or the next run) picks up the answer when it arrives.
  • Optimistic execution with a compensating action. The step runs immediately, and a pre-built reversal (refund, unsend, revert) is queued; a person reviews after the fact and triggers the reversal if needed.

Choosing the pattern

Pattern How it survives the wait Cost Choose it when
Synchronous blocking Simplest: one paused thread, one resume call Throughput drops to the speed of the slowest approver The action is irreversible, high-value, or confidence is low and it is a new agent in shadow mode
Asynchronous queued Needs a durable checkpoint and a callback or signal id More moving parts: a queue, a timeout, a dedupe key Volume is too high for blocking, and the action can safely wait minutes to hours (most operational approvals)
Optimistic + compensating action Needs a tested reversal path for every optimistic action Risk shifts to "did the reversal actually work" The action is reversible within a known window and the business values speed over pre-approval (routine, low-amount cases)

Building the checkpoint-placement matrix

  1. Classify the action by reversibility

    Read-only, reversible write, or irreversible/external (payment, refund, customer message, deletion, permission change) — the same classification our security guide uses for gating, reused here to choose the pause shape, not just whether to gate.

  2. Set the latency budget

    Ask what the business tolerates: seconds (synchronous), minutes to hours (queued), or none at all because the window for reversal is what matters (optimistic).

  3. Score confidence from real cases

    An agent with a measured accuracy on a case type can move from synchronous to queued, and from queued to optimistic, only after that accuracy is shown on cases like the one in front of it, not assumed from a demo.

  4. Map the combination to a pattern

    Low reversibility or low confidence always routes to synchronous, regardless of latency budget; high reversibility and high confidence is the only combination eligible for optimistic execution.

  5. Add a timeout and an escalation path to every asynchronous and optimistic pattern

    A pause with no upper bound is a leak: if no answer arrives in the budgeted window, escalate to a second approver or fail safe to "do not act," never to "act anyway."

State and checkpoint design: what must survive the pause

A synchronous gate can get away with holding state in a running process, because nothing else happens while it waits. Asynchronous and optimistic patterns cannot: the process may restart, the worker may die, and the wait may last hours. Durable execution platforms and agent frameworks converge on the same answer — persist state outside process memory, keyed by a run or thread id, and resume from exactly where the run paused rather than replaying it. Four fields matter most for an agent checkpoint: the inputs the agent had read so far (including any retrieved context, where the run touches a knowledge base as in our RAG architecture guide), the exact proposed tool call and its parameters, a stable run or thread identifier the approval response is bound to, and an idempotency key, because a resumed run commonly re-enters the paused step from the top and must not repeat a side effect such as a charge or a duplicate message. LangChain's interrupt-and-resume model persists a full state snapshot at the pause and restores it on resume; cloud workflow callback patterns (for example AWS Step Functions-style task tokens) keep the paused step visible and billable-idle rather than holding a live thread; durable-promise patterns such as Restate's awakeables bind the approval to a callback id that survives a process restart. None of this is a product endorsement: Netbase works with the major commercial and open-source AI tools and models, chosen per project, and designs the checkpoint around whichever orchestration layer a project already uses.

Where AI and the checkpoint design meet

What delivery record exists, and what does not

  • What exists. Netbase has delivered AI for clients that are not named, including the moderation and chatbot records above, where an agent's output or action waits on or syncs with a person-facing queue or CRM workflow. Netbase follows its own security practices (secure code review, TLS in transit and AES at rest, role-based access control, MFA for admin dashboards, vulnerability scanning and penetration testing) and holds ISO 27001 certification and a SOC 2 Type II attestation for its own operations; see security and compliance.
  • What does not. No published Netbase record states a resume time, a queue depth, an approval rate, or an incident count for a paused agent run, and no record names which orchestration framework or workflow engine a given delivery used. Netbase's certifications cover Netbase's own operations; they are not extended to a client's system.

Alternatives and selection criteria

Approach Strength Weakness Choose it when
Synchronous blocking everywhere Simplest to build and to explain to an approver Throughput caps at approver speed; trains rubber-stamping over time A new agent in shadow mode, or every action is high-stakes
Asynchronous queued by default Keeps throughput while a person still decides every gated case Needs durable storage, a timeout policy and a dedupe key Most operational agents once volume outgrows synchronous review
Optimistic execution with compensating actions Fastest; no wait on the happy path Only as safe as the reversal path; wrong for anything irreversible Routine, reversible, low-amount actions with a tested rollback
No checkpoint, fixed workflow Nothing to design; smallest surface for state loss No agent flexibility The task is predictable enough that a fixed workflow fits better than an agent

The same four criteria from the matrix above decide between these rows: reversibility of the worst case, the latency the business accepts, measured confidence on real cases, and whether a tested compensating action exists at all.

Limits of this guide

  • It describes architecture patterns, not a specific framework's guarantees; every orchestration layer and workflow engine implements interrupt, callback and durable-promise patterns slightly differently, and the cited sources document their own product's behaviour, not a universal standard.
  • It is general engineering guidance, not legal advice; whether an agent's decision needs human oversight under the EU AI Act or the GDPR is a legal assessment covered in our AI agent security and human approval controls guide, not this one.

Plan the next step with a Netbase consultant

Frequently asked questions

No. Read-only and low-stakes reversible actions can run without one; the matrix in this guide routes only actions with low reversibility, a tight latency budget, or unproven confidence to a checkpoint, and picks the pause shape for the rest.

It is the simplest pattern to build, but it caps throughput at approver speed and, left on every action, trains approvers to approve without reading. Move routine, reversible actions to an asynchronous queue or an optimistic pattern once accuracy is proven on real cases.

With state held only in process memory, the pause is lost. A durable checkpoint — keyed by a run or thread id, holding the proposed action and an idempotency key — lets the run resume from where it paused without replaying the step, whichever orchestration layer wrote the checkpoint.

Next step

Share the agent you are architecting, the actions it proposes and how long a decision can safely wait, and we will book a solution review to place its checkpoints and design the state it needs to resume. You can also see AI automation and agents or more Netbase insights.

AI automation and agents that keep people in charge AI automation and agents that keep people in charge

Netbase provides AI automation and agent development for operations teams that want repetitive, multi-step work done by software while people keep approval over the decisions that matter. We combine rule-based workflow automation with AI steps where they add value, design the human approval points in, and measure return against a baseline taken before the build.

Learn More
line
Enterprise AI agent platform: run AI agents with approvals, audit logs and cost limits Enterprise AI agent platform: run AI agents with approvals, audit logs and cost limits

An enterprise AI agent platform is a governed runtime where AI agents plan and carry out multi-step work through approved tools, while approvals, audit logs, access rules and cost limits keep every action accountable. Netbase offers it as a forward-looking solution, built inside your own environment rather than licensed as a product.

Learn More
line
Contact Netbase

Discuss a project

Netbase JSC helps organizations design, build, modernize, and operate digital products and AI-enabled business systems.
Project enquiries

[email protected]

WhatsApp

+84 937 869 689

Office address

91 Nguyen Chi Thanh, Dong Da, Hanoi, Vietnam

Get in touch

Tell us what you want to build, modernize, or operate.

Tell us what you want to build, modernize, or operate.

Contact Netbase