Engineering2026-09-29
Long-Running AI Agents: The Four Async Patterns That Work in Production
Most AI agents fail on tasks that take longer than 30 seconds. Four async patterns — Accept-and-Resume, Checkpoint-Restore, Queue-based execution, and Durable HITL — that fix it in production.
<p class="lead">Most AI agents fail on the tasks that matter most. Not because the model gets the answer wrong — but because the task takes longer than 30 seconds and the whole thing falls over. Long-running AI agents require a different engineering approach: async patterns, durable state, and checkpoints that survive timeouts, crashes, and human approval gates. Here are the four patterns that work in production.</p>
<figure>
<img src="https://images.unsplash.com/photo-1518770660439-4636190af475?w=1200&q=80" alt="Long-running AI agents async patterns for production — four engineering patterns that handle complex multi-step tasks without timeouts, state loss, or blocking" width="1200" height="630" loading="lazy" />
</figure>
<h2>Why Synchronous AI Agents Break on Complex Tasks</h2>
<figure>
<img src="https://images.unsplash.com/photo-1461749280684-dccba630e2f6?w=1200&q=80" alt="Synchronous AI agent timeout failure — 30-second HTTP limits breaking complex multi-step agent workflows before they complete, losing all partial progress" width="1200" height="800" loading="lazy" />
</figure>
<p>The default pattern for most AI agents built on frameworks like n8n, LangChain, or direct API calls is synchronous: a request comes in, the agent processes it, and the response goes back. For tasks that complete in under 30 seconds, this works fine. For anything longer, it creates three compounding failure modes.</p>
<p><strong>HTTP timeouts.</strong> Most web servers, API gateways, and serverless platforms enforce a maximum response time between 30 seconds and five minutes. A contract review agent parsing 40 pages, a proposal writer pulling from five source documents, or a reporting agent querying three separate data sources will hit these limits routinely. When the timeout fires, you get no partial result — just a failure and a frustrated user.</p>
<p><strong>Lost state on failure.</strong> A synchronous agent that fails halfway through a complex task loses everything. The 25 minutes of work it completed is gone. The next attempt starts from scratch, costing the same token budget and clock time all over again.</p>
<p><strong>Blocked threads under load.</strong> When long-running agents occupy server threads, they prevent other requests from processing. A five-person consultancy with three staff submitting reports simultaneously will see the third request queue behind the first two — a classic head-of-line blocking problem that compounds as usage grows.</p>
<blockquote><p>The synchronous agent pattern is sensible for simple tasks. It becomes a reliability tax as your AI operating system grows. Every long-running task you add is a new way for the whole system to fail at the worst moment.</p></blockquote>
<p>The fix is not a faster model or more memory. It is a different architecture. These four patterns address the problem at its root.</p>
<h2>Pattern One — Accept and Resume</h2>
<figure>
<img src="https://images.unsplash.com/photo-1526628953301-3e589a6a8b74?w=1200&q=80" alt="Accept-and-Resume async pattern for AI agents — 202 Accepted response returns task ID immediately while background worker processes the long-running task independently" width="1200" height="800" loading="lazy" />
</figure>
<p>The Accept and Resume pattern is the foundational async fix. Instead of holding a connection open while the agent works, the system accepts the request, returns a task ID immediately, and lets the caller poll for results when ready.</p>
<p>The mechanics are straightforward. When a task arrives — a document to review, a report to generate, a research request — the system returns an HTTP 202 Accepted response with a <code>task_id</code>. The agent runs in the background. The calling system periodically checks a status endpoint (<code>GET /tasks/{task_id}/status</code>) until it receives a completed or failed state. On completion, the result is available at a results endpoint.</p>
<p>This pattern eliminates HTTP timeouts entirely. The 202 response fires immediately — well within any timeout threshold. The actual work happens asynchronously on a background worker that has no connection to hold open and no timeout to hit.</p>
<p>In n8n, this translates to a webhook that triggers a workflow execution and returns the execution ID, with a second workflow that fetches the result when the first completes and routes it to Slack, email, or the requesting system. In a custom Python stack, Celery with Redis provides the queue and the result backend out of the box.</p>
<p>For UK service businesses using agents for tasks like client report generation or contract drafting, this pattern means staff can submit a long-form task and continue working — the result arrives in their inbox or Slack when complete, rather than blocking their screen for minutes at a time. The same principle that makes email more useful than phone calls — you send it and carry on — applies to your AI agents.</p>
<h2>Pattern Two — Checkpoint and Restore</h2>
<figure>
<img src="https://images.unsplash.com/photo-1544197150-b99a580bb7a8?w=1200&q=80" alt="Checkpoint-and-Restore pattern for AI agent state persistence — serialising agent state at workflow boundaries so failed tasks resume from the last stable point rather than starting over" width="1200" height="800" loading="lazy" />
</figure>
<p>Accept and Resume solves the timeout problem. It does not protect you from mid-task failures. If a reporting agent successfully queries two of three data sources and then the third API returns an error, a naive implementation discards the first two results and starts over. Checkpoint and Restore prevents that.</p>
<p>The pattern works by serialising the agent's state at defined checkpoints and writing it to durable storage. If the agent fails after checkpoint three, the next attempt loads from checkpoint three and continues — skipping the work already completed and the token cost that went with it.</p>
<p>Implementation requires three decisions: what constitutes a checkpoint, where state is stored, and how resumption is triggered.</p>
<ul>
<li><strong>Checkpoint placement.</strong> Natural checkpoints are the boundaries between major steps: after data retrieval, after initial analysis, after first draft generation. For a proposal agent — after pulling client brief and past proposals (checkpoint one), after generating the executive summary (checkpoint two), after completing each section (checkpoint three onwards). Checkpoints should be frequent enough that a restart costs no more than one unit of meaningful work.</li>
<li><strong>State storage.</strong> Checkpoint state needs to survive process restarts, so in-memory storage does not work. PostgreSQL, Redis with AOF persistence, or a lightweight SQLite file in object storage are all viable. The state should include the full agent context: what has been retrieved, what has been generated, and what step comes next.</li>
<li><strong>Resumption trigger.</strong> A retry policy in the task queue re-queues failed tasks automatically. On re-queue, the agent checks for an existing checkpoint for that task ID, loads it if present, and continues from the saved step rather than the beginning.</li>
</ul>
<p>The cost of implementing this pattern is real — state serialisation adds engineering overhead. The return is equally real: tasks that previously required full reruns on failure now complete from their last stable point. For agents running expensive multi-step workflows where each step makes costly API calls, the compounding savings on token budget and clock time are significant across a week of production use.</p>
<p>This pattern pairs naturally with the <a href="/blog/ai-agent-fault-tolerance-patterns">fault tolerance patterns</a> covered earlier in this series — particularly circuit breakers, which prevent the agent from repeatedly hammering a failing downstream API during a checkpoint-resume loop.</p>
<h2>Pattern Three — Queue-Based Async Execution</h2>
<figure>
<img src="https://images.unsplash.com/photo-1517694712202-14dd9538aa97?w=1200&q=80" alt="Queue-based async AI agent execution — message queue absorbs task spikes, worker autoscaling handles peak load, dead letter queues surface repeated failures for human review" width="1200" height="800" loading="lazy" />
</figure>
<p>Accept and Resume handles one long-running task at a time. Queue-Based Async Execution handles many — and keeps the system responsive under load.</p>
<p>The core element is a message queue between the task receiver and the task executor. Incoming tasks enter the queue. Workers pull from the queue at a rate they can sustain. The queue absorbs spikes without dropping requests or degrading latency for other concurrent users.</p>
<p>Three properties make this pattern reliable in production:</p>
<ul>
<li><strong>Worker autoscaling.</strong> During peak load (Monday morning report generation, end-of-month client summaries), the number of workers scales up automatically. During quiet periods, it scales to zero, eliminating compute cost entirely. AWS SQS with Lambda, Google Pub/Sub with Cloud Run, or Cloudflare Queues with Workers all support this model cleanly.</li>
<li><strong>Dead letter queues.</strong> Tasks that fail repeatedly — after exhausting their retry budget — move to a dead letter queue for human review rather than silently disappearing. This is the observability surface you need when an agent starts failing on a new class of input that your error handling does not yet cover.</li>
<li><strong>Priority queuing.</strong> Not all tasks are equal. A time-sensitive client approval request should not wait behind 50 routine report jobs. Priority queuing lets you assign weight to different task types — urgent client-facing tasks run immediately, background maintenance tasks run when workers are idle.</li>
</ul>
<p>Benchmark data from production deployments in comparable UK service businesses: queue-based architectures handle four to eight times the task volume of synchronous architectures before showing latency degradation, with failure rates dropping below 2% compared to 12–18% for synchronous agents under moderate concurrent load.</p>
<p>For a UK consultancy running an AI operating system with five or more active agents, the queue becomes the coordination layer between them. The <a href="/blog/event-driven-ai-agents-webhook-architecture">event-driven architecture</a> that routes triggers into agents is the input side; the queue is the execution engine that ensures reliable, ordered, scalable processing regardless of how many tasks arrive at once.</p>
<h2>Pattern Four — Durable State with Human Approval Gates</h2>
<p>The three patterns above address mechanical failures: timeouts, mid-task crashes, and load spikes. This fourth pattern addresses a different class of long-running task — one that is deliberately paused to wait for a human decision.</p>
<p>Many high-value agent tasks in UK service businesses are not fully automated. A proposal agent drafts the document, but the principal signs it off. A contract review agent flags the risk clauses, but the solicitor approves the response. A report agent compiles the data, but the account manager reviews it before it goes to the client. These are not failure scenarios — they are intentional workflow design. The agent handles the repeatable 80%; the professional handles the 20% that requires judgement.</p>
<p>The engineering challenge: the agent should not hold a connection open or continue running while waiting for human input that might arrive in an hour, a day, or three days. Durable State solves this cleanly.</p>
<p>On reaching an approval gate, the agent serialises its complete state to storage, sends a notification — Slack message, email, or webhook to the requesting system — with the pending decision and a unique approval link, and shuts down. No running process, no open connection, no compute cost. When the human responds (approve, reject, or revise), a webhook fires, the agent loads its saved state, applies the decision, and continues from exactly where it paused.</p>
<p>This is the same underlying principle as the HITL approval flows in the <a href="/blog/build-hitl-approval-workflow-n8n">n8n HITL workflow guide</a>. The key engineering addition here is the explicit state serialisation that makes resumption reliable regardless of how long the agent has been paused — an hour or a week. The agent wakes up with the same context it had when it went to sleep.</p>
<blockquote><p>The practical outcome: agents that handle the structured work automatically, pause cleanly for professional judgement, and resume without any manual re-entry or context loss when that judgement is given. The human stays in control of what matters. The agent handles everything else.</p></blockquote>
<h2>Building the Stack</h2>
<p>These four patterns compose. In a production AI operating system for a UK service business, the architecture looks like this: incoming task requests go to a lightweight API that returns a 202 immediately (Pattern One). The task enters a priority queue (Pattern Three). A worker picks it up and checkpoints progress at each major step (Pattern Two). If the task reaches a decision gate, it serialises state and pauses for human approval (Pattern Four). The result arrives via Slack, email, or webhook when complete.</p>
<p>The total infrastructure cost to run this stack for a five-person consultancy — using n8n, Redis, and PostgreSQL on a modest cloud setup — runs between £40 and £80 per month. The reliability improvement over synchronous agents is measurable from the first week: tasks that previously failed 15–20% of the time on complex inputs complete successfully above 97% of the time. Staff stop chasing agents that "got stuck" and start treating them as dependable colleagues.</p>
<p>This is the engineering foundation that separates an AI operating system that genuinely scales from one that performs well in a demo and struggles in production. The agents doing your highest-value work — the ones handling client deliverables, financial documents, and compliance outputs — almost never complete in under 30 seconds. They deserve an architecture that treats that as a first-class design constraint, not an afterthought.</p>
<p>If you are building or scaling an AI operating system for your UK service business and want a practical architecture review, <a href="/contact">get in touch with the Quantum Flow team</a>. We design, build, and run AI systems for UK service businesses under real production conditions — not just in demos.</p>