Synchronous vs. Asynchronous Agent Execution: Architecture Patterns for Production

In this article, you will learn how synchronous and asynchronous execution patterns differ architecturally, and how to choose between them when deploying LLM-based agents to production.

Topics we will cover include:

  • How the synchronous “wait and see” execution pattern works, and the production scenarios where it is appropriate.
  • How the asynchronous, event-driven “fire and forget” pattern decouples task submission from task completion to handle long-running agent workflows.
  • A practical guide to choosing between the two patterns based on task complexity, latency tolerance, and infrastructure requirements.

Synchronous vs. Asynchronous Agent Execution: Architecture Patterns for Production

Introduction

Building local Python scripts where an LLM-based agent loops through several tools is increasingly easy nowadays, with endless high-level libraries and supporting frameworks on the rise. However, there is a severe “deployment gap” in the agentic AI space that deserves further attention.

From a production standpoint, real-world agent workflows entail intricate dependencies, multi-step reasoning loops, and API latency. Failing to design an architecture that accounts for this reality results in system timeouts, dropped user requests, and memory leaks. Two core architectural patterns that bridge these gaps and move toward standard distributed system paradigms are synchronous and asynchronous execution. This article examines these two patterns from an agent execution perspective, illustrating their rationale through two easy-to-run, lightweight notebooks whose underlying logic can be readily translated to production-ready agent-based applications.

Synchronous Agent Execution: “Wait and See”

The synchronous agent execution pattern closely resembles the classic HTTP request-response approach. It is best suited to scenarios that demand immediate feedback, sequential dependencies between tasks, or efficient data retrieval, e.g. in standard RAG pipelines.

It works as follows: an external caller submits a prompt, then the execution thread is blocked while the agent finalizes the whole chain-of-thought process and tool usage.

While simple and effective, this pattern is limited and remarkably fragile when combined with advanced agents. Cloud infrastructure providers like AWS have a default 29-second timeout in their API Gateway. For an agent that takes 45 seconds to plan, search the web, and produce a response, a standard synchronous pipeline would drop the connection, resulting in wasted tokens and lost progress.

The following code shows a runnable example in Google Colab that emulates the synchronous agent execution pattern. It first defines a function, mock_llm_call(prompt), that simulates an agent’s call to an LLM without using an actual LLM. It takes a prompt as input and simulates a small network latency, similar to invoking a real LLM. After that, simple conditional logic checks whether the prompt includes the word “search.” If it does, the function simulates a tool call for performing a web search. Otherwise, it simulates the LLM arriving at a final answer, meaning the model has processed the request and generated a direct response. In both cases, the function returns a dictionary describing an action and its associated metadata.

Meanwhile, an overarching agent function that calls the previous one is also needed. We call this function synchronous_agent(query). It simulates the behavior of an agent that processes a query in several steps. At each step, the agent calls mock_llm_call(). If the mock LLM signals a tool call, the agent simulates the tool execution and adjusts the query for the next step. If the LLM returns a final response, the agent finalizes execution and returns that response as the final result.

Let’s try it out:

Result:

We have just seen how a synchronous loop works in practice. The code above simulated an agent undertaking a reasoning step, calling a tool, and returning the final answer in the subsequent step.

Asynchronous, Event-Driven Execution: “Fire and Forget”

Shifting to an asynchronous execution pattern becomes necessary when addressing complex tasks, such as refactoring codebases, multi-agent debating, or workflows that require frequent HITL (Human-In-The-Loop) approvals.

Under this pattern, task submission and task completion are fully decoupled. Once a client triggers the agent’s execution of a task, the system creates and returns a job_id, moving the task into a queue. Background workers then retrieve tasks from the queue and process them independently. This way, an agent can run for hours without blocking the user interface. Task states are checkpointed frequently to detect and recover from issues like node crashes, allowing the agent to resume where it left off.

The disadvantage of this pattern is the need for a more robust infrastructure, including message brokers like RabbitMQ and Redis, as well as suitable databases for modeling worker node states: PostgreSQL, MongoDB.

Let’s move on to the code example to understand the asynchronous pattern. A key building block here is Python’s asyncio library, whereby a “client” sends a task, receives an acknowledgment, and the task is processed in the background: ideal for time-consuming tasks that would otherwise block the user interface in a deployed application.

The agent_worker() function acts like a background worker node that processes tasks from a queue asynchronously, simulating long-running working steps and updating states in a database.

Next, we have submit_job(prompt). This function acts like an API that receives a task, assigns it an ID, and adds it to the queue, returning the task ID to the client immediately without waiting for the task to complete.

The last function, main(), is responsible for orchestrating the whole agent-based system: initiating the worker node in the background, simulating a task submission by a client, and periodically monitoring the task state until completion.

Result:

Conclusion: When to Use Which

As a general rule, start with a lightweight, synchronous architecture when facing fast, immediate tasks like question-answering. As your agent application grows more capable, you will typically need to transition to asynchronous execution, since the tasks it handles will take longer to process. Asynchronous systems place a higher demand on infrastructure — databases, queues, message brokers — but they are far more resilient against brittle timeouts, and that resilience is ultimately the key to scaling complex agent workflows into production.

No comments yet.

Leave a Reply

Machine Learning Mastery is part of Guiding Tech Media, a leading digital media publisher focused on helping people figure out technology. Visit our corporate website to learn more about our mission and team.