In this article, you will learn what AI agent observability means, why traditional monitoring tools fall short for agentic systems, and how to implement structured logging, distributed tracing, and practical debugging workflows for AI agents.
Topics we will cover include:
- Why AI agents fail in ways that look like success, and why that makes standard monitoring tools insufficient.
- How to implement structured logging and OpenTelemetry-based tracing for agent runs, tool calls, and model inference steps.
- How to read trace waterfalls, track token costs with metrics, and use that data to debug real agent failures.

An agent handling customer support tickets closes one out with a clean, professional, entirely wrong answer. It called the refund-lookup tool once, then called it again with slightly different arguments a few seconds later, then answered confidently based on the second result instead of the first. Nothing crashed. No error fired. The uptime dashboard shows green the entire time. The only reason anyone finds out is a customer replying two days later, confused, and by then nobody can reconstruct what actually happened inside that run.
That failure is the whole reason this article exists. A traditional service either returns an error 200 or throws something you can grep for. An agent can do neither and still be completely wrong, and the tooling built for the first kind of system is close to blind to the second. This article walks through what actually needs to change — logging, tracing, and debugging — one at a time, with real code.
What AI Agent Observability Actually Means
AI agent observability is the practice of capturing every model call, tool execution, and reasoning step an agent makes as structured data, so that when something goes wrong, you can reconstruct exactly what happened and why, rather than guessing or re-running the same prompt and hoping the problem repeats itself.
It borrows from the three pillars observability engineers already know — logs, metrics, and traces — but the reason it needs its own name and its own discipline comes down to how agents actually fail. Aryan Kargwal, a researcher in the field, put it plainly in coverage from Digital Applied’s 2026 observability guide: agentic systems fail in ways that look like success — incorrect but well-formed outputs, unnecessary tool calls, or actions that are syntactically valid but semantically wrong. None of that trips an error handler. A health check reporting “up” tells you almost nothing useful about whether the agent actually did the right thing on any given run.
Why Agents Break the Traditional Monitoring Model
It’s worth being specific about the mechanics here, because “agents are unpredictable” undersells exactly what changes.
The same input doesn’t reliably produce the same behavior anymore. Temperature settings, retrieval results, and which tools happen to be available can all shift the path an agent takes, so the same prompt can trigger a genuinely different sequence of tool calls on two consecutive runs. A single “it worked when I tested it” trace tells you almost nothing about what the distribution of real runs actually looks like.
Cost and latency stop correlating with request count and start correlating with tokens instead. A single “slow” request might be consuming ten times the normal token budget, and a monitoring setup built around requests-per-second is structurally blind to that. Multi-step chains compound the problem: one user request might trigger several model calls, a handful of tool calls, and a couple of retrieval lookups, and each one is an independent point of failure that a single aggregate error metric can’t distinguish between. And prompts themselves routinely carry real personal or confidential information, which means naive logging that dumps full prompt text into a backend creates a genuine compliance problem before it has created any debugging value at all.
A side-by-side comparison of the two worlds makes the shift concrete:
| Signal | Traditional app | LLM / AI agent |
|---|---|---|
| Latency driver | CPU, I/O, network | Token count, model size, context window |
| Cost unit | Requests per second | Tokens consumed |
| Failure mode | Exception, timeout | Hallucination, context overflow, tool error |
| Debug artifact | Stack trace | Prompt, completion, and the reasoning chain between them |
Logging
Start with the most familiar pillar, because it’s still the foundation everything else builds on, just applied differently. For an agent, the events worth logging are specific: which tool got called and with what arguments, what came back, how many tokens a given step consumed, how long each hop took, and any error along the way — and all of it structured rather than written as free-text sentences a human has to parse later.
The detail that actually makes agent logging useful is tying every log line back to the specific run it came from. A log statement that just says “tool call failed” is nearly worthless at 2 am when three different users triggered three different runs in the same minute. Attaching the current trace ID to every log line — something OpenTelemetry does automatically once tracing is set up — is what turns a pile of scattered log statements into something you can filter down to the exact run that broke.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 |
import logging from opentelemetry import trace # Standard Python logging, nothing exotic here logger = logging.getLogger("agent") logging.basicConfig(level=logging.INFO) tracer = trace.get_tracer("agent-service") def call_tool(tool_name: str, arguments: dict): # get_current_span() pulls whatever span is active right now, # so the log line below can be tied back to the exact trace # and step it happened inside span = trace.get_current_span() trace_id = format(span.get_span_context().trace_id, "032x") logger.info( "tool_call_started", extra={ "trace_id": trace_id, "tool_name": tool_name, "arguments": arguments, }, ) try: result = execute_tool(tool_name, arguments) logger.info( "tool_call_succeeded", extra={"trace_id": trace_id, "tool_name": tool_name, "result_length": len(str(result))}, ) return result except Exception as e: logger.error( "tool_call_failed", extra={"trace_id": trace_id, "tool_name": tool_name, "error": str(e)}, ) raise |
A few things worth noticing in that snippet. trace.get_current_span() doesn’t require you to manually pass a trace ID down through every function call — it reads whatever span is active in the current execution context, which is exactly what makes this pattern practical to sprinkle throughout a real codebase without threading an ID parameter through every layer.
Logging the arguments and the result length, rather than the full result content, is a deliberate choice, not an oversight; full tool outputs can be large and can carry sensitive data, and a length or a truncated preview is usually enough to spot a problem without turning every log line into a privacy liability. And logging both a start and an end event for the same tool call, rather than just the outcome, is what lets you later measure exactly how long that specific call took — which becomes the raw material tracing formalizes properly in the next section.
Tracing
Logging tells you what happened at individual points in time. Tracing is what stitches those points into a shape — a full record of one agent run from the first request to the final answer, with every step nested inside the step that triggered it. That nested shape is the actual answer to “why did the agent do that,” because it shows you not just that a tool was called, but which reasoning step decided to call it and what happened immediately before and after.
The vocabulary here comes from the OpenTelemetry GenAI semantic conventions, which define a standard set of gen_ai.* span types and attributes specifically for this. Rather than every team inventing their own span names, the spec defines a handful of operation types worth knowing: create_agent for when an agent is first defined, invoke_agent for a single agent run, invoke_workflow for orchestration across multiple agents handing off to each other, execute_tool for an individual tool call, and chat for the actual model inference call itself. Each one carries a standard set of attributes — gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, and gen_ai.response.finish_reasons among them — so a trace produced by one team’s agent looks structurally the same as one produced by a completely different framework.
Here’s what manually instrumenting a small tool-calling agent actually looks like:
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 |
from opentelemetry import trace from opentelemetry.trace import Status, StatusCode tracer = trace.get_tracer("agent-service") def run_agent(task: str) -> str: # The root span for this entire run; every step below nests # inside it, which is what produces the parent-child tree with tracer.start_as_current_span("invoke_agent") as agent_span: agent_span.set_attributes({ "gen_ai.system": "openai", "agent.name": "support-agent", "gen_ai.request.model": "gpt-4o", }) messages = [ {"role": "system", "content": "You are a support assistant."}, {"role": "user", "content": task}, ] while True: # The model call itself gets its own child span with tracer.start_as_current_span("chat") as chat_span: response = model_client.chat.completions.create( model="gpt-4o", messages=messages, tools=AVAILABLE_TOOLS ) choice = response.choices[0] chat_span.set_attributes({ "gen_ai.response.model": response.model, "gen_ai.usage.input_tokens": response.usage.prompt_tokens, "gen_ai.usage.output_tokens": response.usage.completion_tokens, }) if choice.finish_reason != "tool_calls": agent_span.set_status(Status(StatusCode.OK)) return choice.message.content # Each tool call gets its own child span, nested under # the agent run, not under the chat span, since a tool # call is a sibling step, not a sub-step of inference for tool_call in choice.message.tool_calls: with tracer.start_as_current_span("execute_tool") as tool_span: tool_span.set_attributes({ "gen_ai.tool.name": tool_call.function.name, "gen_ai.tool.call.id": tool_call.id, }) try: result = call_tool(tool_call.function.name, tool_call.function.arguments) except Exception as e: tool_span.record_exception(e) tool_span.set_status(Status(StatusCode.ERROR, str(e))) raise messages.append({ "role": "tool", "content": str(result), "tool_call_id": tool_call.id, }) |
The nesting is doing the real work here. Every chat span and every execute_tool span opens inside the with tracer.start_as_current_span(…) block belonging to the run above it, which is exactly what OpenTelemetry uses to build the parent-child relationship automatically — you never manually wire “this span belongs under that one,” it’s implicit in how the with blocks are structured in your code. record_exception plus set_status(StatusCode.ERROR, …) on the tool span is what makes a failed tool call show up clearly in a trace viewer rather than silently vanishing into the returned string, which matters directly for the debugging section later in this article. And separating token usage attributes onto the chat span specifically, rather than the top-level invoke_agent span, is what lets a trace viewer later show you token cost broken down per model call within a single run, not just a single combined total for the whole thing.

A screenshot of a trace waterfall view in an observability dashboard (click to enlarge)
Image by Author
Chain Visualization: Reading the Trace Waterfall
The spans from the last section don’t mean much as a raw list. What actually makes them useful is a waterfall view — spans stacked by their nesting depth and stretched horizontally by how long each one took, so a whole agent run becomes a single picture you can scan in seconds. It’s worth learning to read one in plain text before ever opening a real dashboard, since the shape is the same either way:
|
1 2 3 4 5 |
invoke_agent (1850ms) ├── chat (920ms) ← decides to call two tools │ ├── execute_tool: refund_lookup (310ms) │ └── execute_tool: refund_lookup (295ms) ← called again, same tool └── chat (530ms) ← final answer |
That waterfall alone tells you more than an hour of guessing would. The two refund_lookup calls sitting as siblings under the same chat span is exactly the kind of redundant tool call that caused the failure this article opened with, and it’s visible at a glance rather than buried in a wall of logs. In a real dashboard, this same structure renders as horizontal bars, and the habits worth building are simple: look for a bar that’s unusually wide relative to its siblings, since that’s where time and token budget are actually going; look for a tool call that repeats when it shouldn’t; and look for a chat span where the model reached for a tool at all when the task didn’t obviously need one. None of that requires reading a single log line. It’s all visible in the shape of the trace itself.
Token Tracking and Cost Metrics
Traces are excellent for understanding one specific run in detail. They’re the wrong tool for spotting a trend across thousands of runs, which is what metrics are for. The two genuinely load-bearing metrics for an agent, per the OpenTelemetry GenAI conventions, are gen_ai.client.token.usage and gen_ai.client.operation.duration, tracked as a counter and a histogram respectively.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 |
from opentelemetry import metrics import time meter = metrics.get_meter("agent-service") token_counter = meter.create_counter( "gen_ai.client.token.usage", unit="token", description="Tokens consumed, broken down by model and input/output", ) duration_histogram = meter.create_histogram( "gen_ai.client.operation.duration", unit="s", description="Duration of each model call, in seconds", ) def tracked_chat_call(messages: list, model: str = "gpt-4o") -> str: attrs = {"gen_ai.system": "openai", "gen_ai.request.model": model} start = time.time() try: response = model_client.chat.completions.create(model=model, messages=messages) # Recording input and output tokens as two separate calls, # not one combined total, is what preserves the actual # cost structure, since input and output tokens are # priced differently on nearly every provider token_counter.add(response.usage.prompt_tokens, {**attrs, "gen_ai.token.type": "input"}) token_counter.add(response.usage.completion_tokens, {**attrs, "gen_ai.token.type": "output"}) return response.choices[0].message.content finally: duration_histogram.record(time.time() - start, attrs) |
The choice to record input and output tokens as two separate counter calls, rather than one combined total_tokens value, matters more than it looks. A single combined counter hides exactly the information you’d need to notice — for instance, that a system prompt has quietly grown so large it’s dwarfing the actual user input on every single call. Splitting them apart is what makes that visible in a dashboard rather than buried inside an average.
With those two metrics flowing, a small set of alert conditions covers most of what actually goes wrong in production, per Uptrace’s recommendations:
| Metric | Alert condition | Why it matters |
|---|---|---|
| Token usage rate | More than double the baseline over 10 minutes | Often a runaway loop or a prompt injection attempt |
| Operation duration, p99 | Above 30 seconds | The model is overloaded, or the context window is too large |
| Error rate | Above 2% over 5 minutes | Rate limiting or quota exhaustion, worth catching before users do |
| Input-to-output token ratio | Consistently above 10 to 1 | The system prompt has likely grown bloated and needs trimming |
Telemetry Pipelines: Getting Traces from Code to a Backend
Everything so far has assumed traces and metrics land somewhere useful, and that plumbing deserves its own attention rather than being an afterthought. In a typical setup, your instrumented application exports telemetry to an OpenTelemetry Collector — a separate process that receives it, optionally transforms or filters it, and forwards it on to wherever you’re actually storing and viewing traces. That middle layer is where two practical concerns get handled without touching a single line of application code.
- The first is sampling. Capturing every single trace at full volume is reasonable in development, but LLM calls are slow and their spans are large, so full capture in production gets expensive fast without adding proportional value. The practical pattern is to sample differently depending on the situation: full capture in development, a modest percentage — often 5 to 10% — of routine successful production calls, and 100% capture for anything genuinely valuable: every error, every high-token request, and every full agent run, since those are exactly the cases you’ll actually want to look back at.
- The second is privacy, and it deserves to be treated as a first-class design decision rather than something bolted on later. Prompt and completion content belongs in span events, not span attributes, since attributes are always indexed and exported with no size limit, while events can be filtered, truncated, or dropped entirely at the Collector level. A Collector configuration can strip or hash prompt content from every span crossing the pipeline before it ever reaches a storage backend, which means a compliance requirement doesn’t have to turn into a code change scattered across every instrumented call site in your application.
|
1 2 3 4 5 6 7 8 9 10 |
processors: transform: trace_statements: - context: spanevent statements: # Strips prompt and completion content from every span # event that passes through the Collector, no application # code changes required - delete_matching_keys(attributes, "gen_ai.prompt.content") - delete_matching_keys(attributes, "gen_ai.completion.content") |
MCP Tracing
This is a genuinely recent addition worth knowing about specifically, since it closes a gap most existing coverage of this topic hasn’t caught up to yet. Before OpenTelemetry’s Model Context Protocol semantic conventions, added in spec version 1.39, the actual protocol mechanics underneath an MCP tool call — which method got invoked, which session it belonged to, and which protocol version was in use — were effectively invisible in a trace. You could see that a tool ran and what it returned, but not the layer beneath that.
The detail worth understanding is how these new attributes get attached. Rather than creating a second, separate span for the protocol layer, MCP instrumentation enriches the existing execute_tool span with attributes like mcp.method.name, mcp.session.id, and mcp.protocol.version, layering the extra detail onto the span you already had instead of doubling the trace’s noise. For anyone building agents that call out to several MCP servers, that’s the difference between a trace that stays readable and one that turns into a wall of near-duplicate spans.
Debugging Workflows
This is where everything built so far actually pays off. A real debugging workflow, once tracing is in place, tends to follow the same shape regardless of the platform behind it: start from the bad output a user reported, pull the trace ID that produced it, open the waterfall, and scan for the span where the run actually went wrong — an unexpectedly wide chat span, a repeated execute_tool call, an error status somewhere in the tree. Once you’ve found the step, the span’s attributes and events tell you exactly what arguments were passed and what came back, which is usually enough to understand the mistake directly, no re-running required, no guessing.
Two more advanced techniques are worth knowing as this practice matures. Replay, sometimes called time-travel debugging, lets you re-run an agent session with point-in-time precision — effectively restoring the exact state the agent was in at a given step and continuing from there, a capability AgentOps is specifically known for. And a newer pattern worth watching is natural-language trace querying, where instead of manually scanning a waterfall, an engineer can directly ask a platform something like “why did the agent enter this loop” and get an answer generated from analyzing the trace data itself — a capability LangSmith has built directly into its product. Neither replaces the fundamentals covered above. Both are what those fundamentals make possible once a team has enough traces flowing to make asking that kind of question worthwhile.
The Tools Top Teams Actually Use in 2026
Building this yourself with raw OpenTelemetry, as shown throughout this article, works and keeps you vendor-neutral, but most teams eventually reach for a platform to store, visualize, and query the traces they’re producing. The decision genuinely comes down to deployment model before it comes down to features, since that single choice eliminates most of the field on its own, per Digital Applied’s 2026 breakdown.
Self-hosted platforms — Langfuse and Arize Phoenix among them — suit teams with real data residency requirements or a need for tight cost control at scale, at the cost of owning the operational overhead yourself. Managed SDKs — LangSmith and Braintrust — trade that ownership for speed: you add an SDK, the vendor runs the backend and storage, and you typically get evaluation tooling bundled in from day one. Proxy gateways — Helicone being the clearest example — sit your traffic behind a routing layer that logs cost and usage across hundreds of models with close to zero code change, with the tradeoff that the gateway itself becomes a single point of failure worth planning real uptime around.
| Platform | Deployment model | Free tier | OpenTelemetry support | 2025 to 2026 signal |
|---|---|---|---|---|
| Langfuse | Self-hosted or cloud | Free self-hosting | Yes | Acquired by ClickHouse, January 2026 |
| Arize Phoenix | Self-hosted | Open source, free | Yes, OTLP-native | Actively growing open-source project |
| LangSmith | Managed SDK | 5,000 traces per month | Yes | Natural-language trace querying built in |
| Braintrust | Managed SDK | 1 million spans per month | Yes | $80M Series B, February 2026 |
| Helicone | Proxy gateway | 10,000 requests per month | Via gateway | Cost tracking across 300+ models |
| AgentOps | SDK | Open source, free | Yes | Known for time-travel replay debugging |
| Datadog LLM Observability | Managed, extends existing APM | 40,000 LLM spans per month | Yes | Bills only LLM spans, not tool or retrieval spans |
Putting It Together
Here’s how all of the pieces above combine into one working script, so the concepts in this article don’t stay disconnected. This wires up tracing, token metrics, and structured logging together around a single small agent.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 |
import logging import time from opentelemetry import trace, metrics from opentelemetry.trace import Status, StatusCode tracer = trace.get_tracer("agent-service") meter = metrics.get_meter("agent-service") logger = logging.getLogger("agent") token_counter = meter.create_counter("gen_ai.client.token.usage", unit="token") duration_histogram = meter.create_histogram("gen_ai.client.operation.duration", unit="s") def run_instrumented_agent(task: str) -> str: with tracer.start_as_current_span("invoke_agent") as agent_span: trace_id = format(agent_span.get_span_context().trace_id, "032x") agent_span.set_attributes({"agent.name": "support-agent", "gen_ai.request.model": "gpt-4o"}) logger.info("agent_run_started", extra={"trace_id": trace_id, "task": task[:100]}) messages = [{"role": "user", "content": task}] while True: with tracer.start_as_current_span("chat") as chat_span: start = time.time() response = model_client.chat.completions.create( model="gpt-4o", messages=messages, tools=AVAILABLE_TOOLS ) duration_histogram.record(time.time() - start, {"gen_ai.request.model": "gpt-4o"}) usage = response.usage token_counter.add(usage.prompt_tokens, {"gen_ai.token.type": "input"}) token_counter.add(usage.completion_tokens, {"gen_ai.token.type": "output"}) chat_span.set_attributes({ "gen_ai.usage.input_tokens": usage.prompt_tokens, "gen_ai.usage.output_tokens": usage.completion_tokens, }) choice = response.choices[0] if choice.finish_reason != "tool_calls": agent_span.set_status(Status(StatusCode.OK)) logger.info("agent_run_completed", extra={"trace_id": trace_id}) return choice.message.content for tool_call in choice.message.tool_calls: with tracer.start_as_current_span("execute_tool") as tool_span: tool_span.set_attribute("gen_ai.tool.name", tool_call.function.name) logger.info( "tool_call_started", extra={"trace_id": trace_id, "tool_name": tool_call.function.name}, ) try: result = call_tool(tool_call.function.name, tool_call.function.arguments) except Exception as e: tool_span.record_exception(e) tool_span.set_status(Status(StatusCode.ERROR, str(e))) logger.error( "tool_call_failed", extra={"trace_id": trace_id, "tool_name": tool_call.function.name}, ) raise messages.append({"role": "tool", "content": str(result), "tool_call_id": tool_call.id}) |
Run this with a Collector configured to export to whichever backend you’ve picked from the table above, and a single call to run_instrumented_agent produces a full trace with nested spans for every tool call, token metrics recorded per model call, and structured log lines carrying the trace ID that ties everything back together — exactly the setup this entire article has been building toward.
Conclusion
An unmonitored agent isn’t a smaller risk than an unmonitored web service; it’s a larger one, precisely because its failures are built to look fine until someone checks closely. The refund lookup called twice, the confident answer built on stale data — none of that trips an alarm on its own. Logging tells you what happened at each step. Tracing shows you how those steps actually connected. Debugging is what turns that record into an answer instead of a guess. Build all three in before an agent is handling anything that actually matters, not after the first customer notices something went wrong.






No comments yet.