AI Agent Observability: Logging, Tracing, and Debugging Explained

In this article, you will learn what AI agent observability means, why traditional monitoring tools fall short for agentic systems, and how to implement structured logging, distributed tracing, and practical debugging workflows for AI agents.

Topics we will cover include:

  • Why AI agents fail in ways that look like success, and why that makes standard monitoring tools insufficient.
  • How to implement structured logging and OpenTelemetry-based tracing for agent runs, tool calls, and model inference steps.
  • How to read trace waterfalls, track token costs with metrics, and use that data to debug real agent failures.

AI Agent Observability: Logging, Tracing, and Debugging Explained

An agent handling customer support tickets closes one out with a clean, professional, entirely wrong answer. It called the refund-lookup tool once, then called it again with slightly different arguments a few seconds later, then answered confidently based on the second result instead of the first. Nothing crashed. No error fired. The uptime dashboard shows green the entire time. The only reason anyone finds out is a customer replying two days later, confused, and by then nobody can reconstruct what actually happened inside that run.

That failure is the whole reason this article exists. A traditional service either returns an error 200 or throws something you can grep for. An agent can do neither and still be completely wrong, and the tooling built for the first kind of system is close to blind to the second. This article walks through what actually needs to change — logging, tracing, and debugging — one at a time, with real code.

What AI Agent Observability Actually Means

AI agent observability is the practice of capturing every model call, tool execution, and reasoning step an agent makes as structured data, so that when something goes wrong, you can reconstruct exactly what happened and why, rather than guessing or re-running the same prompt and hoping the problem repeats itself.

It borrows from the three pillars observability engineers already know — logs, metrics, and traces — but the reason it needs its own name and its own discipline comes down to how agents actually fail. Aryan Kargwal, a researcher in the field, put it plainly in coverage from Digital Applied’s 2026 observability guide: agentic systems fail in ways that look like success — incorrect but well-formed outputs, unnecessary tool calls, or actions that are syntactically valid but semantically wrong. None of that trips an error handler. A health check reporting “up” tells you almost nothing useful about whether the agent actually did the right thing on any given run.

Why Agents Break the Traditional Monitoring Model

It’s worth being specific about the mechanics here, because “agents are unpredictable” undersells exactly what changes.

The same input doesn’t reliably produce the same behavior anymore. Temperature settings, retrieval results, and which tools happen to be available can all shift the path an agent takes, so the same prompt can trigger a genuinely different sequence of tool calls on two consecutive runs. A single “it worked when I tested it” trace tells you almost nothing about what the distribution of real runs actually looks like.

Cost and latency stop correlating with request count and start correlating with tokens instead. A single “slow” request might be consuming ten times the normal token budget, and a monitoring setup built around requests-per-second is structurally blind to that. Multi-step chains compound the problem: one user request might trigger several model calls, a handful of tool calls, and a couple of retrieval lookups, and each one is an independent point of failure that a single aggregate error metric can’t distinguish between. And prompts themselves routinely carry real personal or confidential information, which means naive logging that dumps full prompt text into a backend creates a genuine compliance problem before it has created any debugging value at all.

A side-by-side comparison of the two worlds makes the shift concrete:

Signal Traditional app LLM / AI agent
Latency driver CPU, I/O, network Token count, model size, context window
Cost unit Requests per second Tokens consumed
Failure mode Exception, timeout Hallucination, context overflow, tool error
Debug artifact Stack trace Prompt, completion, and the reasoning chain between them

Logging

Start with the most familiar pillar, because it’s still the foundation everything else builds on, just applied differently. For an agent, the events worth logging are specific: which tool got called and with what arguments, what came back, how many tokens a given step consumed, how long each hop took, and any error along the way — and all of it structured rather than written as free-text sentences a human has to parse later.

The detail that actually makes agent logging useful is tying every log line back to the specific run it came from. A log statement that just says “tool call failed” is nearly worthless at 2 am when three different users triggered three different runs in the same minute. Attaching the current trace ID to every log line — something OpenTelemetry does automatically once tracing is set up — is what turns a pile of scattered log statements into something you can filter down to the exact run that broke.

A few things worth noticing in that snippet. trace.get_current_span() doesn’t require you to manually pass a trace ID down through every function call — it reads whatever span is active in the current execution context, which is exactly what makes this pattern practical to sprinkle throughout a real codebase without threading an ID parameter through every layer.

Logging the arguments and the result length, rather than the full result content, is a deliberate choice, not an oversight; full tool outputs can be large and can carry sensitive data, and a length or a truncated preview is usually enough to spot a problem without turning every log line into a privacy liability. And logging both a start and an end event for the same tool call, rather than just the outcome, is what lets you later measure exactly how long that specific call took — which becomes the raw material tracing formalizes properly in the next section.

Tracing

Logging tells you what happened at individual points in time. Tracing is what stitches those points into a shape — a full record of one agent run from the first request to the final answer, with every step nested inside the step that triggered it. That nested shape is the actual answer to “why did the agent do that,” because it shows you not just that a tool was called, but which reasoning step decided to call it and what happened immediately before and after.

The vocabulary here comes from the OpenTelemetry GenAI semantic conventions, which define a standard set of gen_ai.* span types and attributes specifically for this. Rather than every team inventing their own span names, the spec defines a handful of operation types worth knowing: create_agent for when an agent is first defined, invoke_agent for a single agent run, invoke_workflow for orchestration across multiple agents handing off to each other, execute_tool for an individual tool call, and chat for the actual model inference call itself. Each one carries a standard set of attributes — gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, and gen_ai.response.finish_reasons among them — so a trace produced by one team’s agent looks structurally the same as one produced by a completely different framework.

Here’s what manually instrumenting a small tool-calling agent actually looks like:

The nesting is doing the real work here. Every chat span and every execute_tool span opens inside the with tracer.start_as_current_span(…) block belonging to the run above it, which is exactly what OpenTelemetry uses to build the parent-child relationship automatically — you never manually wire “this span belongs under that one,” it’s implicit in how the with blocks are structured in your code. record_exception plus set_status(StatusCode.ERROR, …) on the tool span is what makes a failed tool call show up clearly in a trace viewer rather than silently vanishing into the returned string, which matters directly for the debugging section later in this article. And separating token usage attributes onto the chat span specifically, rather than the top-level invoke_agent span, is what lets a trace viewer later show you token cost broken down per model call within a single run, not just a single combined total for the whole thing.

A screenshot of a trace waterfall view in an observability dashboard

A screenshot of a trace waterfall view in an observability dashboard (click to enlarge)
Image by Author

Chain Visualization: Reading the Trace Waterfall

The spans from the last section don’t mean much as a raw list. What actually makes them useful is a waterfall view — spans stacked by their nesting depth and stretched horizontally by how long each one took, so a whole agent run becomes a single picture you can scan in seconds. It’s worth learning to read one in plain text before ever opening a real dashboard, since the shape is the same either way:

That waterfall alone tells you more than an hour of guessing would. The two refund_lookup calls sitting as siblings under the same chat span is exactly the kind of redundant tool call that caused the failure this article opened with, and it’s visible at a glance rather than buried in a wall of logs. In a real dashboard, this same structure renders as horizontal bars, and the habits worth building are simple: look for a bar that’s unusually wide relative to its siblings, since that’s where time and token budget are actually going; look for a tool call that repeats when it shouldn’t; and look for a chat span where the model reached for a tool at all when the task didn’t obviously need one. None of that requires reading a single log line. It’s all visible in the shape of the trace itself.

Token Tracking and Cost Metrics

Traces are excellent for understanding one specific run in detail. They’re the wrong tool for spotting a trend across thousands of runs, which is what metrics are for. The two genuinely load-bearing metrics for an agent, per the OpenTelemetry GenAI conventions, are gen_ai.client.token.usage and gen_ai.client.operation.duration, tracked as a counter and a histogram respectively.

The choice to record input and output tokens as two separate counter calls, rather than one combined total_tokens value, matters more than it looks. A single combined counter hides exactly the information you’d need to notice — for instance, that a system prompt has quietly grown so large it’s dwarfing the actual user input on every single call. Splitting them apart is what makes that visible in a dashboard rather than buried inside an average.

With those two metrics flowing, a small set of alert conditions covers most of what actually goes wrong in production, per Uptrace’s recommendations:

Metric Alert condition Why it matters
Token usage rate More than double the baseline over 10 minutes Often a runaway loop or a prompt injection attempt
Operation duration, p99 Above 30 seconds The model is overloaded, or the context window is too large
Error rate Above 2% over 5 minutes Rate limiting or quota exhaustion, worth catching before users do
Input-to-output token ratio Consistently above 10 to 1 The system prompt has likely grown bloated and needs trimming

Telemetry Pipelines: Getting Traces from Code to a Backend

Everything so far has assumed traces and metrics land somewhere useful, and that plumbing deserves its own attention rather than being an afterthought. In a typical setup, your instrumented application exports telemetry to an OpenTelemetry Collector — a separate process that receives it, optionally transforms or filters it, and forwards it on to wherever you’re actually storing and viewing traces. That middle layer is where two practical concerns get handled without touching a single line of application code.

  1. The first is sampling. Capturing every single trace at full volume is reasonable in development, but LLM calls are slow and their spans are large, so full capture in production gets expensive fast without adding proportional value. The practical pattern is to sample differently depending on the situation: full capture in development, a modest percentage — often 5 to 10% — of routine successful production calls, and 100% capture for anything genuinely valuable: every error, every high-token request, and every full agent run, since those are exactly the cases you’ll actually want to look back at.
  2. The second is privacy, and it deserves to be treated as a first-class design decision rather than something bolted on later. Prompt and completion content belongs in span events, not span attributes, since attributes are always indexed and exported with no size limit, while events can be filtered, truncated, or dropped entirely at the Collector level. A Collector configuration can strip or hash prompt content from every span crossing the pipeline before it ever reaches a storage backend, which means a compliance requirement doesn’t have to turn into a code change scattered across every instrumented call site in your application.

MCP Tracing

This is a genuinely recent addition worth knowing about specifically, since it closes a gap most existing coverage of this topic hasn’t caught up to yet. Before OpenTelemetry’s Model Context Protocol semantic conventions, added in spec version 1.39, the actual protocol mechanics underneath an MCP tool call — which method got invoked, which session it belonged to, and which protocol version was in use — were effectively invisible in a trace. You could see that a tool ran and what it returned, but not the layer beneath that.

The detail worth understanding is how these new attributes get attached. Rather than creating a second, separate span for the protocol layer, MCP instrumentation enriches the existing execute_tool span with attributes like mcp.method.name, mcp.session.id, and mcp.protocol.version, layering the extra detail onto the span you already had instead of doubling the trace’s noise. For anyone building agents that call out to several MCP servers, that’s the difference between a trace that stays readable and one that turns into a wall of near-duplicate spans.

Debugging Workflows

This is where everything built so far actually pays off. A real debugging workflow, once tracing is in place, tends to follow the same shape regardless of the platform behind it: start from the bad output a user reported, pull the trace ID that produced it, open the waterfall, and scan for the span where the run actually went wrong — an unexpectedly wide chat span, a repeated execute_tool call, an error status somewhere in the tree. Once you’ve found the step, the span’s attributes and events tell you exactly what arguments were passed and what came back, which is usually enough to understand the mistake directly, no re-running required, no guessing.

Two more advanced techniques are worth knowing as this practice matures. Replay, sometimes called time-travel debugging, lets you re-run an agent session with point-in-time precision — effectively restoring the exact state the agent was in at a given step and continuing from there, a capability AgentOps is specifically known for. And a newer pattern worth watching is natural-language trace querying, where instead of manually scanning a waterfall, an engineer can directly ask a platform something like “why did the agent enter this loop” and get an answer generated from analyzing the trace data itself — a capability LangSmith has built directly into its product. Neither replaces the fundamentals covered above. Both are what those fundamentals make possible once a team has enough traces flowing to make asking that kind of question worthwhile.

The Tools Top Teams Actually Use in 2026

Building this yourself with raw OpenTelemetry, as shown throughout this article, works and keeps you vendor-neutral, but most teams eventually reach for a platform to store, visualize, and query the traces they’re producing. The decision genuinely comes down to deployment model before it comes down to features, since that single choice eliminates most of the field on its own, per Digital Applied’s 2026 breakdown.

Self-hosted platforms — Langfuse and Arize Phoenix among them — suit teams with real data residency requirements or a need for tight cost control at scale, at the cost of owning the operational overhead yourself. Managed SDKs — LangSmith and Braintrust — trade that ownership for speed: you add an SDK, the vendor runs the backend and storage, and you typically get evaluation tooling bundled in from day one. Proxy gateways — Helicone being the clearest example — sit your traffic behind a routing layer that logs cost and usage across hundreds of models with close to zero code change, with the tradeoff that the gateway itself becomes a single point of failure worth planning real uptime around.

Platform Deployment model Free tier OpenTelemetry support 2025 to 2026 signal
Langfuse Self-hosted or cloud Free self-hosting Yes Acquired by ClickHouse, January 2026
Arize Phoenix Self-hosted Open source, free Yes, OTLP-native Actively growing open-source project
LangSmith Managed SDK 5,000 traces per month Yes Natural-language trace querying built in
Braintrust Managed SDK 1 million spans per month Yes $80M Series B, February 2026
Helicone Proxy gateway 10,000 requests per month Via gateway Cost tracking across 300+ models
AgentOps SDK Open source, free Yes Known for time-travel replay debugging
Datadog LLM Observability Managed, extends existing APM 40,000 LLM spans per month Yes Bills only LLM spans, not tool or retrieval spans

Putting It Together

Here’s how all of the pieces above combine into one working script, so the concepts in this article don’t stay disconnected. This wires up tracing, token metrics, and structured logging together around a single small agent.

Run this with a Collector configured to export to whichever backend you’ve picked from the table above, and a single call to run_instrumented_agent produces a full trace with nested spans for every tool call, token metrics recorded per model call, and structured log lines carrying the trace ID that ties everything back together — exactly the setup this entire article has been building toward.

Conclusion

An unmonitored agent isn’t a smaller risk than an unmonitored web service; it’s a larger one, precisely because its failures are built to look fine until someone checks closely. The refund lookup called twice, the confident answer built on stale data — none of that trips an alarm on its own. Logging tells you what happened at each step. Tracing shows you how those steps actually connected. Debugging is what turns that record into an answer instead of a guess. Build all three in before an agent is handling anything that actually matters, not after the first customer notices something went wrong.

No comments yet.

Leave a Reply

Machine Learning Mastery is part of Guiding Tech Media, a leading digital media publisher focused on helping people figure out technology. Visit our corporate website to learn more about our mission and team.