Gable Blog | AI Agent Observability: Traces, Metrics, and Evals
.avif)
Get the ultimate guide to Data Contracts Deep Dive
Get the ultimate guide to Data Contracts as Code
Discover where your data really comes from.
Ultimate Guide to Data Contracts
What AI agent observability actually is
Agent observability provides step-by-step visibility into an agent's execution. It records which tools the agent called, what data it retrieved, where its reasoning held together, and where it diverged from the intended path. The standard observability vocabulary, drawn from the broader practice, is MELT data: metrics, events, logs, and traces. Agent observability uses that same foundation and adds signals unique to LLM-driven systems, such as token usage, tool interactions, and the agent's decision path.
It helps to set the expectation honestly up front: observability is a visibility tool, not a reliability mechanism. It tells you what an agent did after it did it.
Why agents need more than traditional monitoring
Traditional application performance monitoring captures request-response cycles. It shows that a request came in, a response went out, and how long it took. For a deterministic service, that's often enough. For an agent, it shows the wrapper and misses everything that matters inside.
Consider a support agent that invokes a billing tool with a malformed argument, loops while trying to recover, and then returns a confident but incorrect answer about a refund. Standard monitoring records a successful request: a response was returned in two seconds. Only step-level tracing reveals the hallucinated tool parameter, the retry loop, and the reasoning step where the agent committed to a wrong conclusion.
The three pillars: traces, metrics, and evals
Most agent observability practice organizes around three kinds of signal. Each answers a different question about the agent.
Traces and spans
A trace is the full execution tree of a single agent run, broken into spans. Each span represents one unit of work: an LLM call, a tool invocation, a retrieval step.
Metrics
Metrics quantify agent behavior over many runs. The agent-specific ones that matter most are token usage and the cost it drives, latency per step, and error rates.
Evaluations
Evaluations measure whether the agent is doing a good job, not just what it did.
What to instrument: agent observability best practices
Capturing the right signals is what separates a usable trace from noise. Effective instrumentation captures a consistent set of things at every step:
- LLM calls: the model used, the inputs and prompt version, the output completion, and the input and output token counts.
- Tool calls: which tool was selected, the arguments passed, the result returned, and how long the call took.
- Retrieval steps: the queries sent to a vector store or knowledge base, the documents returned, and any relevance signals available.
- Reasoning transitions: how the agent decided to move from one step to the next, including intermediate reasoning where it's exposed.
- State changes: for stateful agents, what memory was read and written, and how that state shaped later decisions.
The blind spot every observability setup shares
Everything to this point watches the agent. Its calls, its reasoning, its outputs. None of it watches the data the agent depends on, and that's where a large share of real-world agent failures actually originate.
Catching failures before the agent ever runs
Closing that gap means moving enforcement upstream, to the point where the data is produced. That's the role of data contracts: enforceable agreements on schema, semantics, ownership, and constraints between the data producers who generate data and the systems, including agents, that consume it.
Reliable agents need reliable inputs
Visibility into an agent is necessary, and the traces, metrics, and evals above are how you get it. But that visibility is bounded. It ends at the agent's inputs, and a meaningful share of what gets logged as an agent failure is really an upstream data failure wearing an agent costume: the reasoning was sound, the tools fired correctly, and the answer was still wrong because the data was wrong.