How to Add Tracing and Observability to AI Agents: A Complete Implementation Guide

You can add tracing and observability to AI agents by instrumenting your code with OpenTelemetry, exporting spans to backends like Jaeger, and exposing metrics via Prometheus, enabling full visibility into LLM calls, tool executions, and latency bottlenecks.

Adding tracing and observability to AI agents is essential for monitoring production workloads, diagnosing failures, and optimizing latency in LLM-driven applications. This guide demonstrates how to implement comprehensive observability in the Shubhamsaboo/awesome-llm-apps repository using OpenTelemetry, covering everything from automatic instrumentation to custom span creation.

Why Observability Matters for AI Agents

AI agents interact with users, external APIs, and vector databases in complex, multi-step workflows. Without observability, debugging production issues becomes guesswork. Instrumenting your agents allows you to:

  • Monitor execution flow across LLM calls, tool invocations, and retrieval steps
  • Diagnose failures by examining stack traces and error contexts within specific trace spans
  • Identify latency bottlenecks by measuring duration of individual operations like embedding generation or API requests
  • Collect usage metrics such as token consumption, request rates, and error rates

High-Level Architecture for AI Agent Observability

A production-grade observability stack for AI agents consists of five distinct layers. In the awesome-llm-apps repository, these layers are implemented through dedicated modules and configuration files.

1. Instrumentation Layer

The instrumentation layer inserts trace spans, logs, and metrics directly into your agent code. Use the OpenTelemetry SDK for Python, LangChain Callbacks, or custom decorators. In the repository, this logic resides in src/agent.py where you initialize the tracer provider and wrap agent execution.

2. Trace Export Layer

This layer sends collected spans to a backend for storage and analysis. Configure an OTLP exporter to send data to Jaeger, Zipkin, Tempo, Azure Monitor, AWS X-Ray, or GCP Trace. The src/observability.py module handles exporter configuration and span processors.

3. Metrics Backend

Metrics backends aggregate counters, histograms, and gauges. Use the Prometheus client or OpenTelemetry metrics SDK to expose an HTTP /metrics endpoint. The repository includes a PrometheusMetricReader in src/observability.py that exposes metrics on port 8000.

4. Log Aggregation

Centralized logging requires structured JSON logs shipped to Loki, Elasticsearch, or CloudWatch. Configure your Python logger to emit JSON format and pipe output to your chosen sink. This ensures logs correlate with trace IDs for complete context.

5. Dashboard and Alerting

Visualization layers like Grafana or Datadog query your trace and metrics backends to display latency histograms, error rates, and trace waterfalls. The repository provides a docker-compose.yml that spins up Jaeger, Prometheus, and Grafana with pre-configured dashboards.

Instrumentation Patterns for AI Agents

Automatic Instrumentation with LLM Frameworks

Many LLM frameworks already emit OpenTelemetry-compatible trace events. LangChain and LlamaIndex provide callback handlers that automatically create spans for LLM calls, retrieval operations, and tool executions. Enable these by attaching the appropriate callback handler during agent initialization.

Manual Span Creation for Custom Logic

For custom business logic or external API calls not covered by automatic instrumentation, wrap code blocks with manual spans. This captures timing and metadata for database lookups, vector searches, or proprietary tool invocations.

with tracer.start_as_current_span("fetch_external_data"):
    data = fetch_external_api()
    # Add attributes to the span for filtering

    span = trace.get_current_span()
    span.set_attribute("api.endpoint", "https://api.example.com/data")

Context Propagation Across Service Boundaries

When your agent calls external microservices or APIs, propagate the trace context using W3C Trace Context headers. This ensures distributed traces span multiple services, allowing you to view the complete request lifecycle from user input through backend processing.

Implementing OpenTelemetry in Python AI Agents

The following implementation demonstrates how to add tracing and observability to AI agents in the awesome-llm-apps repository using OpenTelemetry.

Agent Implementation with Tracing

Create or modify src/agent.py to initialize the tracer provider and wrap agent execution with spans. This example shows automatic HTTP instrumentation and manual span creation for agent logic.


# src/agent.py

from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor, ConsoleSpanExporter
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.requests import RequestsInstrumentor
from langchain.callbacks import OpenAICallbackHandler

# 1️⃣ Set up tracer provider

resource = Resource(attributes={"service.name": "awesome-llm-agent"})
provider = TracerProvider(resource=resource)
trace.set_tracer_provider(provider)

# 2️⃣ Export spans to Jaeger via OTLP

otlp_exporter = OTLPSpanExporter(endpoint="http://localhost:4317", insecure=True)
provider.add_span_processor(BatchSpanProcessor(otlp_exporter))

# Optional: also log to console while developing

provider.add_span_processor(BatchSpanProcessor(ConsoleSpanExporter()))

tracer = trace.get_tracer(__name__)

# 3️⃣ Auto‑instrument HTTP requests

RequestsInstrumentor().instrument()

# 4️⃣ Agent logic with manual spans

def run_agent(prompt: str):
    with tracer.start_as_current_span("agent.run"):
        # LLM call – LangChain already emits its own spans when callbacks are attached

        callback = OpenAICallbackHandler()
        response = llm.invoke(prompt, callbacks=[callback])
        # Custom tool call

        with tracer.start_as_current_span("fetch_external_data"):
            data = fetch_external_api()
        return response

Observability Configuration Module

Create src/observability.py to centralize metrics export and telemetry configuration. This module sets up Prometheus metrics export and provides reusable counters for agent requests.


# src/observability.py

from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.exporter.prometheus import PrometheusMetricReader
from prometheus_client import start_http_server

# Export Prometheus metrics on port 8000

reader = PrometheusMetricReader()
provider = MeterProvider(metric_readers=[reader])
meter = provider.get_meter("awesome-llm-agent")

# Example counter

request_counter = meter.create_counter(
    "agent_requests_total",
    description="Number of agent requests processed",
)

def record_request():
    request_counter.add(1)

# Start the Prometheus endpoint (run once when the service starts)

if __name__ == "__main__":
    start_http_server(8000)

Repository Structure for Awesome LLM Apps

When integrating observability into the Shubhamsaboo/awesome-llm-apps repository, organize files to separate concerns and enable reuse across demo applications.

  • src/agent.py – Core agent implementation where you initialize the tracer provider and wrap execution logic with OpenTelemetry spans.

  • src/observability.py – Reusable module containing tracer configuration, OTLP exporters, and Prometheus metrics setup.

  • docker-compose.yml – Orchestration file that spins up Jaeger for trace storage, Prometheus for metrics aggregation, and Grafana for visualization.

  • README.md – Documentation entry point that explains how to enable observability features and links to the tracing guide.

Keep observability configuration environment-agnostic by reading endpoints such as OTEL_EXPORTER_OTLP_ENDPOINT and PROMETHEUS_PORT from environment variables. This ensures the same code runs locally during development and in production cloud deployments without modification.

Setting Up the Complete Observability Stack

Deploy a local observability stack using Docker Compose to visualize traces and metrics during development.

Docker Compose Configuration

Create a docker-compose.yml file that launches Jaeger, Prometheus, and Grafana with pre-configured data sources.

version: '3.8'
services:
  jaeger:
    image: jaegertracing/all-in-one:latest
    ports:
      - "16686:16686"
      - "4317:4317"
    environment:
      - COLLECTOR_OTLP_ENABLED=true

  prometheus:
    image: prom/prometheus:latest
    ports:
      - "9090:9090"
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml

  grafana:
    image: grafana/grafana:latest
    ports:
      - "3000:3000"
    volumes:
      - ./grafana/provisioning:/etc/grafana/provisioning

Grafana Dashboard Provisioning

Configure Grafana to automatically load dashboards for LLM agent monitoring.


# grafana/provisioning/dashboards/llm-agent.yml

apiVersion: 1
providers:
  - name: 'LLM Agent'
    folder: 'LLM'
    type: file
    options:
      path: /var/lib/grafana/dashboards

Import the Jaeger and Prometheus data sources into Grafana. Use the "Trace to logs" feature to jump directly from a trace span to the corresponding JSON log entry for complete debugging context.

Testing the Observability Setup

Verify the complete stack by running the agent and checking that traces and metrics appear in the respective backends.


# 1️⃣ Start observability stack

docker compose up -d  # brings up jaeger, prometheus, grafana

# 2️⃣ Run the agent

python -m src.agent "Summarize the latest news"

# 3️⃣ Verify

# - Open http://localhost:16686 (Jaeger UI) → see the trace.

# - Open http://localhost:9090/metrics → see `agent_requests_total`.

# - Open Grafana (http://localhost:3000) → view the LLM Agent dashboard.

When the agent receives a request, you will see a trace consisting of agent.run, the internal LLM call, and any custom tool spans. Prometheus exposes the agent_requests_total metric, and all logs can be shipped to a centralized logger for correlation with trace IDs.

Summary

  • Instrument early by adding tracing at the entry point of the agent and around every external call to capture complete execution flow.

  • Leverage existing callbacks from LangChain, LlamaIndex, and other LLM wrappers that already emit OpenTelemetry-compatible spans.

  • Export to standards-based backends using OTLP for traces (Jaeger, Tempo) and Prometheus for metrics to ensure your stack evolves without code changes.

  • Make observability optional by guarding imports with environment variables or feature flags so core demos remain lightweight while production deployments gain full visibility.

By following these patterns and committing the observability.py helper module to the Shubhamsaboo/awesome-llm-apps repository, you establish a reusable foundation for tracing and observability across all showcased AI agents.

Frequently Asked Questions

What is the difference between tracing and logging for AI agents?

Tracing captures the complete lifecycle of a request as it flows through various components of your agent, showing timing and relationships between operations like LLM calls and tool executions. Logging provides discrete event records with timestamps and messages. While logs tell you what happened at specific points, traces show you how the entire request propagated through the system, making them essential for debugging latency issues in complex agent workflows.

How do I choose between Jaeger, Zipkin, and Tempo for trace storage?

Jaeger provides a complete UI and storage solution with strong community support for OpenTelemetry, making it ideal for local development and small to medium deployments. Zipkin offers a simpler architecture with excellent Java ecosystem integration but fewer features for complex queries. Tempo, from Grafana Labs, is designed for high-volume cloud-native environments and works best when you already use Grafana for metrics visualization. For the awesome-llm-apps repository, Jaeger offers the fastest setup with Docker Compose.

Can I add observability to existing AI agents without rewriting the core logic?

Yes, you can add observability to existing agents through non-invasive instrumentation. Use OpenTelemetry's automatic instrumentation libraries for HTTP requests and database calls, and attach LangChain callbacks or LlamaIndex event handlers to capture LLM-specific telemetry without modifying business logic. Create a separate observability.py module that initializes the tracer provider, then import and use it as a decorator or context manager in your existing agent.py file. This approach keeps observability code isolated from agent functionality.

What metrics should I track for production AI agents?

Track request latency (histograms of end-to-end response times), token consumption (counters for input and output tokens per model), error rates (counters for failed LLM calls or tool executions), and queue depth (if using async processing). Additionally, monitor cost per request by multiplying token counts by model pricing, and cache hit rates for RAG applications. Expose these via Prometheus metrics in your observability.py module and create Grafana dashboards with alerts for p95 latency thresholds and error rate spikes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →