How to Trace and Debug AI Agents with LangSmith: Complete Setup Guide

Enable LangSmith by setting LANGCHAIN_TRACING_V2=true and LANGCHAIN_PROJECT environment variables, then execute your LangGraph agent to automatically capture every LLM call, tool invocation, and state transition as a structured trace without modifying your application code.

Tracing AI agents in production requires granular visibility into decision-making logic, external API calls, and state mutations. This guide demonstrates how to trace and debug AI agents with LangSmith using the reference implementation in the NirDiamant/agents-towards-production repository. By configuring a few environment variables, you can instrument a LangGraph agent to stream detailed execution traces to the LangSmith observability platform.

Installation and Project Setup

Before configuring tracing, install the required dependencies from the tutorial located at tutorials/tracing-with-langsmith/langsmith_basics.ipynb. These packages provide the LLM client, graph execution engine, and tracing backend.

pip install -U langchain-core langchain-openai langgraph langsmith requests

The langsmith package contains the client library that serializes trace data, while langgraph provides the state-machine framework that LangSmith automatically instruments.

Configuring Environment Variables for Autonomous Tracing

LangSmith operates on a zero-code-instrumentation model. The SDK reads specific environment variables at import time to configure the tracer. Set these variables before importing LangChain components:

import os

os.environ["OPENAI_API_KEY"] = "sk-..."
os.environ["LANGCHAIN_API_KEY"] = "ls-..."
os.environ["LANGCHAIN_TRACING_V2"] = "true"
os.environ["LANGCHAIN_PROJECT"] = "langsmith-tutorial-demo"
  • LANGCHAIN_TRACING_V2: Must be set to "true" to enable the flight-recorder mode.
  • LANGCHAIN_PROJECT: Organizes traces into named dashboards (e.g., langsmith-tutorial-demo).
  • LANGCHAIN_API_KEY: Authenticates your requests to the LangSmith API.

When these variables are present, the LangChain SDK injects a Tracer callback into every LLM, Tool, and StateGraph operation, streaming serialized request/response payloads asynchronously to LangSmith.

Building an Observable Agent with Typed State

To maximize the debugging value of traces, define your agent state using a TypedDict schema. This allows LangSmith to surface the exact data flowing through your workflow rather than opaque objects.

Define the State Schema

from typing import TypedDict

class AgentState(TypedDict):
    user_question: str
    needs_search: bool
    search_result: str
    final_answer: str
    reasoning: str

Each field in AgentState appears as a structured payload in the LangSmith trace view, enabling you to inspect intermediate values like needs_search or search_result at each graph step.

Create Instrumented Tools

Define tools using the @tool decorator. LangSmith automatically wraps these functions, logging start/end timestamps, arguments, return values, and exception traces:

from langchain_core.tools import tool
import requests

@tool
def wikipedia_search(query: str) -> str:
    """Search Wikipedia and return a concise snippet."""
    resp = requests.get(
        "https://en.wikipedia.org/w/api.php",
        params={
            "action": "query",
            "list": "search",
            "srsearch": query,
            "format": "json"
        },
        timeout=10,
    )
    resp.raise_for_status()
    results = resp.json()["query"]["search"]
    if not results:
        return "No relevant result found."
    return results[0]["snippet"]

Assemble the StateGraph

Construct your agent logic using StateGraph from LangGraph. Each node function becomes a distinct span in the LangSmith trace:

from langchain_openai import ChatOpenAI
from langgraph.graph import StateGraph, END

# Deterministic LLM for reproducible traces

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)

def router(state: AgentState):
    """Decide whether to search or answer directly."""
    state["needs_search"] = "who" in state["user_question"].lower()
    return "search" if state["needs_search"] else "answer"

def search(state: AgentState):
    """Execute Wikipedia search."""
    state["search_result"] = wikipedia_search(state["user_question"])
    return "answer"

def answer(state: AgentState):
    """Generate final response."""
    prompt = (
        f"Question: {state['user_question']}\n"
        f"Search result: {state.get('search_result', '')}\n"
        "Provide a concise answer."
    )
    response = llm.invoke(prompt)
    state["final_answer"] = response.content
    return END

# Build graph

graph = StateGraph(AgentState)
graph.add_node("router", router)
graph.add_node("search", search)
graph.add_node("answer", answer)
graph.set_entry_point("router")
graph.add_conditional_edges("router", lambda s: "search" if s["needs_search"] else "answer")
graph.add_edge("search", "answer")

graph_compiled = graph.compile()

Executing Traced Runs

Invoke the compiled graph to generate a trace. The entire execution is recorded as a single parent trace containing nested sub-spans for every LLM call and tool invocation:

result = graph_compiled.invoke({
    "user_question": "Who invented the internet?"
})
print(result["final_answer"])

Under the hood, the LangChain Tracer captures:

  • The input/output of the router decision node
  • The HTTP latency of the wikipedia_search tool call
  • Token usage and latency for the answer node's LLM invocation
  • State diffs between each transition

Analyzing Traces in the LangSmith Dashboard

Navigate to https://smith.langchain.com and select your project name (langsmith-tutorial-demo). The dashboard renders several diagnostic views:

  • Node Graph: A visual flow diagram showing the execution path (router → search → answer).
  • Latency Breakdown: Per-step timing metrics revealing bottlenecks in tool calls or LLM latency.
  • Token Economics: Automatic calculation of prompt and completion token counts for cost optimization.
  • Error Inspection: Stack traces and exception messages for failed tool executions or LLM errors.

Because the tracer operates at the library level, you receive semantic observability (e.g., "Node search performed Wikipedia lookup") rather than unstructured text logs.

Summary

  • Zero-code instrumentation: Set LANGCHAIN_TRACING_V2=true and LANGCHAIN_PROJECT to capture traces without sprinkling print statements throughout your codebase.
  • Structured state visibility: Use TypedDict schemas for AgentState so LangSmith can display the exact data moving through your workflow at each step.
  • Granular span capture: Every tool execution, LLM request, and graph transition appears as a distinct, timed span with automatic error tracking.
  • Performance insights: LangSmith reports token counts and latency metrics out-of-the-box, enabling precise cost and speed optimization.

Frequently Asked Questions

Do I need to modify my agent code to enable LangSmith tracing?

No. Tracing is activated purely through environment variables (LANGCHAIN_TRACING_V2 and LANGCHAIN_PROJECT). The LangChain SDK automatically installs callbacks on LLM and Tool classes at import time, provided the variables are set before you instantiate components.

What latency overhead does LangSmith tracing add?

The tracer uses asynchronous background threads to POST trace data to LangSmith's API, so the overhead on your agent's execution time is typically negligible (under 5ms per span). The critical path of your agent remains unblocked while telemetry streams in the background.

Can I trace custom Python functions that are not LangChain tools?

Yes. While LangSmith automatically traces @tool decorated functions and LangChain primitives, you can manually instrument arbitrary functions using the LangSmith SDK's @traceable decorator or by creating custom callbacks. However, for most agent debugging scenarios, wrapping utilities as LangChain tools provides sufficient visibility.

How do I organize traces across development and production environments?

Use the LANGCHAIN_PROJECT environment variable to segregate traces by environment. Set distinct project names like agent-dev, agent-staging, and agent-prod. The LangSmith dashboard filters by project, allowing you to isolate debugging sessions and prevent test data from polluting production analytics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →