# How to Set Up Langfuse Tracing with v3 SDK for RAG Pipeline Observability

> Learn to set up Langfuse tracing with the v3 SDK for RAG pipeline observability. Enhance your application's visibility and debug complex workflows efficiently.

- Repository: [jamwithai/production-agentic-rag-course](https://github.com/jamwithai/production-agentic-rag-course)
- Tags: how-to-guide
- Published: 2026-03-23

---

**TL;DR:** Install `langfuse>=3.0.0`, configure environment variables with your project credentials, initialize a **singleton `LangfuseTracer`** via the factory in [`src/services/langfuse/factory.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/factory.py), attach it to FastAPI's application state in [`src/main.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/main.py), then use the **context-manager API** (`trace_request`, `trace_embedding`, `trace_search`, `trace_generation`) inside your RAG workflow to capture full trace hierarchies.

The `jamwithai/production-agentic-rag-course` repository demonstrates a production-ready integration of the Langfuse v3 SDK for end-to-end observability in agentic RAG systems. This guide covers the exact implementation patterns used to trace embeddings, retrievals, and LLM generations with full hierarchical context.

## Install and Configure the Langfuse v3 SDK

Start by installing the v3 SDK and verifying your environment uses the correct major version.

```bash
pip install "langfuse>=3.0.0"

```

The repository pins this requirement in [`requirements.txt`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/requirements.txt), but explicitly installing ensures compatibility with the v3 API surface used throughout the codebase.

Next, add your Langfuse credentials to a `.env` file at the repository root. The `LangfuseSettings` class in [`src/config.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/config.py) reads these using the `LANGFUSE__` prefix convention.

```dotenv
LANGFUSE__PUBLIC_KEY=pk-lf-your-public-key
LANGFUSE__SECRET_KEY=sk-lf-your-secret-key
LANGFUSE__HOST=http://localhost:3000
LANGFUSE__ENABLED=true
LANGFUSE__FLUSH_AT=15
LANGFUSE__FLUSH_INTERVAL=1.0

```

- **`LANGFUSE__ENABLED`**: Set to `false` to disable tracing without code changes.
- **`LANGFUSE__FLUSH_AT`**: Number of events to batch before auto-flushing.
- **`LANGFUSE__FLUSH_INTERVAL`**: Seconds between background flushes.

## Initialize the Singleton Tracer

The repository implements a **singleton pattern** to ensure one `LangfuseTracer` instance exists per application lifecycle. This prevents duplicate client initialization and maintains trace context across requests.

In [`src/services/langfuse/factory.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/factory.py), the `make_langfuse_tracer()` function uses `functools.lru_cache` to guarantee a single instance:

```python

# src/services/langfuse/factory.py

from functools import lru_cache
from src.config import get_settings
from src.services.langfuse.client import LangfuseTracer

@lru_cache(maxsize=1)
def make_langfuse_tracer() -> LangfuseTracer:
    settings = get_settings()
    return LangfuseTracer(settings)

```

The `LangfuseTracer` class in [`src/services/langfuse/client.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/client.py) wraps the native v3 client (`langfuse.Langfuse`), handling authentication, batching configuration, and span lifecycle management.

## Integrate with FastAPI Lifecycle

Wire the tracer into FastAPI's startup sequence so it becomes available to all request handlers. In [`src/main.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/main.py), the lifespan context manager attaches the tracer to `app.state`:

```python

# src/main.py (inside lifespan context manager)

from src.services.langfuse.factory import make_langfuse_tracer

async def lifespan(app: FastAPI):
    # Startup

    app.state.langfuse_tracer = make_langfuse_tracer()
    logger.info("Langfuse tracer initialized")
    yield
    # Shutdown

    app.state.langfuse_tracer.shutdown()

```

Create a dependency in [`src/dependencies.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/dependencies.py) to inject the tracer into endpoints:

```python

# src/dependencies.py

from typing import Annotated
from fastapi import Depends, Request
from src.services.langfuse.client import LangfuseTracer

def get_langfuse_tracer(request: Request) -> LangfuseTracer:
    return request.app.state.langfuse_tracer

LangfuseDep = Annotated[LangfuseTracer, Depends(get_langfuse_tracer)]

```

## Instrument the RAG Pipeline

With the tracer available via dependency injection, instrument your RAG orchestration logic. The `AgenticRAGService` in [`src/services/agents/agentic_rag.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/agents/agentic_rag.py) demonstrates the full integration pattern.

### Open Top-Level Request Traces

Begin each RAG request by creating a root trace using the v3 SDK's `start_as_current_span` method:

```python

# src/services/agents/agentic_rag.py

class AgenticRAGService:
    def __init__(self, langfuse_tracer: LangfuseTracer):
        self.langfuse_tracer = langfuse_tracer
    
    async def ask(self, question: str, user_id: str):
        # Open top-level trace (v3 SDK)

        trace = self.langfuse_tracer.client.start_as_current_span(
            name="agentic_rag_request",
            input={"question": question, "user_id": user_id}
        )
        # ... workflow execution ...

        trace.end()
        self.langfuse_tracer.flush()

```

### Auto-Instrument LLM Calls with CallbackHandler

Use Langfuse's `CallbackHandler` to automatically capture every LangChain or LangGraph LLM invocation as a child span. Pass it through the runtime configuration:

```python

# Inside AgenticRAGService._run_workflow()

from langfuse.langchain import CallbackHandler

if self.langfuse_tracer:
    callback_handler = CallbackHandler()  # Auto-links to active trace

    config = {"callbacks": [callback_handler], "recursion_limit": 100}
    # Pass config to your LangGraph app or chain

    result = await app.ainvoke(inputs, config=config)

```

### Create Custom Spans for RAG Steps

For granular observability of embeddings, retrieval, and prompt construction, use the helper methods in [`src/services/langfuse/tracer.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/tracer.py). The `RAGTracer` class provides context managers that handle span creation, timing, and metadata attachment:

```python

# src/services/langfuse/tracer.py (excerpt)

from contextlib import contextmanager
import time

class RAGTracer:
    def __init__(self, tracer: LangfuseTracer):
        self.tracer = tracer
    
    @contextmanager
    def trace_embedding(self, trace, query: str):
        start = time.time()
        span = self.tracer.create_span(
            trace=trace,
            name="query_embedding",
            input_data={"query": query, "length": len(query)}
        )
        try:
            yield span
        finally:
            duration = time.time() - start
            self.tracer.update_span(
                span=span,
                output={"duration_ms": round(duration * 1000, 2)}
            )
            span.end()

```

Use these helpers inside your RAG nodes to create hierarchical spans:

```python

# Example usage inside a retrieval node

async def retrieve_node(state, runtime):
    tracer = runtime.context.langfuse_tracer
    rag_tracer = RAGTracer(tracer)
    query = state["messages"][-1].content
    
    with rag_tracer.trace_embedding(trace, query) as span:
        embeddings = await embedding_model.embed(query)
    
    with rag_tracer.trace_search(trace, query, top_k=10) as span:
        docs = await vector_store.search(embeddings, k=10)
        rag_tracer.end_search(span, docs, total_hits=len(docs))
    
    return {"documents": docs}

```

The `RAGTracer` also exposes `trace_prompt_construction`, `trace_generation`, and `trace_request` for complete pipeline coverage.

## Complete Implementation Example

Here is a minimal FastAPI endpoint that returns the Langfuse trace ID alongside the generated answer, demonstrating the full integration:

```python

# src/routers/ask.py

from fastapi import APIRouter
from src.dependencies import LangfuseDep, AgenticRAGDep

ask_router = APIRouter()

@ask_router.post("/api/v1/ask")
async def ask_endpoint(
    question: str,
    langfuse_tracer: LangfuseDep,
    rag_service: AgenticRAGDep
):
    # Execute RAG with full tracing enabled

    result = await rag_service.ask(question, user_id="api_user")
    
    # Retrieve trace ID for UI inspection

    trace_id = langfuse_tracer.get_trace_id()
    
    return {
        "answer": result["answer"],
        "trace_id": trace_id,
        "sources": result.get("sources", [])
    }

```

**Response format:**

```json
{
  "answer": "Langfuse v3 uses OpenTelemetry-compatible span tracking...",
  "trace_id": "abc123-def456",
  "sources": ["arxiv:2401.12345"]
}

```

## Summary

- **Install** `langfuse>=3.0.0` and configure credentials via `LANGFUSE__*` environment variables in [`src/config.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/config.py).
- **Initialize** a singleton `LangfuseTracer` using the factory in [`src/services/langfuse/factory.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/factory.py) to ensure one client instance per application.
- **Attach** the tracer to FastAPI's `app.state` in [`src/main.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/main.py) and expose it via `LangfuseDep` in [`src/dependencies.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/dependencies.py).
- **Trace** each request by opening a top-level span with `client.start_as_current_span()` in your RAG service.
- **Auto-capture** LLM calls by injecting `CallbackHandler` into LangChain/LangGraph configurations.
- **Instrument** specific RAG steps (embedding, retrieval, generation) using the `RAGTracer` context managers in [`src/services/langfuse/tracer.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/tracer.py).

## Frequently Asked Questions

### How do I disable Langfuse tracing in development without removing code?

Set `LANGFUSE__ENABLED=false` in your `.env` file. The `LangfuseTracer` initialization checks this flag and returns no-op implementations, allowing your application to run without attempting authentication or network calls to the Langfuse server.

### What is the difference between `LangfuseTracer` and `RAGTracer`?

**`LangfuseTracer`** (in [`src/services/langfuse/client.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/client.py)) is a thin wrapper around the native v3 `Langfuse` client, handling authentication, batch flushing, and low-level span CRUD operations. **`RAGTracer`** (in [`src/services/langfuse/tracer.py`](https://github.com/jamwithai/production-agentic-rag-course/blob/main/src/services/langfuse/tracer.py)) is a higher-level abstraction providing semantic context managers like `trace_embedding` and `trace_search` that automatically capture RAG-specific metadata and timing.

### Why use a singleton factory instead of creating the client per request?

The Langfuse v3 SDK maintains internal buffering and background flushing threads. Creating multiple instances per request would leak resources and fragment trace context. The `make_langfuse_tracer()` factory with `@lru_cache(maxsize=1)` ensures all requests share one client with proper lifecycle management (flush on shutdown).

### How do I correlate spans across asynchronous RAG nodes?

The v3 SDK propagates trace context automatically through the `start_as_current_span()` API. When you call `trace = self.langfuse_tracer.client.start_as_current_span()`, subsequent span creations (including those inside `CallbackHandler` or `RAGTracer` methods) automatically become children of that trace without manual ID passing, provided they execute within the same async context.