# How to Configure Observability with Langfuse in WeKnora

> Learn how to configure Langfuse observability in Tencent WeKnora. Easily track LLM generations with automatic activation and minimal runtime overhead. Get started today.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: how-to-guide
- Published: 2026-09-13

---

**WeKnora provides built-in Langfuse observability that activates automatically when `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY` are set, using an isolated OpenTelemetry tracer provider and middleware to capture LLM generation spans without runtime overhead when disabled.**

WeKnora is an open-source AI orchestration framework that ships with optional Langfuse-based observability for comprehensive LLM tracing. Configuring observability with Langfuse in WeKnora is entirely environment-driven and leverages a dedicated internal tracing package to isolate telemetry from your existing OpenTelemetry instrumentation. The integration captures detailed generation spans for chat, embedding, rerank, vision, and speech models while ensuring zero performance impact when deactivated.

## Core Architecture Components

WeKnora's observability stack is implemented in `internal/tracing/langfuse/` and designed to operate independently of other telemetry systems. The architecture consists of four primary components that work together to ingest traces into Langfuse.

### Langfuse Manager and Exporter

The **Langfuse Manager** ([`internal/tracing/langfuse/manager.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/manager.go)) creates a private `TracerProvider` specifically for WeKnora's internal use. Crucially, it does not invoke `otel.SetTracerProvider`, which prevents it from interfering with any existing OpenTelemetry instrumentation in your process. When Langfuse is disabled, this manager returns a no-op tracer that makes all subsequent calls computationally free.

The **Langfuse Exporter** ([`internal/tracing/langfuse/exporter.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/exporter.go)) implements an OTLP/HTTP client that transmits spans to `POST <host>/api/public/otel/v1/traces`. It adds required headers including `x-langfuse-ingestion-version: 4` to comply with the Langfuse v3/LiteFuse ingestion protocol.

### Middleware and Model Wrappers

The middleware layer ([`internal/tracing/langfuse/middleware.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/middleware.go)) provides `GinMiddleware` for HTTP handlers and `AsynqMiddleware` for background task processors. These wrappers propagate trace context across request boundaries and can skip health-check endpoints to reduce noise.

**Model wrappers** located in `internal/models/*/langfuse_wrapper.go` (covering chat, embedding, rerank, VLM, and ASR) automatically create Generation spans. These wrappers capture model names, input parameters, token usage metrics, and error states using attribute keys defined in [`internal/tracing/langfuse/events.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/events.go).

## Enabling Langfuse via Environment Variables

Langfuse activates automatically when both required keys are present. No code changes are necessary to enable or disable the integration.

### Required Variables

- **LANGFUSE_PUBLIC_KEY** – Your Langfuse project public key.
- **LANGFUSE_SECRET_KEY** – Your Langfuse project secret key.

When both variables are non-empty, [`internal/tracing/langfuse/config.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/config.go) initializes the exporter and tracer provider.

### Optional Configuration Tuning

| Variable | Description | Default |
|----------|-------------|---------|
| `LANGFUSE_HOST` | Base URL of the Langfuse API | `https://cloud.langfuse.com` |
| `LANGFUSE_ENABLED` | Explicit boolean toggle (`true`/`false`) | `true` (when keys present) |
| `LANGFUSE_RELEASE` | Release version tag shown in UI | — |
| `LANGFUSE_ENVIRONMENT` | Environment label (e.g., `prod`, `dev`) | — |
| `LANGFUSE_SAMPLE_RATE` | Sampling fraction (0.0 to 1.0) | `1.0` |
| `LANGFUSE_FLUSH_INTERVAL` | Export batch flush frequency | — |
| `LANGFUSE_QUEUE_SIZE` | Pending span queue capacity | — |
| `LANGFUSE_REQUEST_TIMEOUT` | HTTP client timeout for exports | — |
| `LANGFUSE_DEBUG` | Enable verbose exporter logging | `false` |

For self-hosted Langfuse instances, set `LANGFUSE_HOST` to your internal endpoint (e.g., `http://langfuse-web:3000`).

## Development and Production Deployment

WeKnora provides automation scripts and Docker Compose profiles to simplify running Langfuse alongside your application.

### Local Development with dev.sh

The [`scripts/dev.sh`](https://github.com/Tencent/WeKnora/blob/main/scripts/dev.sh) utility manages the full stack including a self-hosted Langfuse instance.

```bash

# Start WeKnora with Langfuse enabled (default)

./scripts/dev.sh start

# Start without Langfuse observability

./scripts/dev.sh start --no-langfuse

```

When Langfuse is enabled, the script automatically adds the `--profile langfuse` flag to Docker Compose and sets `ENABLED_SERVICES="langfuse"`.

### Docker Compose Profile

For production or isolated testing, use the dedicated Compose profile:

```bash

# Start only the observability stack

docker compose --profile langfuse up -d

```

Documentation for cloud image deployment is available in [`scripts/cloud-image/README.md`](https://github.com/Tencent/WeKnora/blob/main/scripts/cloud-image/README.md).

## Implementing Custom Tracing in Code

While WeKnora automatically traces all model interactions, you can manually instrument custom logic using the Langfuse manager. Initialize the singleton manager once at application startup, then use model wrappers to capture generations.

```go
package main

import (
    "github.com/Tencent/WeKnora/internal/tracing/langfuse"
    "github.com/Tencent/WeKnora/internal/models/chat"
)

func init() {
    // Initialize once; reads configuration from environment variables.
    _ = langfuse.NewManager()
}

func handleChatRequest(prompt string) (string, error) {
    // LangfuseWrapper creates a Generation span automatically.
    resp, err := chat.LangfuseWrapper(
        "gpt-4o-mini",
        map[string]any{"temperature": 0.7, "max_tokens": 150},
        func() (string, error) {
            // Your actual LLM call implementation
            return callExternalAPI(prompt)
        },
    )
    return resp, err
}

```

The `LangfuseWrapper` functions attach attributes defined in [`internal/tracing/langfuse/events.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/events.go) (such as `langfuse.observation.input` and `langfuse.observation.usage_details`) to ensure compatibility with the official Langfuse Python SDK schema.

## Summary

- **Zero-overhead design**: When `LANGFUSE_PUBLIC_KEY` or `LANGFUSE_SECRET_KEY` are absent, the `Manager` returns a no-op tracer that eliminates runtime cost.
- **Isolated telemetry**: The implementation in [`internal/tracing/langfuse/manager.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/manager.go) uses a private `TracerProvider` that does not overwrite global OpenTelemetry settings.
- **Environment-driven configuration**: All settings are controlled via environment variables parsed in [`internal/tracing/langfuse/config.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/config.go).
- **Automatic model tracing**: Wrappers in `internal/models/*/langfuse_wrapper.go` capture generation spans for chat, embedding, rerank, VLM, and ASR models without manual instrumentation.
- **Flexible deployment**: Use [`./scripts/dev.sh`](https://github.com/Tencent/WeKnora/blob/main/./scripts/dev.sh) for local development or `docker compose --profile langfuse` for production stacks.

## Frequently Asked Questions

### What environment variables are required to enable Langfuse in WeKnora?

You must set both `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY`. When these are present, WeKnora automatically enables tracing via the configuration logic in [`internal/tracing/langfuse/config.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/config.go). All other variables are optional and control sampling, timeouts, or metadata tags.

### How does WeKnora isolate Langfuse tracing from existing OpenTelemetry instrumentation?

The `LangfuseManager` in [`internal/tracing/langfuse/manager.go`](https://github.com/Tencent/WeKnora/blob/main/internal/tracing/langfuse/manager.go) instantiates a private `TracerProvider` that is not registered globally with `otel.SetTracerProvider`. This design pattern ensures that WeKnora's LLM spans are exported to Langfuse while leaving any existing OpenTelemetry collectors or tracer providers unaffected.

### Can I disable Langfuse observability without modifying source code?

Yes. Set `LANGFUSE_ENABLED=false` or simply remove the `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY` environment variables. The `NewManager()` function detects missing credentials and returns a no-op implementation, ensuring all tracing calls become cheap null operations with no external network requests.

### Which model interactions are automatically traced by WeKnora?

WeKnora automatically traces all interactions wrapped by the dedicated Langfuse wrapper functions. This includes chat completions ([`internal/models/chat/langfuse_wrapper.go`](https://github.com/Tencent/WeKnora/blob/main/internal/models/chat/langfuse_wrapper.go)), text embeddings, reranking, vision-language models (VLM), and automatic speech recognition (ASR). Each wrapper emits a Generation span containing model parameters, token usage, latency metrics, and error details to the Langfuse ingestion endpoint.