# How to Configure OpenTelemetry Spans in Switchyard to Trace Routing Decisions per Turn

> Configure OpenTelemetry spans in Switchyard to trace routing decisions per turn. Learn how Switchyard instruments each turn and captures routing details in spans.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Switchyard instruments every turn with OpenTelemetry spans through the `tracing` facade, capturing routing decisions in the `libsy.run` span's `switchyard.route` attribute after configuring a global OpenTelemetry provider and subscriber.**

The NVIDIA-NeMo/Switchyard routing engine provides built-in observability hooks that emit OpenTelemetry spans for every request and algorithm execution. By properly configuring these **OpenTelemetry spans in Switchyard**, you can trace which model route is selected for each turn, track session correlations, and analyze fallback behavior across distributed traces. The instrumentation resides in the `tracing` facade layer, requiring only a configured exporter and subscriber to begin capturing telemetry.

## Understanding the Span Architecture

Switchyard creates a hierarchical span structure that follows the request lifecycle from HTTP ingestion to LLM response. Each span type captures specific attributes that enable per-turn routing analysis.

### Request-Level Spans

The entry point instrumentation is defined in [`crates/switchyard-server/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/observability.rs) within the `request_span` function. This creates a `switchyard.request` span that acts as the root for each incoming HTTP request.

The span automatically extracts incoming **W3C trace context** from the `traceparent` header and records the `openinference.span.kind="CHAIN"` attribute. If the routing decision is known at the middleware level, it also populates the `switchyard.route` attribute with the target model identifier.

### Algorithm Execution Spans

Inside [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs), the `run_span` function generates the `libsy.run` span—the canonical location for routing decision telemetry. This span captures:

- `algorithm`: The routing algorithm name (e.g., "router")
- `switchyard.route`: The selected model ID (e.g., "myorg/gpt-4")
- `session_id`, `agent_id`, `task_id`, `correlation_id`: Correlation identifiers for the conversation
- `outcome_id` and `outcome`: Final execution status

When the algorithm completes, the `record_outcome` function appends the `selected_model_ids` array, documenting the fallback-ordered list of models attempted during the turn.

### LLM Call Spans

Nested within the algorithm span, instrumented via `#[tracing::instrument]` in `Algorithm::run_stream`, the `libsy.llm_call` span inherits the parent context and adds `selected_model_ids` alongside optional evidence fields. This nesting allows you to correlate token-level metrics with the high-level routing decision.

## Configuring the OpenTelemetry Provider in Rust

Switchyard does not bundle an exporter; you must initialize a global **OpenTelemetry** provider before starting the server. The following pattern configures an OTLP gRPC exporter compatible with Jaeger, Tempo, or the OpenTelemetry Collector.

```rust
use opentelemetry::{global, sdk::trace as sdktrace};
use opentelemetry_otlp::WithExportConfig;
use tracing_subscriber::{layer::SubscriberExt, Registry};

fn init_tracing() -> Result<(), Box<dyn std::error::Error>> {
    // Configure OTLP gRPC exporter targeting localhost:4317
    let exporter = opentelemetry_otlp::new_exporter()
        .tonic()
        .with_endpoint("http://localhost:4317")
        .build_exporter()?;

    let tracer = sdktrace::TracerProvider::builder()
        .with_batch_exporter(exporter, sdktrace::BatchConfig::default())
        .build()
        .get_tracer("switchyard", None);

    // Install as the global tracer provider
    global::set_tracer_provider(tracer.provider());

    // Bridge tracing spans to OpenTelemetry
    let otel_layer = tracing_opentelemetry::layer().with_tracer(tracer);
    let subscriber = Registry::default().with(otel_layer);
    tracing::subscriber::set_global_default(subscriber)?;
    Ok(())
}

```

Initialize this function before constructing the server state. The `switchyard-server` crate automatically attaches the `request_span` to incoming requests once the subscriber is active.

## Configuring OpenTelemetry via Python Bindings

When using Switchyard's Python bindings, configure the provider using the OpenTelemetry Python SDK. The Rust internals will forward spans through the same `tracing` facade.

```python
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace.export import BatchSpanProcessor

def init_otel():
    exporter = OTLPSpanExporter(endpoint="http://localhost:4317")
    provider = TracerProvider()
    provider.add_span_processor(BatchSpanProcessor(exporter))
    trace.set_tracer_provider(provider)

init_otel()

# Subsequent switchyard.run() calls automatically generate libsy.run spans

import switchyard_py as switchyard
response = switchyard.run(
    model="myorg/gpt-4",
    prompt="Explain observability.",
    metadata={"session_id": "sess-123"}
)

```

## Capturing Routing Decisions per Turn

The routing decision is recorded dynamically during algorithm execution. In [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs), the `run_span` function initializes the span with empty fields, then populates `switchyard.route` once the model selection logic resolves:

```rust
let span = tracing::info_span!(
    target: TRACING_TARGET,
    "libsy.run",
    algorithm,
    switchyard.route = tracing::field::Empty,
    // ... other correlation fields ...
);

if let Some(route) = request.model_id() {
    span.record("switchyard.route", route.as_ref());
}

```

When the algorithm finishes, `record_outcome` finalizes the span:

```rust
pub fn record_outcome(span: &tracing::Span, outcome: &Outcome) {
    span.record("outcome_id", &outcome.id());
    span.record("outcome", outcome.status());
    span.record("selected_model_ids", &format_model_ids(&outcome.models));
}

```

This creates a queryable attribute `switchyard.route` that represents the final model selected for that specific turn, even when fallbacks occur.

## Context Propagation and Distributed Tracing

Switchyard-server automatically extracts W3C **traceparent** headers from incoming HTTP requests in the `request_span` function. To maintain end-to-end visibility across services:

1. Ensure upstream clients include the `traceparent` header following the W3C Trace Context specification.
2. The extracted context becomes the parent of the `switchyard.request` span, linking Switchyard's internal spans with external trace roots.
3. Environment variables such as `OTEL_EXPORTER_OTLP_ENDPOINT` and `OTEL_EXPORTER_OTLP_HEADERS` configure the exporter behavior without modifying Switchyard code.

## Summary

- Switchyard emits three primary span types: `switchyard.request`, `libsy.run`, and `libsy.llm_call`, defined in [`crates/switchyard-server/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/observability.rs) and [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs).
- The **routing decision** is stored in the `switchyard.route` attribute of the `libsy.run` span, populated dynamically during algorithm execution.
- Configuration requires initializing a global OpenTelemetry **tracer provider** and a `tracing` subscriber before starting the server, with support for both Rust and Python initialization patterns.
- **W3C trace context** propagation enables distributed tracing across service boundaries by extracting `traceparent` headers at the server entry point.

## Frequently Asked Questions

### What span attribute contains the selected model route?

The `switchyard.route` attribute within the `libsy.run` span contains the model identifier (e.g., "myorg/gpt-4") chosen for the turn. This attribute is recorded in [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs) after the routing algorithm resolves the target model.

### How do I correlate spans across multiple Switchyard instances?

Include a W3C-compliant `traceparent` header in requests between services. Switchyard-server extracts this header in the `request_span` function of [`crates/switchyard-server/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/observability.rs), automatically parenting the incoming request span to the external trace context. This creates a continuous trace across distributed Switchyard deployments.

### Can I use Jaeger or Zipkin instead of OTLP?

Yes. Any OpenTelemetry-compatible exporter works with Switchyard's spans. Simply substitute the OTLP exporter in the initialization code with the Jaeger or Zipkin exporter from the `opentelemetry` ecosystem. The `tracing` facade forwards spans to whichever global provider you configure.

### What is the difference between `switchyard.request` and `libsy.run` spans?

The `switchyard.request` span in [`crates/switchyard-server/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/observability.rs) represents the HTTP request lifecycle and handles inbound trace context extraction. The `libsy.run` span in [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs) represents the algorithm execution lifecycle and contains the specific routing decision (`switchyard.route`) and outcome metadata. The run span is typically a child of the request span.