How to Configure OpenTelemetry Spans in Switchyard to Trace Routing Decisions per Turn

Switchyard instruments every turn with OpenTelemetry spans through the tracing facade, capturing routing decisions in the libsy.run span's switchyard.route attribute after configuring a global OpenTelemetry provider and subscriber.

The NVIDIA-NeMo/Switchyard routing engine provides built-in observability hooks that emit OpenTelemetry spans for every request and algorithm execution. By properly configuring these OpenTelemetry spans in Switchyard, you can trace which model route is selected for each turn, track session correlations, and analyze fallback behavior across distributed traces. The instrumentation resides in the tracing facade layer, requiring only a configured exporter and subscriber to begin capturing telemetry.

Understanding the Span Architecture

Switchyard creates a hierarchical span structure that follows the request lifecycle from HTTP ingestion to LLM response. Each span type captures specific attributes that enable per-turn routing analysis.

Request-Level Spans

The entry point instrumentation is defined in crates/switchyard-server/src/observability.rs within the request_span function. This creates a switchyard.request span that acts as the root for each incoming HTTP request.

The span automatically extracts incoming W3C trace context from the traceparent header and records the openinference.span.kind="CHAIN" attribute. If the routing decision is known at the middleware level, it also populates the switchyard.route attribute with the target model identifier.

Algorithm Execution Spans

Inside crates/libsy/src/observability.rs, the run_span function generates the libsy.run span—the canonical location for routing decision telemetry. This span captures:

  • algorithm: The routing algorithm name (e.g., "router")
  • switchyard.route: The selected model ID (e.g., "myorg/gpt-4")
  • session_id, agent_id, task_id, correlation_id: Correlation identifiers for the conversation
  • outcome_id and outcome: Final execution status

When the algorithm completes, the record_outcome function appends the selected_model_ids array, documenting the fallback-ordered list of models attempted during the turn.

LLM Call Spans

Nested within the algorithm span, instrumented via #[tracing::instrument] in Algorithm::run_stream, the libsy.llm_call span inherits the parent context and adds selected_model_ids alongside optional evidence fields. This nesting allows you to correlate token-level metrics with the high-level routing decision.

Configuring the OpenTelemetry Provider in Rust

Switchyard does not bundle an exporter; you must initialize a global OpenTelemetry provider before starting the server. The following pattern configures an OTLP gRPC exporter compatible with Jaeger, Tempo, or the OpenTelemetry Collector.

use opentelemetry::{global, sdk::trace as sdktrace};
use opentelemetry_otlp::WithExportConfig;
use tracing_subscriber::{layer::SubscriberExt, Registry};

fn init_tracing() -> Result<(), Box<dyn std::error::Error>> {
    // Configure OTLP gRPC exporter targeting localhost:4317
    let exporter = opentelemetry_otlp::new_exporter()
        .tonic()
        .with_endpoint("http://localhost:4317")
        .build_exporter()?;

    let tracer = sdktrace::TracerProvider::builder()
        .with_batch_exporter(exporter, sdktrace::BatchConfig::default())
        .build()
        .get_tracer("switchyard", None);

    // Install as the global tracer provider
    global::set_tracer_provider(tracer.provider());

    // Bridge tracing spans to OpenTelemetry
    let otel_layer = tracing_opentelemetry::layer().with_tracer(tracer);
    let subscriber = Registry::default().with(otel_layer);
    tracing::subscriber::set_global_default(subscriber)?;
    Ok(())
}

Initialize this function before constructing the server state. The switchyard-server crate automatically attaches the request_span to incoming requests once the subscriber is active.

Configuring OpenTelemetry via Python Bindings

When using Switchyard's Python bindings, configure the provider using the OpenTelemetry Python SDK. The Rust internals will forward spans through the same tracing facade.

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace.export import BatchSpanProcessor

def init_otel():
    exporter = OTLPSpanExporter(endpoint="http://localhost:4317")
    provider = TracerProvider()
    provider.add_span_processor(BatchSpanProcessor(exporter))
    trace.set_tracer_provider(provider)

init_otel()

# Subsequent switchyard.run() calls automatically generate libsy.run spans

import switchyard_py as switchyard
response = switchyard.run(
    model="myorg/gpt-4",
    prompt="Explain observability.",
    metadata={"session_id": "sess-123"}
)

Capturing Routing Decisions per Turn

The routing decision is recorded dynamically during algorithm execution. In crates/libsy/src/observability.rs, the run_span function initializes the span with empty fields, then populates switchyard.route once the model selection logic resolves:

let span = tracing::info_span!(
    target: TRACING_TARGET,
    "libsy.run",
    algorithm,
    switchyard.route = tracing::field::Empty,
    // ... other correlation fields ...
);

if let Some(route) = request.model_id() {
    span.record("switchyard.route", route.as_ref());
}

When the algorithm finishes, record_outcome finalizes the span:

pub fn record_outcome(span: &tracing::Span, outcome: &Outcome) {
    span.record("outcome_id", &outcome.id());
    span.record("outcome", outcome.status());
    span.record("selected_model_ids", &format_model_ids(&outcome.models));
}

This creates a queryable attribute switchyard.route that represents the final model selected for that specific turn, even when fallbacks occur.

Context Propagation and Distributed Tracing

Switchyard-server automatically extracts W3C traceparent headers from incoming HTTP requests in the request_span function. To maintain end-to-end visibility across services:

  1. Ensure upstream clients include the traceparent header following the W3C Trace Context specification.
  2. The extracted context becomes the parent of the switchyard.request span, linking Switchyard's internal spans with external trace roots.
  3. Environment variables such as OTEL_EXPORTER_OTLP_ENDPOINT and OTEL_EXPORTER_OTLP_HEADERS configure the exporter behavior without modifying Switchyard code.

Summary

  • Switchyard emits three primary span types: switchyard.request, libsy.run, and libsy.llm_call, defined in crates/switchyard-server/src/observability.rs and crates/libsy/src/observability.rs.
  • The routing decision is stored in the switchyard.route attribute of the libsy.run span, populated dynamically during algorithm execution.
  • Configuration requires initializing a global OpenTelemetry tracer provider and a tracing subscriber before starting the server, with support for both Rust and Python initialization patterns.
  • W3C trace context propagation enables distributed tracing across service boundaries by extracting traceparent headers at the server entry point.

Frequently Asked Questions

What span attribute contains the selected model route?

The switchyard.route attribute within the libsy.run span contains the model identifier (e.g., "myorg/gpt-4") chosen for the turn. This attribute is recorded in crates/libsy/src/observability.rs after the routing algorithm resolves the target model.

How do I correlate spans across multiple Switchyard instances?

Include a W3C-compliant traceparent header in requests between services. Switchyard-server extracts this header in the request_span function of crates/switchyard-server/src/observability.rs, automatically parenting the incoming request span to the external trace context. This creates a continuous trace across distributed Switchyard deployments.

Can I use Jaeger or Zipkin instead of OTLP?

Yes. Any OpenTelemetry-compatible exporter works with Switchyard's spans. Simply substitute the OTLP exporter in the initialization code with the Jaeger or Zipkin exporter from the opentelemetry ecosystem. The tracing facade forwards spans to whichever global provider you configure.

What is the difference between switchyard.request and libsy.run spans?

The switchyard.request span in crates/switchyard-server/src/observability.rs represents the HTTP request lifecycle and handles inbound trace context extraction. The libsy.run span in crates/libsy/src/observability.rs represents the algorithm execution lifecycle and contains the specific routing decision (switchyard.route) and outcome metadata. The run span is typically a child of the request span.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →