How to Configure OpenTelemetry Spans in Switchyard to Trace Routing Decisions per Turn
Switchyard instruments every turn with OpenTelemetry spans through the tracing facade, capturing routing decisions in the libsy.run span's switchyard.route attribute after configuring a global OpenTelemetry provider and subscriber.
The NVIDIA-NeMo/Switchyard routing engine provides built-in observability hooks that emit OpenTelemetry spans for every request and algorithm execution. By properly configuring these OpenTelemetry spans in Switchyard, you can trace which model route is selected for each turn, track session correlations, and analyze fallback behavior across distributed traces. The instrumentation resides in the tracing facade layer, requiring only a configured exporter and subscriber to begin capturing telemetry.
Understanding the Span Architecture
Switchyard creates a hierarchical span structure that follows the request lifecycle from HTTP ingestion to LLM response. Each span type captures specific attributes that enable per-turn routing analysis.
Request-Level Spans
The entry point instrumentation is defined in crates/switchyard-server/src/observability.rs within the request_span function. This creates a switchyard.request span that acts as the root for each incoming HTTP request.
The span automatically extracts incoming W3C trace context from the traceparent header and records the openinference.span.kind="CHAIN" attribute. If the routing decision is known at the middleware level, it also populates the switchyard.route attribute with the target model identifier.
Algorithm Execution Spans
Inside crates/libsy/src/observability.rs, the run_span function generates the libsy.run span—the canonical location for routing decision telemetry. This span captures:
algorithm: The routing algorithm name (e.g., "router")switchyard.route: The selected model ID (e.g., "myorg/gpt-4")session_id,agent_id,task_id,correlation_id: Correlation identifiers for the conversationoutcome_idandoutcome: Final execution status
When the algorithm completes, the record_outcome function appends the selected_model_ids array, documenting the fallback-ordered list of models attempted during the turn.
LLM Call Spans
Nested within the algorithm span, instrumented via #[tracing::instrument] in Algorithm::run_stream, the libsy.llm_call span inherits the parent context and adds selected_model_ids alongside optional evidence fields. This nesting allows you to correlate token-level metrics with the high-level routing decision.
Configuring the OpenTelemetry Provider in Rust
Switchyard does not bundle an exporter; you must initialize a global OpenTelemetry provider before starting the server. The following pattern configures an OTLP gRPC exporter compatible with Jaeger, Tempo, or the OpenTelemetry Collector.
use opentelemetry::{global, sdk::trace as sdktrace};
use opentelemetry_otlp::WithExportConfig;
use tracing_subscriber::{layer::SubscriberExt, Registry};
fn init_tracing() -> Result<(), Box<dyn std::error::Error>> {
// Configure OTLP gRPC exporter targeting localhost:4317
let exporter = opentelemetry_otlp::new_exporter()
.tonic()
.with_endpoint("http://localhost:4317")
.build_exporter()?;
let tracer = sdktrace::TracerProvider::builder()
.with_batch_exporter(exporter, sdktrace::BatchConfig::default())
.build()
.get_tracer("switchyard", None);
// Install as the global tracer provider
global::set_tracer_provider(tracer.provider());
// Bridge tracing spans to OpenTelemetry
let otel_layer = tracing_opentelemetry::layer().with_tracer(tracer);
let subscriber = Registry::default().with(otel_layer);
tracing::subscriber::set_global_default(subscriber)?;
Ok(())
}
Initialize this function before constructing the server state. The switchyard-server crate automatically attaches the request_span to incoming requests once the subscriber is active.
Configuring OpenTelemetry via Python Bindings
When using Switchyard's Python bindings, configure the provider using the OpenTelemetry Python SDK. The Rust internals will forward spans through the same tracing facade.
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace.export import BatchSpanProcessor
def init_otel():
exporter = OTLPSpanExporter(endpoint="http://localhost:4317")
provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(exporter))
trace.set_tracer_provider(provider)
init_otel()
# Subsequent switchyard.run() calls automatically generate libsy.run spans
import switchyard_py as switchyard
response = switchyard.run(
model="myorg/gpt-4",
prompt="Explain observability.",
metadata={"session_id": "sess-123"}
)
Capturing Routing Decisions per Turn
The routing decision is recorded dynamically during algorithm execution. In crates/libsy/src/observability.rs, the run_span function initializes the span with empty fields, then populates switchyard.route once the model selection logic resolves:
let span = tracing::info_span!(
target: TRACING_TARGET,
"libsy.run",
algorithm,
switchyard.route = tracing::field::Empty,
// ... other correlation fields ...
);
if let Some(route) = request.model_id() {
span.record("switchyard.route", route.as_ref());
}
When the algorithm finishes, record_outcome finalizes the span:
pub fn record_outcome(span: &tracing::Span, outcome: &Outcome) {
span.record("outcome_id", &outcome.id());
span.record("outcome", outcome.status());
span.record("selected_model_ids", &format_model_ids(&outcome.models));
}
This creates a queryable attribute switchyard.route that represents the final model selected for that specific turn, even when fallbacks occur.
Context Propagation and Distributed Tracing
Switchyard-server automatically extracts W3C traceparent headers from incoming HTTP requests in the request_span function. To maintain end-to-end visibility across services:
- Ensure upstream clients include the
traceparentheader following the W3C Trace Context specification. - The extracted context becomes the parent of the
switchyard.requestspan, linking Switchyard's internal spans with external trace roots. - Environment variables such as
OTEL_EXPORTER_OTLP_ENDPOINTandOTEL_EXPORTER_OTLP_HEADERSconfigure the exporter behavior without modifying Switchyard code.
Summary
- Switchyard emits three primary span types:
switchyard.request,libsy.run, andlibsy.llm_call, defined incrates/switchyard-server/src/observability.rsandcrates/libsy/src/observability.rs. - The routing decision is stored in the
switchyard.routeattribute of thelibsy.runspan, populated dynamically during algorithm execution. - Configuration requires initializing a global OpenTelemetry tracer provider and a
tracingsubscriber before starting the server, with support for both Rust and Python initialization patterns. - W3C trace context propagation enables distributed tracing across service boundaries by extracting
traceparentheaders at the server entry point.
Frequently Asked Questions
What span attribute contains the selected model route?
The switchyard.route attribute within the libsy.run span contains the model identifier (e.g., "myorg/gpt-4") chosen for the turn. This attribute is recorded in crates/libsy/src/observability.rs after the routing algorithm resolves the target model.
How do I correlate spans across multiple Switchyard instances?
Include a W3C-compliant traceparent header in requests between services. Switchyard-server extracts this header in the request_span function of crates/switchyard-server/src/observability.rs, automatically parenting the incoming request span to the external trace context. This creates a continuous trace across distributed Switchyard deployments.
Can I use Jaeger or Zipkin instead of OTLP?
Yes. Any OpenTelemetry-compatible exporter works with Switchyard's spans. Simply substitute the OTLP exporter in the initialization code with the Jaeger or Zipkin exporter from the opentelemetry ecosystem. The tracing facade forwards spans to whichever global provider you configure.
What is the difference between switchyard.request and libsy.run spans?
The switchyard.request span in crates/switchyard-server/src/observability.rs represents the HTTP request lifecycle and handles inbound trace context extraction. The libsy.run span in crates/libsy/src/observability.rs represents the algorithm execution lifecycle and contains the specific routing decision (switchyard.route) and outcome metadata. The run span is typically a child of the request span.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →