How to Debug Issues in Switchyard: OpenTelemetry Metrics and Tracing Guide
Enable verbose tracing with RUST_LOG=switchyard_server=debug,libsy=debug, attach a tracing subscriber, and query OpenTelemetry metrics to inspect routing decisions and model selections in real-time.
Switchyard is a Rust-based LLM routing engine built from modular crates that expose OpenTelemetry metrics and tracing spans for every routing decision. To effectively debug issues in Switchyard, you configure environment-based logging, attach tracing subscribers, and analyze exported telemetry data to see exactly how algorithms behave, why specific models were chosen, and where failures occurred. The system records all activity through the libsy library, making debugging possible without modifying core routing logic.
Understanding Switchyard's Observability Architecture
The Core libsy Layer
The libsy crate in crates/libsy/src/observability.rs provides the central telemetry functions that capture algorithm activity. The run_span function (lines 68-94) creates a span that captures request metadata including model ID, session ID, and correlation ID. Helper functions like record_outcome, record_llm_call, and record_decision (lines 16-33 and 81-89) add attributes such as selected models, evidence scores, and token usage to the current span. The observe_run wrapper (lines 83-98) counts in-flight runs and measures execution timing.
All these functions use the global OpenTelemetry meter under the switchyard scope. If no SDK is installed, the calls are no-ops, allowing you to debug issues in Switchyard without changing any library code.
Algorithm Tracing Instrumentation
The algorithm entry point at Algorithm::run_stream in crates/libsy/src/core/algorithm.rs (lines 170-182) is annotated with #[tracing::instrument]. This macro automatically creates a child span for each routing step—such as CallModel and Done—and propagates the parent libsy.run span created by run_span. This hierarchical structure lets you trace a single request through every decision point in the routing tree.
Server-Side Telemetry
When running the standalone switchyard-server, the server adds HTTP handling spans and publishes Prometheus counters. The server's observability module in crates/switchyard-server/src/observability.rs (lines 45-55) records request-level errors and OpenTelemetry exporter status. The server initialization in crates/switchyard-server/src/lib.rs (lines 47-56) shows how the global subscriber is installed and how the metrics endpoint is exposed.
Step-by-Step Debugging Workflow
1. Enable Verbose Environment Logging
Set the RUST_LOG environment variable to activate debug output before launching the server or your Rust binary:
export RUST_LOG=switchyard_server=debug,libsy=debug
For maximum detail during complex routing investigations, use trace instead of debug. This setting logs routing decisions, classifier failures, and metric increments to stderr.
2. Attach a Tracing Subscriber
In your Rust application, initialize a subscriber to print the spans generated by run_span and the algorithm's instrumentation macros:
use tracing_subscriber::{fmt, EnvFilter};
fn main() {
tracing_subscriber::registry()
.with(fmt::layer())
.with(EnvFilter::from_default_env())
.init();
// Switchyard algorithm execution begins here
}
This subscriber reads the RUST_LOG variable and formats the span hierarchy, showing exactly which algorithm step is executing.
3. Configure OpenTelemetry Exporters
To persist metrics beyond stdout, install an OpenTelemetry SDK before calling any Switchyard code:
use opentelemetry::global;
use opentelemetry_jaeger::JaegerPipelineBuilder;
fn init_otel() {
JaegerPipelineBuilder::default()
.with_service_name("switchyard")
.install_simple()
.expect("Jaeger exporter failed");
global::set_meter_provider(opentelemetry::sdk::metrics::MeterProvider::builder().build());
}
After initialization, metrics such as switchyard.run_duration_ms, switchyard.decisions, and switchyard.runs become visible in your observability backend.
4. Query Prometheus Counters
The server exposes counters that reveal hot paths and systematic failures. Query these endpoints to debug issues in Switchyard at scale:
# Total routing decisions made
curl http://localhost:9090/api/v1/query?query=switchyard_decisions
# Current in-flight algorithm runs (should idle to zero)
curl http://localhost:9090/api/v1/query?query=switchyard_algorithms_in_flight
# LLM call attempts and classifier fallback events
curl http://localhost:9090/api/v1/query?query=switchyard_llm_calls
curl http://localhost:9090/api/v1/query?query=switchyard_classifier_fail_open
5. Validate with the Test Harness
Verify that your telemetry pipeline is functional by running the built-in observability tests:
cargo test --package libsy-llm-client --test observability
This test in crates/libsy-llm-client/tests/observability.rs prints a sample tracing tree and asserts that counters are recorded, confirming your exporter receives data.
6. Enable Server Request Logging
Configure the standalone server to log every inbound request and selected model by setting minimum_severity = "debug" in the server configuration. As documented in the switchyard-server README, this setting correlates specific client requests with span data and Prometheus metrics, allowing you to trace a single HTTP request through the entire routing pipeline.
7. Instrument Individual Algorithm Steps
For deep debugging of specific routing logic, insert temporary debug statements in algorithm implementations such as crates/libsy/src/algorithms/fall_through.rs:
tracing::debug!(
algorithm = self.name,
target = %score.target,
confidence = %score.confidence,
"Model selected in fall-through router"
);
After recompiling, the debug log shows the exact target and confidence scores that led to the routing decision, helping you verify that score_signal and pick_tier functions yield expected values.
Avoiding Common Debugging Pitfalls
Missing OpenTelemetry Provider: If you never install an OpenTelemetry SDK, all metric calls become no-ops, making it appear that no routing activity is occurring. Always install a provider such as opentelemetry-jaeger before initializing Switchyard.
Span Mis-Parenting: The run_span is attached to a task rather than a thread. If you await inside the span without tracing::Instrument, child spans may attach to the wrong executor thread. Use .instrument(span.clone()) on futures as shown in crates/libsy-llm-client/src/run.rs (lines 218-236).
High Cardinality Attributes: Only a bounded set of attributes (algorithm, selected_model, outcome) are recorded by default. Adding dynamic strings (such as raw user prompts) as attributes causes the exporter to drop them. Keep custom data in the span's event payload rather than attributes.
Practical Implementation Examples
Running the Server with Debug Logging
export RUST_LOG=switchyard_server=debug,libsy=debug
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
Expected output includes:
INFO libsy.run algorithm="stage_router" ... outcome_id="ok"
DEBUG libsy.llm_call algorithm="stage_router" selected_model="z-ai/glm-5.2" outcome="ok"
Querying Routing Metrics via Prometheus
# Check for systematic classifier failures
curl http://localhost:9090/api/v1/query?query=sum(rate(switchyard_classifier_fail_open[5m]))
This reveals if the routing system is consistently failing open to default models due to upstream errors.
Summary
- Switchyard exposes OpenTelemetry metrics and tracing spans through the
libsycrate for every routing decision. - Set
RUST_LOG=switchyard_server=debug,libsy=debugto enable verbose logging without code changes. - Install a tracing subscriber and OpenTelemetry SDK to capture
switchyard.runs,switchyard.run_duration_ms, andswitchyard.decisionsin backends like Jaeger or Prometheus. - Use
#[tracing::instrument]spans incrates/libsy/src/core/algorithm.rsto follow requests throughCallModelandDonesteps. - Avoid span mis-parenting by using
.instrument(span.clone())on futures incrates/libsy-llm-client/src/run.rs. - Query
switchyard_algorithms_in_flightandswitchyard_classifier_fail_opencounters to identify overloads and fallback patterns.
Frequently Asked Questions
How do I see which model Switchyard selected for a specific request?
Enable server-side request logging by setting minimum_severity = "debug" in the server configuration, or attach a tracing subscriber in your Rust code. The span created by run_span in crates/libsy/src/observability.rs includes selected_model as an attribute, which appears in logs when RUST_LOG=libsy=debug is set.
Why are my Switchyard metrics not appearing in Prometheus?
This typically occurs when no OpenTelemetry SDK is installed. Without a provider, calls to the global meter in crates/libsy/src/observability.rs become no-ops. Install an exporter like opentelemetry-jaeger or opentelemetry-otlp and call global::set_meter_provider() before executing any Switchyard algorithms.
How can I debug a routing algorithm that picks the wrong model?
Insert tracing::debug! statements in the specific algorithm file (such as crates/libsy/src/algorithms/fall_through.rs) to print intermediate scores and confidences. Recompile with RUST_LOG=libsy=trace to see the decision tree, or query the switchyard_decisions counter to see how often each branch is taken.
What causes child spans to appear disconnected from their parent request?
This is span mis-parenting caused by asynchronous execution. In crates/libsy-llm-client/src/run.rs (lines 218-236), futures must use .instrument(span.clone()) to ensure the span context propagates across await points. Without this, child spans attach to the executor thread rather than the logical task, breaking the trace hierarchy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →