# How to Debug Issues in Switchyard: OpenTelemetry Metrics and Tracing Guide

> Debug Switchyard issues effectively. Utilize OpenTelemetry metrics and tracing to inspect routing and model selections in real-time. Enable verbose logging for deeper insights.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Enable verbose tracing with `RUST_LOG=switchyard_server=debug,libsy=debug`, attach a tracing subscriber, and query OpenTelemetry metrics to inspect routing decisions and model selections in real-time.**

Switchyard is a Rust-based LLM routing engine built from modular crates that expose OpenTelemetry metrics and tracing spans for every routing decision. To effectively debug issues in Switchyard, you configure environment-based logging, attach tracing subscribers, and analyze exported telemetry data to see exactly how algorithms behave, why specific models were chosen, and where failures occurred. The system records all activity through the `libsy` library, making debugging possible without modifying core routing logic.

## Understanding Switchyard's Observability Architecture

### The Core libsy Layer

The `libsy` crate in [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs) provides the central telemetry functions that capture algorithm activity. The `run_span` function (lines 68-94) creates a span that captures request metadata including model ID, session ID, and correlation ID. Helper functions like `record_outcome`, `record_llm_call`, and `record_decision` (lines 16-33 and 81-89) add attributes such as selected models, evidence scores, and token usage to the current span. The `observe_run` wrapper (lines 83-98) counts in-flight runs and measures execution timing.

All these functions use the **global OpenTelemetry meter** under the `switchyard` scope. If no SDK is installed, the calls are no-ops, allowing you to debug issues in Switchyard without changing any library code.

### Algorithm Tracing Instrumentation

The algorithm entry point at `Algorithm::run_stream` in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) (lines 170-182) is annotated with `#[tracing::instrument]`. This macro automatically creates a child span for each routing step—such as `CallModel` and `Done`—and propagates the parent `libsy.run` span created by `run_span`. This hierarchical structure lets you trace a single request through every decision point in the routing tree.

### Server-Side Telemetry

When running the standalone `switchyard-server`, the server adds HTTP handling spans and publishes Prometheus counters. The server's observability module in [`crates/switchyard-server/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/observability.rs) (lines 45-55) records request-level errors and OpenTelemetry exporter status. The server initialization in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) (lines 47-56) shows how the global subscriber is installed and how the metrics endpoint is exposed.

## Step-by-Step Debugging Workflow

### 1. Enable Verbose Environment Logging

Set the `RUST_LOG` environment variable to activate debug output before launching the server or your Rust binary:

```bash
export RUST_LOG=switchyard_server=debug,libsy=debug

```

For maximum detail during complex routing investigations, use `trace` instead of `debug`. This setting logs routing decisions, classifier failures, and metric increments to stderr.

### 2. Attach a Tracing Subscriber

In your Rust application, initialize a subscriber to print the spans generated by `run_span` and the algorithm's instrumentation macros:

```rust
use tracing_subscriber::{fmt, EnvFilter};

fn main() {
    tracing_subscriber::registry()
        .with(fmt::layer())
        .with(EnvFilter::from_default_env())
        .init();
    
    // Switchyard algorithm execution begins here
}

```

This subscriber reads the `RUST_LOG` variable and formats the span hierarchy, showing exactly which algorithm step is executing.

### 3. Configure OpenTelemetry Exporters

To persist metrics beyond stdout, install an OpenTelemetry SDK before calling any Switchyard code:

```rust
use opentelemetry::global;
use opentelemetry_jaeger::JaegerPipelineBuilder;

fn init_otel() {
    JaegerPipelineBuilder::default()
        .with_service_name("switchyard")
        .install_simple()
        .expect("Jaeger exporter failed");
    
    global::set_meter_provider(opentelemetry::sdk::metrics::MeterProvider::builder().build());
}

```

After initialization, metrics such as `switchyard.run_duration_ms`, `switchyard.decisions`, and `switchyard.runs` become visible in your observability backend.

### 4. Query Prometheus Counters

The server exposes counters that reveal hot paths and systematic failures. Query these endpoints to debug issues in Switchyard at scale:

```bash

# Total routing decisions made

curl http://localhost:9090/api/v1/query?query=switchyard_decisions

# Current in-flight algorithm runs (should idle to zero)

curl http://localhost:9090/api/v1/query?query=switchyard_algorithms_in_flight

# LLM call attempts and classifier fallback events

curl http://localhost:9090/api/v1/query?query=switchyard_llm_calls
curl http://localhost:9090/api/v1/query?query=switchyard_classifier_fail_open

```

### 5. Validate with the Test Harness

Verify that your telemetry pipeline is functional by running the built-in observability tests:

```bash
cargo test --package libsy-llm-client --test observability

```

This test in [`crates/libsy-llm-client/tests/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/tests/observability.rs) prints a sample tracing tree and asserts that counters are recorded, confirming your exporter receives data.

### 6. Enable Server Request Logging

Configure the standalone server to log every inbound request and selected model by setting `minimum_severity = "debug"` in the server configuration. As documented in the `switchyard-server` README, this setting correlates specific client requests with span data and Prometheus metrics, allowing you to trace a single HTTP request through the entire routing pipeline.

### 7. Instrument Individual Algorithm Steps

For deep debugging of specific routing logic, insert temporary debug statements in algorithm implementations such as [`crates/libsy/src/algorithms/fall_through.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/fall_through.rs):

```rust
tracing::debug!(
    algorithm = self.name,
    target = %score.target,
    confidence = %score.confidence,
    "Model selected in fall-through router"
);

```

After recompiling, the debug log shows the exact target and confidence scores that led to the routing decision, helping you verify that `score_signal` and `pick_tier` functions yield expected values.

## Avoiding Common Debugging Pitfalls

**Missing OpenTelemetry Provider:** If you never install an OpenTelemetry SDK, all metric calls become no-ops, making it appear that no routing activity is occurring. Always install a provider such as `opentelemetry-jaeger` before initializing Switchyard.

**Span Mis-Parenting:** The `run_span` is attached to a task rather than a thread. If you `await` inside the span without `tracing::Instrument`, child spans may attach to the wrong executor thread. Use `.instrument(span.clone())` on futures as shown in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs) (lines 218-236).

**High Cardinality Attributes:** Only a bounded set of attributes (`algorithm`, `selected_model`, `outcome`) are recorded by default. Adding dynamic strings (such as raw user prompts) as attributes causes the exporter to drop them. Keep custom data in the span's event payload rather than attributes.

## Practical Implementation Examples

### Running the Server with Debug Logging

```bash
export RUST_LOG=switchyard_server=debug,libsy=debug
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

Expected output includes:

```

INFO  libsy.run algorithm="stage_router" ... outcome_id="ok"
DEBUG libsy.llm_call algorithm="stage_router" selected_model="z-ai/glm-5.2" outcome="ok"

```

### Querying Routing Metrics via Prometheus

```bash

# Check for systematic classifier failures

curl http://localhost:9090/api/v1/query?query=sum(rate(switchyard_classifier_fail_open[5m]))

```

This reveals if the routing system is consistently failing open to default models due to upstream errors.

## Summary

- Switchyard exposes **OpenTelemetry metrics** and **tracing spans** through the `libsy` crate for every routing decision.
- Set `RUST_LOG=switchyard_server=debug,libsy=debug` to enable verbose logging without code changes.
- Install a tracing subscriber and OpenTelemetry SDK to capture `switchyard.runs`, `switchyard.run_duration_ms`, and `switchyard.decisions` in backends like Jaeger or Prometheus.
- Use `#[tracing::instrument]` spans in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) to follow requests through `CallModel` and `Done` steps.
- Avoid span mis-parenting by using `.instrument(span.clone())` on futures in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs).
- Query `switchyard_algorithms_in_flight` and `switchyard_classifier_fail_open` counters to identify overloads and fallback patterns.

## Frequently Asked Questions

### How do I see which model Switchyard selected for a specific request?

Enable server-side request logging by setting `minimum_severity = "debug"` in the server configuration, or attach a tracing subscriber in your Rust code. The span created by `run_span` in [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs) includes `selected_model` as an attribute, which appears in logs when `RUST_LOG=libsy=debug` is set.

### Why are my Switchyard metrics not appearing in Prometheus?

This typically occurs when no OpenTelemetry SDK is installed. Without a provider, calls to the global meter in [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs) become no-ops. Install an exporter like `opentelemetry-jaeger` or `opentelemetry-otlp` and call `global::set_meter_provider()` before executing any Switchyard algorithms.

### How can I debug a routing algorithm that picks the wrong model?

Insert `tracing::debug!` statements in the specific algorithm file (such as [`crates/libsy/src/algorithms/fall_through.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/fall_through.rs)) to print intermediate scores and confidences. Recompile with `RUST_LOG=libsy=trace` to see the decision tree, or query the `switchyard_decisions` counter to see how often each branch is taken.

### What causes child spans to appear disconnected from their parent request?

This is span mis-parenting caused by asynchronous execution. In [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs) (lines 218-236), futures must use `.instrument(span.clone())` to ensure the span context propagates across await points. Without this, child spans attach to the executor thread rather than the logical task, breaking the trace hierarchy.