# How Switchyard Collects and Exposes Prometheus Metrics: A Deep Dive into the Rust Implementation

> Learn how Switchyard collects and exposes Prometheus metrics using OpenTelemetry and an Axum endpoint. Discover its Rust implementation for histograms, counters, and gauges.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-08-21

---

**Switchyard collects and exposes Prometheus metrics through an OpenTelemetry-based registry initialized at server startup, exposing them via a native Rust Axum endpoint at `/metrics` that renders histograms, counters, and gauges in Prometheus text format.**

The **NVIDIA-NeMo/Switchyard** repository implements a high-performance routing layer for LLM inference that relies on comprehensive observability to monitor system health. Understanding how Switchyard collects and exposes Prometheus metrics is essential for operators who need to track routing overhead, upstream attempt outcomes, and latency distributions in production environments.

## Architecture Overview

Switchyard’s observability stack is built on **OpenTelemetry** with a **Prometheus exporter** embedded directly into the native Rust server (`switchyard-server`). This architecture eliminates external sidecar dependencies, reducing operational complexity while maintaining compatibility with standard Prometheus scraping protocols.

### OpenTelemetry Integration

The implementation uses the `opentelemetry` crate alongside `prometheus` to create a unified metrics pipeline. When configured, Switchyard can simultaneously export metrics via the Prometheus pull model and optionally push metrics via OTLP if the `METRICS` OTLP flag is enabled.

## Metric Registry Initialization

The foundation of Switchyard’s metrics collection resides in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs), which establishes a global registry during server startup.

### Global Registry Setup in [`metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/metrics.rs)

When the server initializes, the code creates a global `prometheus::Registry` and an `SdkMeterProvider`. The registry is stored in a `OnceLock` static variable named `METRICS`, ensuring thread-safe, exactly-once initialization across asynchronous workers. The provider is configured with custom bucket views specifically tailored for **routing-overhead** and **LLM-latency** histograms, allowing precise latency tracking at millisecond granularity.

### Seeding Default Counters with `seed_outcome_metrics()`

To prevent "missing metric" gaps in dashboards, Switchyard proactively seeds counters before traffic arrives. The `seed_outcome_metrics()` function pre-creates essential counters including:

- `switchyard.upstream_attempts`
- `switchyard.client_responses`
- `switchyard.router_retry_recovered`

This initialization guarantees that Prometheus queries return zero values rather than gaps when no events have occurred yet.

## Runtime Metric Collection

Once the registry is established, various components throughout the codebase record events using helper functions defined in the metrics module.

### Recording Client Responses

Throughout request processing, Switchyard invokes `record_client_response(status)` to increment the `switchyard.client_responses` counter with appropriate outcome labels. This captures the final disposition of every client request, enabling error-rate calculations and throughput monitoring.

### LLM Client Metrics

The LLM client implementation (`switchyard_llm_client`) registers additional counters when `initialize_metrics()` is invoked. These metrics track upstream communication health separately from the router's internal operations, providing granular visibility into backend inference latency and connection attempts.

## Exposing Metrics via HTTP Endpoint

Switchyard serves metrics through a dedicated HTTP endpoint integrated into its Axum-based web server.

### The `/metrics` Route in [`lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/lib.rs)

In [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), the server registers an Axum route at **`GET /metrics`**. This endpoint is automatically available when starting the native server without requiring additional configuration flags.

### Prometheus Text Encoding

The route handler invokes `metrics::encode(&registry)`, which utilizes `prometheus::TextEncoder` to serialize the current metric snapshot. As implemented in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs) (lines 36-47), the encoder returns data with the content-type header `text/plain; version=0.0.4; charset=utf-8`, adhering to the Prometheus exposition format specification.

## Scraping Switchyard Metrics

To integrate with your monitoring stack, configure Prometheus to scrape the Switchyard endpoint.

Start the server with metrics enabled:

```bash
switchyard-server --config routes.toml --port 4000

```

Verify the endpoint returns metrics:

```bash
curl -s http://localhost:4000/metrics | head -n 20

```

Example output includes histogram buckets for routing overhead and counters for client responses:

```text

# HELP switchyard.routing_overhead_ms Routing overhead in milliseconds

# TYPE switchyard.routing_overhead_ms histogram

switchyard.routing_overhead_ms_bucket{le="0.1"} 0
switchyard.routing_overhead_ms_bucket{le="0.25"} 5

# HELP switchyard.client_responses Total client responses

# TYPE switchyard.client_responses counter

switchyard.client_responses{status="success"} 42

```

Configure your [`prometheus.yml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/prometheus.yml):

```yaml
scrape_configs:
  - job_name: 'switchyard'
    static_configs:
      - targets: ['localhost:4000']
    metrics_path: /metrics

```

## Summary

- Switchyard initializes a global `prometheus::Registry` stored in a `OnceLock` during server startup in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs).
- The `seed_outcome_metrics()` function pre-creates counters like `switchyard.client_responses` to ensure dashboard visibility before traffic arrives.
- Runtime events are recorded via helpers such as `record_client_response()`, while the LLM client initializes separate metrics through `initialize_metrics()`.
- The Axum server exposes metrics at `GET /metrics` using `prometheus::TextEncoder` to render OpenTelemetry data in standard Prometheus text format.
- Operators scrape `http://<host>:<port>/metrics` to collect histograms for routing overhead, LLM latency, and counter-based outcome tracking.

## Frequently Asked Questions

### How does Switchyard ensure metrics are available immediately after startup?

Switchyard calls `seed_outcome_metrics()` during initialization to pre-register counters including `switchyard.upstream_attempts` and `switchyard.router_retry_recovered`. This prevents the "missing metric" problem where Prometheus queries return no data until the first event occurs.

### Can Switchyard export metrics via OTLP in addition to Prometheus?

Yes. While the Prometheus exporter runs by default in the native Rust server, you can enable OTLP export by setting the `METRICS` OTLP flag. This allows the `SdkMeterProvider` to configure an OTLP exporter alongside the Prometheus registry for hybrid observability pipelines.

### What latency distributions does Switchyard track?

Switchyard configures custom bucket views for two critical histograms: **routing-overhead** (measuring internal routing decision time) and **LLM-latency** (tracking backend inference duration). These are defined in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs) during the `SdkMeterProvider` setup.

### Where is the `/metrics` endpoint handler defined?

The Axum route for `GET /metrics` is registered in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs). The handler calls `metrics::encode(&registry)`, which uses `prometheus::TextEncoder` to generate the text exposition format at lines 36-47 of [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs).