How Switchyard Collects and Exposes Prometheus Metrics: A Deep Dive into the Rust Implementation

Switchyard collects and exposes Prometheus metrics through an OpenTelemetry-based registry initialized at server startup, exposing them via a native Rust Axum endpoint at /metrics that renders histograms, counters, and gauges in Prometheus text format.

The NVIDIA-NeMo/Switchyard repository implements a high-performance routing layer for LLM inference that relies on comprehensive observability to monitor system health. Understanding how Switchyard collects and exposes Prometheus metrics is essential for operators who need to track routing overhead, upstream attempt outcomes, and latency distributions in production environments.

Architecture Overview

Switchyard’s observability stack is built on OpenTelemetry with a Prometheus exporter embedded directly into the native Rust server (switchyard-server). This architecture eliminates external sidecar dependencies, reducing operational complexity while maintaining compatibility with standard Prometheus scraping protocols.

OpenTelemetry Integration

The implementation uses the opentelemetry crate alongside prometheus to create a unified metrics pipeline. When configured, Switchyard can simultaneously export metrics via the Prometheus pull model and optionally push metrics via OTLP if the METRICS OTLP flag is enabled.

Metric Registry Initialization

The foundation of Switchyard’s metrics collection resides in crates/switchyard-server/src/metrics.rs, which establishes a global registry during server startup.

Global Registry Setup in metrics.rs

When the server initializes, the code creates a global prometheus::Registry and an SdkMeterProvider. The registry is stored in a OnceLock static variable named METRICS, ensuring thread-safe, exactly-once initialization across asynchronous workers. The provider is configured with custom bucket views specifically tailored for routing-overhead and LLM-latency histograms, allowing precise latency tracking at millisecond granularity.

Seeding Default Counters with seed_outcome_metrics()

To prevent "missing metric" gaps in dashboards, Switchyard proactively seeds counters before traffic arrives. The seed_outcome_metrics() function pre-creates essential counters including:

  • switchyard.upstream_attempts
  • switchyard.client_responses
  • switchyard.router_retry_recovered

This initialization guarantees that Prometheus queries return zero values rather than gaps when no events have occurred yet.

Runtime Metric Collection

Once the registry is established, various components throughout the codebase record events using helper functions defined in the metrics module.

Recording Client Responses

Throughout request processing, Switchyard invokes record_client_response(status) to increment the switchyard.client_responses counter with appropriate outcome labels. This captures the final disposition of every client request, enabling error-rate calculations and throughput monitoring.

LLM Client Metrics

The LLM client implementation (switchyard_llm_client) registers additional counters when initialize_metrics() is invoked. These metrics track upstream communication health separately from the router's internal operations, providing granular visibility into backend inference latency and connection attempts.

Exposing Metrics via HTTP Endpoint

Switchyard serves metrics through a dedicated HTTP endpoint integrated into its Axum-based web server.

The /metrics Route in lib.rs

In crates/switchyard-server/src/lib.rs, the server registers an Axum route at GET /metrics. This endpoint is automatically available when starting the native server without requiring additional configuration flags.

Prometheus Text Encoding

The route handler invokes metrics::encode(&registry), which utilizes prometheus::TextEncoder to serialize the current metric snapshot. As implemented in crates/switchyard-server/src/metrics.rs (lines 36-47), the encoder returns data with the content-type header text/plain; version=0.0.4; charset=utf-8, adhering to the Prometheus exposition format specification.

Scraping Switchyard Metrics

To integrate with your monitoring stack, configure Prometheus to scrape the Switchyard endpoint.

Start the server with metrics enabled:

switchyard-server --config routes.toml --port 4000

Verify the endpoint returns metrics:

curl -s http://localhost:4000/metrics | head -n 20

Example output includes histogram buckets for routing overhead and counters for client responses:


# HELP switchyard.routing_overhead_ms Routing overhead in milliseconds

# TYPE switchyard.routing_overhead_ms histogram

switchyard.routing_overhead_ms_bucket{le="0.1"} 0
switchyard.routing_overhead_ms_bucket{le="0.25"} 5

# HELP switchyard.client_responses Total client responses

# TYPE switchyard.client_responses counter

switchyard.client_responses{status="success"} 42

Configure your prometheus.yml:

scrape_configs:
  - job_name: 'switchyard'
    static_configs:
      - targets: ['localhost:4000']
    metrics_path: /metrics

Summary

  • Switchyard initializes a global prometheus::Registry stored in a OnceLock during server startup in crates/switchyard-server/src/metrics.rs.
  • The seed_outcome_metrics() function pre-creates counters like switchyard.client_responses to ensure dashboard visibility before traffic arrives.
  • Runtime events are recorded via helpers such as record_client_response(), while the LLM client initializes separate metrics through initialize_metrics().
  • The Axum server exposes metrics at GET /metrics using prometheus::TextEncoder to render OpenTelemetry data in standard Prometheus text format.
  • Operators scrape http://<host>:<port>/metrics to collect histograms for routing overhead, LLM latency, and counter-based outcome tracking.

Frequently Asked Questions

How does Switchyard ensure metrics are available immediately after startup?

Switchyard calls seed_outcome_metrics() during initialization to pre-register counters including switchyard.upstream_attempts and switchyard.router_retry_recovered. This prevents the "missing metric" problem where Prometheus queries return no data until the first event occurs.

Can Switchyard export metrics via OTLP in addition to Prometheus?

Yes. While the Prometheus exporter runs by default in the native Rust server, you can enable OTLP export by setting the METRICS OTLP flag. This allows the SdkMeterProvider to configure an OTLP exporter alongside the Prometheus registry for hybrid observability pipelines.

What latency distributions does Switchyard track?

Switchyard configures custom bucket views for two critical histograms: routing-overhead (measuring internal routing decision time) and LLM-latency (tracking backend inference duration). These are defined in crates/switchyard-server/src/metrics.rs during the SdkMeterProvider setup.

Where is the /metrics endpoint handler defined?

The Axum route for GET /metrics is registered in crates/switchyard-server/src/lib.rs. The handler calls metrics::encode(&registry), which uses prometheus::TextEncoder to generate the text exposition format at lines 36-47 of crates/switchyard-server/src/metrics.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →