Prometheus Counters Emitted by Switchyard-Server: Routing Overhead vs Latency Breakdown

Switchyard-server publishes distinct Prometheus histograms that isolate routing overhead (switchyard.routing_overhead_ms) from model latency (switchyard.model_call_latency_ms), while counters track request attempts, responses, and algorithm-specific events.

The NVIDIA-NeMo/Switchyard repository provides a high-performance LLM routing layer that exposes detailed Prometheus metrics for observability. Understanding how these Prometheus counters emitted by switchyard-server separate internal routing costs from external model latency is critical for optimizing inference pipelines and debugging performance bottlenecks.

Request Lifecycle Counters

Switchyard-server tracks the complete request lifecycle through several core counters defined in metrics.rs. These metrics capture volume and outcomes without measuring time.

Upstream and Client Metrics

The switchyard.upstream_attempts counter, defined at [lines 119-128 in metrics.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L119), increments for every attempt to call an upstream LLM, including retries. It carries labels for outcome (e.g., ok, retryable_error, other_error, client_disconnected) and HTTP code.

Correspondingly, switchyard.client_responses (lines 130-138) tracks completed client responses with the same outcome labels, allowing you to correlate internal failures with external results.

The switchyard.router_retry_recovered counter (lines 140-143) specifically measures resilience by incrementing when a failed routing decision is successfully recovered through a retry mechanism.

Error and Token Tracking

Internal errors are captured by switchyard.errors in [usage_metrics.rs (lines 9-12)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs#L9), which counts server-side failures during response streaming, labeled by model.

Token usage counters—switchyard.prompt_tokens, switchyard.completion_tokens, switchyard.cached_tokens, switchyard.cache_creation_tokens, and switchyard.reasoning_tokens—are defined at [lines 60-75 in usage_metrics.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs#L60). These counters enable cost tracking by model but do not reflect temporal latency.

Algorithm-Specific Counters

Routing algorithms within switchyard-server emit specialized counters to track decision volumes.

The stage router algorithm increments switchyard.stage_router.routing_decisions at [lines 215-267 in stage_router.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs#L215), providing a raw count of routing selections without timing data.

For the advisor gate algorithm, [advisor_gate.rs (lines 194-218)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/advisor_gate.rs#L194) defines four counters: switchyard.advisor_gate.reviews, switchyard.advisor_gate.consult_failures, switchyard.advisor_gate.discarded_turns, and switchyard.advisor_gate.discarded_tokens. These measure algorithmic activity but, like other counters, do not capture duration.

Latency Histograms: Separating Routing from Model Time

While counters track event frequency, three distinct histograms separate routing overhead from model inference time.

The Three Latency Metrics

switchyard.routing_overhead_ms ([metrics.rs lines 19-25, 85-96](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L19)) records time spent inside the routing layer before any LLM is contacted. This includes algorithm evaluation and cache eligibility checks. Buckets range from 0.1 ms to 5 seconds as defined in ROUTING_OVERHEAD_BUCKETS_MS.

switchyard.model_call_latency_ms ([metrics.rs lines 26-34](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L26)) captures the duration of the actual LLM call, from request dispatch to response receipt. This uses LLM_LATENCY_BUCKETS_MS ranging from 0 ms to 5 minutes.

switchyard.total_latency_ms ([metrics.rs lines 99-112](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L99)) measures end-to-end latency from initial client request to final response, using the same buckets as the model latency histogram.

How Differentiation Works

The separation occurs through careful timing instrumentation in the request path:

  • Routing overhead is measured and recorded in routing_overhead_ms before any upstream model is contacted. This isolates internal decision-making costs.
  • Model latency is recorded only during the window when the request is actively being processed by the upstream LLM.
  • Total latency encompasses both phases plus network serialization overhead.

Because routing_overhead_ms and total_latency_ms share compatible bucket boundaries, you can compute the routing fraction by comparing histograms or querying both metrics to determine what percentage of request time is spent in switchyard-server's logic versus waiting for model inference.

Implementation Examples

The following Rust patterns demonstrate how these metrics are instrumented throughout the codebase.

Recording routing decisions:

// From stage_router.rs - increment routing decision counter
global::meter("switchyard")
    .u64_counter("switchyard.stage_router.routing_decisions")
    .build()
    .add(1, &[]);

Capturing timing data:

// Record routing overhead (measured before LLM contact)
global::meter("switchyard")
    .f64_histogram("switchyard.routing_overhead_ms")
    .build()
    .record(routing_duration_ms, &[]);

// Record model call latency (measured during upstream call)
global::meter("switchyard")
    .f64_histogram("switchyard.model_call_latency_ms")
    .build()
    .record(model_latency_ms, &[]);

Tracking client outcomes:

// Increment client response counter with outcome label
global::meter("switchyard")
    .u64_counter("switchyard.client_responses")
    .build()
    .add(1, &[KeyValue::new("outcome", "ok")]);

When the server is running, scrape metrics via the /metrics endpoint (typically on port 8080):

curl http://localhost:8080/metrics

Example output showing the histogram differentiation:


# HELP switchyard_routing_overhead_ms Time spent in routing layer before LLM call

# TYPE switchyard_routing_overhead_ms histogram

switchyard_routing_overhead_ms_bucket{le="0.1"} 12
switchyard_routing_overhead_ms_bucket{le="0.25"} 30

# HELP switchyard_model_call_latency_ms Latency of upstream LLM calls

# TYPE switchyard_model_call_latency_ms histogram

switchyard_model_call_latency_ms_bucket{le="1000"} 45
switchyard_model_call_latency_ms_bucket{le="5000"} 120

# HELP switchyard_total_latency_ms End-to-end request latency

# TYPE switchyard_total_latency_ms histogram

switchyard_total_latency_ms_bucket{le="1000"} 40
switchyard_total_latency_ms_bucket{le="5000"} 125

Summary

  • Switchyard-server emits counters for request volume (upstream_attempts, client_responses), algorithm activity (routing_decisions, advisor_gate metrics), and resource usage (token counters).
  • Routing overhead is isolated in the switchyard.routing_overhead_ms histogram, measured before any model contact.
  • Model latency is captured separately in switchyard.model_call_latency_ms, covering only the upstream LLM call duration.
  • Total latency in switchyard.total_latency_ms provides the aggregate view, enabling operators to calculate the proportion of time spent in routing versus inference by comparing histograms.
  • Source files including metrics.rs, usage_metrics.rs, stage_router.rs, and advisor_gate.rs define these instruments using the OpenTelemetry Prometheus exporter.

Frequently Asked Questions

How can I calculate the percentage of time spent in routing versus model inference?

Query both switchyard.routing_overhead_ms and switchyard.total_latency_ms histograms for your time window. Because both use compatible bucket boundaries, you can sum the routing_overhead_ms observations and divide by the total_latency_ms observations to derive the routing overhead percentage. Alternatively, subtract model_call_latency_ms from total_latency_ms to isolate non-model time.

What is the difference between switchyard.upstream_attempts and switchyard.client_responses?

switchyard.upstream_attempts increments every time switchyard-server tries to call a backend LLM, including retries after failures. switchyard.client_responses increments only when the server sends the final response back to the client. Comparing these counters reveals how often internal retries are required before successfully serving a client.

Why are routing decisions tracked as counters instead of histograms?

The switchyard.stage_router.routing_decisions and advisor gate counters measure volume of algorithmic work (how many decisions were made), not duration. The time spent making those decisions is captured separately in the routing_overhead_ms histogram. This separation allows operators to correlate high decision volumes with increased routing latency while keeping the metrics orthogonal.

Which file contains the Prometheus registry setup for switchyard-server?

The main Prometheus registry and histogram definitions reside in [crates/switchyard-server/src/metrics.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs), while token-specific and error counters are defined in [crates/switchyard-server/src/usage_metrics.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs). The /metrics HTTP endpoint is registered in the server's lib.rs to expose these metrics for scraping.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →