# Prometheus Counters Emitted by Switchyard-Server: Routing Overhead vs Latency Breakdown

> Explore Switchyard-server's Prometheus counters. Understand how switchyard.routing_overhead_ms and switchyard.model_call_latency_ms differentiate routing overhead from model latency for optimized performance.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: performance
- Published: 2026-09-13

---

**Switchyard-server publishes distinct Prometheus histograms that isolate routing overhead (`switchyard.routing_overhead_ms`) from model latency (`switchyard.model_call_latency_ms`), while counters track request attempts, responses, and algorithm-specific events.**

The NVIDIA-NeMo/Switchyard repository provides a high-performance LLM routing layer that exposes detailed Prometheus metrics for observability. Understanding how these Prometheus counters emitted by switchyard-server separate internal routing costs from external model latency is critical for optimizing inference pipelines and debugging performance bottlenecks.

## Request Lifecycle Counters

Switchyard-server tracks the complete request lifecycle through several core counters defined in [`metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/metrics.rs). These metrics capture volume and outcomes without measuring time.

### Upstream and Client Metrics

The `switchyard.upstream_attempts` counter, defined at [lines 119-128 in [`metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/metrics.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L119), increments for every attempt to call an upstream LLM, including retries. It carries labels for `outcome` (e.g., *ok*, *retryable_error*, *other_error*, *client_disconnected*) and HTTP `code`.

Correspondingly, `switchyard.client_responses` ([lines 130-138](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L130)) tracks completed client responses with the same outcome labels, allowing you to correlate internal failures with external results.

The `switchyard.router_retry_recovered` counter ([lines 140-143](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L140)) specifically measures resilience by incrementing when a failed routing decision is successfully recovered through a retry mechanism.

### Error and Token Tracking

Internal errors are captured by `switchyard.errors` in [[`usage_metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/usage_metrics.rs) (lines 9-12)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs#L9), which counts server-side failures during response streaming, labeled by `model`.

Token usage counters—`switchyard.prompt_tokens`, `switchyard.completion_tokens`, `switchyard.cached_tokens`, `switchyard.cache_creation_tokens`, and `switchyard.reasoning_tokens`—are defined at [lines 60-75 in [`usage_metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/usage_metrics.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs#L60). These counters enable cost tracking by model but do not reflect temporal latency.

## Algorithm-Specific Counters

Routing algorithms within switchyard-server emit specialized counters to track decision volumes.

The stage router algorithm increments `switchyard.stage_router.routing_decisions` at [lines 215-267 in [`stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/stage_router.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs#L215), providing a raw count of routing selections without timing data.

For the advisor gate algorithm, [[`advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/advisor_gate.rs) (lines 194-218)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/advisor_gate.rs#L194) defines four counters: `switchyard.advisor_gate.reviews`, `switchyard.advisor_gate.consult_failures`, `switchyard.advisor_gate.discarded_turns`, and `switchyard.advisor_gate.discarded_tokens`. These measure algorithmic activity but, like other counters, do not capture duration.

## Latency Histograms: Separating Routing from Model Time

While counters track event frequency, three distinct histograms separate routing overhead from model inference time.

### The Three Latency Metrics

**`switchyard.routing_overhead_ms`** ([[`metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/metrics.rs) lines 19-25, 85-96](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L19)) records time spent **inside the routing layer** before any LLM is contacted. This includes algorithm evaluation and cache eligibility checks. Buckets range from 0.1 ms to 5 seconds as defined in `ROUTING_OVERHEAD_BUCKETS_MS`.

**`switchyard.model_call_latency_ms`** ([[`metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/metrics.rs) lines 26-34](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L26)) captures the duration of the **actual LLM call**, from request dispatch to response receipt. This uses `LLM_LATENCY_BUCKETS_MS` ranging from 0 ms to 5 minutes.

**`switchyard.total_latency_ms`** ([[`metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/metrics.rs) lines 99-112](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs#L99)) measures **end-to-end latency** from initial client request to final response, using the same buckets as the model latency histogram.

### How Differentiation Works

The separation occurs through careful timing instrumentation in the request path:

- **Routing overhead** is measured and recorded in `routing_overhead_ms` **before** any upstream model is contacted. This isolates internal decision-making costs.
- **Model latency** is recorded only during the window when the request is actively being processed by the upstream LLM.
- **Total latency** encompasses both phases plus network serialization overhead.

Because `routing_overhead_ms` and `total_latency_ms` share compatible bucket boundaries, you can compute the routing fraction by comparing histograms or querying both metrics to determine what percentage of request time is spent in switchyard-server's logic versus waiting for model inference.

## Implementation Examples

The following Rust patterns demonstrate how these metrics are instrumented throughout the codebase.

Recording routing decisions:

```rust
// From stage_router.rs - increment routing decision counter
global::meter("switchyard")
    .u64_counter("switchyard.stage_router.routing_decisions")
    .build()
    .add(1, &[]);

```

Capturing timing data:

```rust
// Record routing overhead (measured before LLM contact)
global::meter("switchyard")
    .f64_histogram("switchyard.routing_overhead_ms")
    .build()
    .record(routing_duration_ms, &[]);

// Record model call latency (measured during upstream call)
global::meter("switchyard")
    .f64_histogram("switchyard.model_call_latency_ms")
    .build()
    .record(model_latency_ms, &[]);

```

Tracking client outcomes:

```rust
// Increment client response counter with outcome label
global::meter("switchyard")
    .u64_counter("switchyard.client_responses")
    .build()
    .add(1, &[KeyValue::new("outcome", "ok")]);

```

When the server is running, scrape metrics via the `/metrics` endpoint (typically on port 8080):

```bash
curl http://localhost:8080/metrics

```

Example output showing the histogram differentiation:

```text

# HELP switchyard_routing_overhead_ms Time spent in routing layer before LLM call

# TYPE switchyard_routing_overhead_ms histogram

switchyard_routing_overhead_ms_bucket{le="0.1"} 12
switchyard_routing_overhead_ms_bucket{le="0.25"} 30

# HELP switchyard_model_call_latency_ms Latency of upstream LLM calls

# TYPE switchyard_model_call_latency_ms histogram

switchyard_model_call_latency_ms_bucket{le="1000"} 45
switchyard_model_call_latency_ms_bucket{le="5000"} 120

# HELP switchyard_total_latency_ms End-to-end request latency

# TYPE switchyard_total_latency_ms histogram

switchyard_total_latency_ms_bucket{le="1000"} 40
switchyard_total_latency_ms_bucket{le="5000"} 125

```

## Summary

- Switchyard-server emits **counters** for request volume (`upstream_attempts`, `client_responses`), algorithm activity (`routing_decisions`, `advisor_gate` metrics), and resource usage (token counters).
- **Routing overhead** is isolated in the `switchyard.routing_overhead_ms` histogram, measured before any model contact.
- **Model latency** is captured separately in `switchyard.model_call_latency_ms`, covering only the upstream LLM call duration.
- **Total latency** in `switchyard.total_latency_ms` provides the aggregate view, enabling operators to calculate the proportion of time spent in routing versus inference by comparing histograms.
- Source files including [`metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/metrics.rs), [`usage_metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/usage_metrics.rs), [`stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/stage_router.rs), and [`advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/advisor_gate.rs) define these instruments using the OpenTelemetry Prometheus exporter.

## Frequently Asked Questions

### How can I calculate the percentage of time spent in routing versus model inference?

Query both `switchyard.routing_overhead_ms` and `switchyard.total_latency_ms` histograms for your time window. Because both use compatible bucket boundaries, you can sum the `routing_overhead_ms` observations and divide by the `total_latency_ms` observations to derive the routing overhead percentage. Alternatively, subtract `model_call_latency_ms` from `total_latency_ms` to isolate non-model time.

### What is the difference between `switchyard.upstream_attempts` and `switchyard.client_responses`?

`switchyard.upstream_attempts` increments every time switchyard-server tries to call a backend LLM, including retries after failures. `switchyard.client_responses` increments only when the server sends the final response back to the client. Comparing these counters reveals how often internal retries are required before successfully serving a client.

### Why are routing decisions tracked as counters instead of histograms?

The `switchyard.stage_router.routing_decisions` and advisor gate counters measure **volume** of algorithmic work (how many decisions were made), not duration. The **time** spent making those decisions is captured separately in the `routing_overhead_ms` histogram. This separation allows operators to correlate high decision volumes with increased routing latency while keeping the metrics orthogonal.

### Which file contains the Prometheus registry setup for switchyard-server?

The main Prometheus registry and histogram definitions reside in [[`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs), while token-specific and error counters are defined in [[`crates/switchyard-server/src/usage_metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs). The `/metrics` HTTP endpoint is registered in the server's [`lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/lib.rs) to expose these metrics for scraping.