# How to Configure Prometheus Metrics and Enable the /metrics and /v1/stats Endpoints in Switchyard

> Configure Switchyard for Prometheus metrics and enable /metrics and /v1/stats endpoints. Simply set metrics and stats to true in your TOML file and restart.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Enable Prometheus metrics and the JSON statistics API in Switchyard by setting `metrics = true` and `stats = true` in your server configuration TOML file, then restart the server to expose `/metrics` and `/v1/stats` on the configured address.**

Switchyard is an open-source LLM routing and dispatch framework maintained by NVIDIA. To effectively monitor production deployments, you must configure Prometheus metrics and enable the `/metrics` and `/v1/stats` endpoints in Switchyard, which expose process-level OpenTelemetry data and routing algorithm statistics respectively.

## Enable Observability via TOML Configuration

The Switchyard server reads observability settings from the `[server]` table and the optional `[observability]` table in your configuration file. In [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), the `ServerConfig` struct parses the `metrics` and `stats` boolean flags to conditionally register the Axum routes for each endpoint.

### Core Configuration Flags

Add these keys to your [`config.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/config.toml) to enable the endpoints:

- `metrics`: Enables the Prometheus text endpoint at `GET /metrics`
- `stats`: Enables the JSON statistics endpoint at `GET /v1/stats`

```toml
[server]
listen = "127.0.0.1:8000"
metrics = true
stats = true

```

### Optional Observability Settings

You can further customize behavior under the `[observability]` table:

- `metrics_addr`: Binds the metrics endpoint to a separate socket address (defaults to the main `listen` address)
- `metric_namespace`: Sets a custom prefix for all Prometheus metrics (defaults to `switchyard`)

```toml
[server]
listen = "0.0.0.0:8080"
metrics = true
stats = true
metrics_addr = "127.0.0.1:9100"

[observability]
metric_namespace = "my_switchyard"

```

## How the Endpoints Work

Understanding the implementation details helps when integrating with scrapers and dashboards.

### Prometheus Metrics Endpoint (/metrics)

When `metrics = true`, the server registers a handler defined in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs). The `prometheus_metrics` function pulls data from the process-wide OpenTelemetry registry and renders it in Prometheus exposition format.

```bash
curl http://localhost:8080/metrics

```

The endpoint returns counters, gauges, and histograms collected by Switchyard's internal instrumentation, prefixed according to your `metric_namespace` setting.

### Statistics API Endpoint (/v1/stats)

When `stats = true`, the server exposes `GET /v1/stats` as implemented in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs). This endpoint returns a JSON payload containing algorithm-specific routing decisions and global request statistics.

```json
{
  "algorithm_stats": {
    "stage_router": {
      "routing_decisions": {
        "override": { "targets": { "model/fast": 12 } },
        "default": { "targets": { "model/slow": 23 } }
      }
    }
  },
  "global_stats": {
    "requests": 45,
    "errors": 3,
    "latency_ms": 210.4
  }
}

```

## Practical Configuration Examples

### Basic Enablement

Create a minimal [`config.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/config.toml) to enable both endpoints on the default port:

```toml
[server]
listen = "0.0.0.0:8080"
metrics = true
stats = true

```

Run the server:

```bash
switchyard-runner --config config.toml

```

### Isolated Metrics Port

For security, bind metrics to localhost only while exposing the main API externally:

```toml
[server]
listen = "0.0.0.0:8080"
metrics = true
stats = true
metrics_addr = "127.0.0.1:9090"

```

Now Prometheus scrapes `127.0.0.1:9090/metrics` while application traffic uses port 8080.

### Accessing Stats via Python

If using the Switchyard Python bindings, fetch statistics programmatically:

```python
import httpx

response = httpx.get("http://localhost:8080/v1/stats")
data = response.json()
print(data["global_stats"]["requests"])

```

## Verification Steps

After restarting the server with the new configuration, verify endpoints respond correctly:

1. Check Prometheus metrics are exposed:

```bash
curl -s http://127.0.0.1:8080/metrics | grep switchyard_requests_total

```

2. Validate JSON statistics return proper schema:

```bash
curl -s http://127.0.0.1:8080/v1/stats | jq '.algorithm_stats'

```

Both commands should return non-empty, properly formatted data confirming that you successfully configured Prometheus metrics and enabled the `/metrics` and `/v1/stats` endpoints in Switchyard.

## Summary

- Set `metrics = true` in `[server]` to expose Prometheus metrics at `/metrics` via the handler in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs).
- Set `stats = true` in `[server]` to enable the JSON statistics API at `/v1/stats` as defined in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs).
- Use `metrics_addr` to bind the metrics endpoint to a separate address for network isolation.
- Customize metric names with `metric_namespace` under the `[observability]` table.
- Both endpoints are disabled by default and require explicit configuration before server startup.

## Frequently Asked Questions

### What format does the /metrics endpoint return?

The `/metrics` endpoint returns Prometheus exposition format (text/plain). The `prometheus_metrics` function in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs) scrapes the internal OpenTelemetry registry and renders counters, gauges, and histograms as Prometheus-compatible text.

### Can I disable the /v1/stats endpoint while keeping /metrics enabled?

Yes. Set `stats = false` (or omit the key) while keeping `metrics = true` in your [`config.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/config.toml). The `ServerConfig` parser in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) registers each route independently based on these boolean flags.

### How do I change the metric prefix exposed to Prometheus?

Add the `metric_namespace` key under the `[observability]` table in your TOML file. This string prepends to all metric names (defaulting to `switchyard`). For example, setting `metric_namespace = "prod_router"` changes `switchyard_requests_total` to `prod_router_requests_total`.

### Is authentication required for these endpoints?

Switchyard does not implement built-in authentication for the metrics or stats endpoints. If you expose these ports externally, place them behind a reverse proxy (like nginx or Envoy) or bind to localhost (`127.0.0.1`) using the `metrics_addr` configuration option to restrict access.