# HTTP Endpoints Exposed by the Standalone Switchyard Proxy: Complete API Reference

> Explore the standalone Switchyard proxy's HTTP endpoints including OpenAI compatible chat completions, Anthropic APIs, and Prometheus metrics. Discover each endpoint's function in our API reference.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: api-reference
- Published: 2026-09-13

---

**The standalone Switchyard proxy exposes fifteen HTTP routes—including OpenAI-compatible chat completions, Anthropic message APIs, diagnostic utilities, and Prometheus metrics—all defined in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) and registered via the `build_switchyard_router` function.**

The NVIDIA-NeMo/Switchyard repository provides a high-performance LLM routing proxy designed to sit between client applications and backend language models. The standalone **switchyard-server** binary implements a comprehensive HTTP endpoint surface that supports drop-in compatibility with existing OpenAI and Anthropic SDKs while providing Switchyard-specific introspection and monitoring capabilities. These endpoints enable request routing, token counting, runtime statistics, and transparent proxy fallback for unsupported paths.

## OpenAI-Compatible Endpoints

The proxy implements the standard OpenAI API surface to support existing clients without modification.

### /v1/chat/completions

**POST** — Handles OpenAI-compatible chat completion requests. Located at line 492 in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), this endpoint receives standard OpenAI chat request payloads, forwards them to the selected LLM backend via Switchyard's routing algorithm, and streams responses back in OpenAI Chat format.

### /v1/responses

**POST** — Provides the legacy OpenAI "responses" endpoint at line 494. It forwards requests to the chosen model and returns responses formatted according to the older OpenAI response schema for backward compatibility with legacy clients.

### /v1/responses/input_tokens

**POST** — Returns the number of input tokens for an OpenAI-style response payload without performing inference (lines 74–76). This endpoint aids in usage-based billing calculations and pre-flight cost estimation.

### /v1/responses/compact

**POST** — Produces a compacted version of an OpenAI response by removing unnecessary fields (line 77). This reduces payload size for downstream efficiency in bandwidth-constrained environments.

### /v1/models

**GET** — Lists the model IDs that the router is capable of serving based on currently configured algorithms (line 78). Clients use this endpoint to discover available models and their identifiers.

## Anthropic-Compatible Endpoints

Switchyard provides native support for Anthropic's message-based API format.

### /v1/messages

**POST** — Implements the Anthropic-compatible messages endpoint at line 493. It receives Anthropic-style request payloads, routes them to the appropriate model backend, and returns responses in Anthropic message format.

### /v1/messages/count_tokens

**POST** — Calculates the number of tokens a message would consume on the target model without performing actual inference (line 72). This Anthropic-specific helper enables clients to estimate costs and context window usage before sending full requests.

## Switchyard-Specific Diagnostic Endpoints

These endpoints provide introspection into the routing layer and runtime behavior.

### /v1/decision

**POST** — Returns the routing decision for a given request without invoking the model (line 71). This endpoint reveals which backend Switchyard would select for a specific payload, enabling debugging of routing algorithms and configuration validation.

### /v1/stats

**GET** — Exposes aggregated runtime statistics including request counts, latency histograms, and error rates (line 79). The data is collected by the usage metrics subsystem implemented in [`crates/switchyard-server/src/usage_metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs).

### /v1/stats/reset

**POST** — Resets the in-memory statistics counters back to zero (line 80). Use this endpoint to start fresh measurement windows for benchmark testing or periodic metric rotation.

### /v1/routing/session-stats

**GET** — Returns per-session routing statistics when the server is started with a routing log enabled (line 84). This optional endpoint traces routing decisions over the lifetime of a specific session, providing detailed audit trails for debugging complex routing scenarios.

## Observability and Health Endpoints

### /metrics

**GET** — Exposes Prometheus-compatible metrics for the entire server instance (line 81). The metrics collection logic resides in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs), providing standardized time-series data for monitoring systems like Grafana or Datadog.

### /health

**GET** — Simple health-check endpoint returning HTTP 200 OK when the server is operational (line 82). Orchestration tools like Kubernetes use this for liveness and readiness probes.

## Fallback Proxy Behavior

Any request that does not match the defined routes is handled by the fallback proxy handler (lines 50–73 in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs)). This **ANY** method catch-all forwards requests unchanged to the configured `fallback_url`, enabling Switchyard to act as a transparent proxy for unsupported API paths or custom extensions.

## Implementation Details

All routes are constructed in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) within the `build_switchyard_router` function. The server initialization occurs in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs), which starts the Axum server using the constructed router. Response formatting for both OpenAI and Anthropic protocols is handled in [`crates/switchyard-server/src/response.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/response.rs), while usage tracking relies on [`crates/switchyard-server/src/usage_metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/usage_metrics.rs).

### Example: Calling Chat Completions

```python
import requests
import json

url = "http://localhost:8000/v1/chat/completions"
payload = {
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
}
resp = requests.post(url, json=payload)
print(resp.json())

```

### Example: Listing Available Models

```bash
curl http://localhost:8000/v1/models

```

### Example: Fetching Routing Statistics

```bash
curl http://localhost:8000/v1/stats | jq .

```

## Summary

- **Standard Compatibility**: The standalone proxy exposes `/v1/chat/completions` and `/v1/messages` for OpenAI and Anthropic compatibility, alongside legacy response formats.
- **Diagnostic Tools**: Endpoints like `/v1/decision`, `/v1/stats`, and `/v1/routing/session-stats` provide visibility into routing logic and runtime performance.
- **Utility Functions**: Token counting endpoints (`/v1/messages/count_tokens`, `/v1/responses/input_tokens`) enable pre-flight cost estimation.
- **Observability**: Prometheus metrics via `/metrics` and health checks via `/health` support production monitoring.
- **Transparent Proxying**: Unmatched routes fall back to the configured upstream API root, ensuring compatibility with non-standard paths.

## Frequently Asked Questions

### How do I check which backend Switchyard selected for my request?

Send a POST request to `/v1/decision` with your request payload. As implemented at line 71 of [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), this endpoint returns the routing decision without invoking the actual model, showing exactly which backend would handle the request under current routing rules.

### Can I use standard OpenAI client libraries with Switchyard?

Yes. The `/v1/chat/completions` endpoint (line 492) implements the full OpenAI chat completions protocol, allowing you to point existing OpenAI SDKs to `http://localhost:8000/v1` as the base URL without code changes.

### What metrics format does the /metrics endpoint provide?

The `/metrics` endpoint (line 81) exposes Prometheus-compatible text format. The underlying implementation in [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs) tracks request latency, throughput, and backend health status for scraping by Prometheus or compatible monitoring systems.

### How do I reset statistics to start a fresh measurement window?

POST to `/v1/stats/reset` (line 80) to clear all in-memory counters. This is useful when running benchmark tests or when you need to isolate metrics for specific time periods without restarting the server.