# Switchyard Server HTTP Endpoints: Complete API Reference

> Explore Switchyard server HTTP endpoints including OpenAI chat, Anthropic messages, and Prometheus metrics. Discover the complete API reference for efficient integration.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: api-reference
- Published: 2026-08-21

---

**The Switchyard server exposes 12 HTTP routes—including OpenAI-compatible chat completions, Anthropic-style messages, and Prometheus metrics—through the `build_switchyard_router` function in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs).**

The NVIDIA-NeMo Switchyard repository implements an intelligent request routing layer for large language models. Understanding the Switchyard HTTP endpoints enables seamless integration with existing AI SDKs and robust production monitoring of the Axum-based server.

## Core Router Architecture

All Switchyard HTTP endpoints are registered in the **`build_switchyard_router`** function within [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) (lines 1484-1492). This function constructs an Axum `Router` that maps HTTP methods and paths to specific handler functions implementing OpenAI-compatible and Anthropic-compatible semantics.

Every route shares a common middleware stack defined immediately after router construction (lines 1499-1502) that enforces a **32 MiB** request body limit (`DEFAULT_MAX_REQUEST_BODY_BYTES`) and stamps request start times for latency measurement.

## OpenAI-Compatible Endpoints

Switchyard provides drop-in replacements for standard OpenAI API routes, allowing existing clients to redirect traffic without code changes.

### Chat Completions (POST /v1/chat/completions)

The **`openai_chat_completions`** handler processes chat completion requests at `POST /v1/chat/completions`. This endpoint accepts standard OpenAI request payloads and returns streamed or non-streamed responses based on the `stream` parameter.

```bash
curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard/general","messages":[{"role":"user","content":"Hello"}]}'

```

### Response Objects (POST /v1/responses)

The **`openai_responses`** handler at `POST /v1/responses` supports legacy OpenAI response object formats for backward compatibility with older client implementations.

### List Models (GET /v1/models)

The **`models`** handler responds to `GET /v1/models` with a JSON array listing all configured model IDs and their capabilities, enabling client-side model discovery.

```bash
curl http://localhost:4000/v1/models

```

## Anthropic-Compatible Endpoints

For applications built against the Anthropic API, Switchyard offers equivalent message handling endpoints.

### Messages (POST /v1/messages)

The **`anthropic_messages`** handler at `POST /v1/messages` implements Anthropic's message format, translating requests into Switchyard's internal routing logic before returning Anthropic-compatible response structures.

```bash
curl -X POST http://localhost:4000/v1/messages \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard/general","messages":[{"role":"user","content":"Explain Rust"}]}'

```

### Token Counting (POST /v1/messages/count_tokens)

The **`anthropic_count_tokens`** handler provides token estimation at `POST /v1/messages/count_tokens` without executing the actual model inference, useful for pre-request quota management.

```bash
curl -X POST http://localhost:4000/v1/messages/count_tokens \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard/anthropic","messages":[{"role":"user","content":"Count my tokens"}]}'

```

## Routing and Administrative Endpoints

Beyond standard LLM APIs, Switchyard exposes specialized endpoints for routing decisions and server management.

### Decision Routing (POST /v1/decision)

The **`decision`** handler at `POST /v1/decision` returns routing metadata—including target model selection and latency predictions—without invoking an answer model. This endpoint supports dry-run testing of routing policies.

### Statistics Management

Switchyard aggregates routing metrics accessible through two endpoints:

- **`GET /v1/stats`**: The `get_stats` handler returns aggregated routing statistics, including request counts, latency percentiles, and cache hit rates.
- **`POST /v1/stats/reset`**: The `reset_stats` handler clears all accumulated statistics counters.

### Session Statistics (GET /v1/routing/session-stats)

When routing-log configuration is enabled, the **`get_session_stats`** handler exposes `GET /v1/routing/session-stats` for retrieving per-session routing logs. This endpoint is conditionally available and only served when the server is started with appropriate observability flags.

## Observability and Health Endpoints

Production deployments rely on dedicated monitoring routes outside the standard API paths.

### Prometheus Metrics (GET /metrics)

The **`prometheus_metrics`** handler at `GET /metrics` exposes server metrics in Prometheus exposition format, compatible with standard monitoring stacks like Grafana and Datadog.

```bash
curl http://localhost:4000/metrics

```

### Health Check (GET /health)

The **`health`** handler provides a lightweight liveness probe at `GET /health`, returning `{"status":"ok"}` for load balancers and container orchestrators.

```bash
curl http://localhost:4000/health

```

### Fallback Handler

Any unmatched paths trigger the **`not_found`** handler, returning a 404-style error response with appropriate JSON error formatting.

## Programmatic Client Examples

### Python Async Client

For Python applications, use `httpx` for non-blocking requests to the Switchyard HTTP API:

```python
import httpx
import asyncio

async def chat():
    async with httpx.AsyncClient(base_url="http://localhost:4000") as client:
        resp = await client.post(
            "/v1/chat/completions",
            json={
                "model": "switchyard/general",
                "messages": [{"role": "user", "content": "Hello"}],
            },
        )
        print(resp.json())

asyncio.run(chat())

```

## Key Implementation Files

The HTTP surface is implemented across the following source files in the `crates/switchyard-server` directory:

- **[`src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/lib.rs)**: Defines the Axum router, all endpoint handlers including `build_switchyard_router`, and middleware configuration (lines 1484-1502).
- **[`src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/main.rs)**: Entry point that parses CLI arguments, creates `ServerState`, and launches the HTTP server.
- **[`src/cli.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/cli.rs)**: Implements the command-line interface with flags for bind address, TLS configuration, and TOML config paths.
- **[`src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/config.rs)**: Parses the TOML configuration describing routes, target models, and client credentials.
- **[`src/response.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/response.rs)**: Translates internal response objects into HTTP responses respecting requested wire formats.
- **[`src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/metrics.rs)**: Exposes Prometheus metrics collected from request handling and routing statistics.
- **[`src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/observability.rs)**: Provides tracing spans and request-level observability hooks used by all handlers.

## Summary

- The Switchyard HTTP server exposes **12 routes** through the `build_switchyard_router` function in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs).
- **OpenAI-compatible** endpoints include `POST /v1/chat/completions`, `POST /v1/responses`, and `GET /v1/models`.
- **Anthropic-compatible** endpoints include `POST /v1/messages` and `POST /v1/messages/count_tokens`.
- Administrative routes provide routing decisions (`POST /v1/decision`), statistics management (`GET /v1/stats`, `POST /v1/stats/reset`), and conditional session logging.
- **Observability** is supported through Prometheus metrics (`GET /metrics`) and health checks (`GET /health`).
- All endpoints enforce a **32 MiB request body limit** and automatic latency tracking via Axum middleware.

## Frequently Asked Questions

### What base URL path prefix do Switchyard endpoints use?

All Switchyard API endpoints use the `/v1/` prefix for versioned routes (e.g., `/v1/chat/completions`), with the exception of system-level endpoints like `/health` and `/metrics` which mount at the root path. Configure your client’s `base_url` to `http://localhost:4000` (or your configured host) without the `/v1` suffix, as the endpoints include this prefix natively.

### Can I use the standard OpenAI Python SDK with Switchyard?

Yes. Because Switchyard implements the OpenAI-compatible `POST /v1/chat/completions` and `GET /v1/models` endpoints, you can point the OpenAI SDK to your Switchyard server by setting the `base_url` parameter to `http://localhost:4000/v1`. Standard chat completion calls will function without modifying your application code, as the request and response schemas match the OpenAI specification exactly.

### What is the maximum payload size for requests?

The Switchyard server enforces a **32 MiB** maximum request body size through the `DEFAULT_MAX_REQUEST_BODY_BYTES` middleware constant defined in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) (lines 1499-1502). Requests exceeding this limit receive an HTTP 413 Payload Too Large response before reaching any handler logic.

### How do I access per-session routing decisions for debugging?

When the server is configured with routing-log enabled, the `GET /v1/routing/session-stats` endpoint exposes detailed per-session routing logs through the `get_session_stats` handler. This endpoint is conditionally available based on your TOML configuration in [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs), providing visibility into model selection logic and latency predictions for individual requests without exposing sensitive request content.