Switchyard Server HTTP Endpoints: Complete API Reference
The Switchyard server exposes 12 HTTP routes—including OpenAI-compatible chat completions, Anthropic-style messages, and Prometheus metrics—through the build_switchyard_router function in crates/switchyard-server/src/lib.rs.
The NVIDIA-NeMo Switchyard repository implements an intelligent request routing layer for large language models. Understanding the Switchyard HTTP endpoints enables seamless integration with existing AI SDKs and robust production monitoring of the Axum-based server.
Core Router Architecture
All Switchyard HTTP endpoints are registered in the build_switchyard_router function within crates/switchyard-server/src/lib.rs (lines 1484-1492). This function constructs an Axum Router that maps HTTP methods and paths to specific handler functions implementing OpenAI-compatible and Anthropic-compatible semantics.
Every route shares a common middleware stack defined immediately after router construction (lines 1499-1502) that enforces a 32 MiB request body limit (DEFAULT_MAX_REQUEST_BODY_BYTES) and stamps request start times for latency measurement.
OpenAI-Compatible Endpoints
Switchyard provides drop-in replacements for standard OpenAI API routes, allowing existing clients to redirect traffic without code changes.
Chat Completions (POST /v1/chat/completions)
The openai_chat_completions handler processes chat completion requests at POST /v1/chat/completions. This endpoint accepts standard OpenAI request payloads and returns streamed or non-streamed responses based on the stream parameter.
curl -X POST http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"switchyard/general","messages":[{"role":"user","content":"Hello"}]}'
Response Objects (POST /v1/responses)
The openai_responses handler at POST /v1/responses supports legacy OpenAI response object formats for backward compatibility with older client implementations.
List Models (GET /v1/models)
The models handler responds to GET /v1/models with a JSON array listing all configured model IDs and their capabilities, enabling client-side model discovery.
curl http://localhost:4000/v1/models
Anthropic-Compatible Endpoints
For applications built against the Anthropic API, Switchyard offers equivalent message handling endpoints.
Messages (POST /v1/messages)
The anthropic_messages handler at POST /v1/messages implements Anthropic's message format, translating requests into Switchyard's internal routing logic before returning Anthropic-compatible response structures.
curl -X POST http://localhost:4000/v1/messages \
-H "Content-Type: application/json" \
-d '{"model":"switchyard/general","messages":[{"role":"user","content":"Explain Rust"}]}'
Token Counting (POST /v1/messages/count_tokens)
The anthropic_count_tokens handler provides token estimation at POST /v1/messages/count_tokens without executing the actual model inference, useful for pre-request quota management.
curl -X POST http://localhost:4000/v1/messages/count_tokens \
-H "Content-Type: application/json" \
-d '{"model":"switchyard/anthropic","messages":[{"role":"user","content":"Count my tokens"}]}'
Routing and Administrative Endpoints
Beyond standard LLM APIs, Switchyard exposes specialized endpoints for routing decisions and server management.
Decision Routing (POST /v1/decision)
The decision handler at POST /v1/decision returns routing metadata—including target model selection and latency predictions—without invoking an answer model. This endpoint supports dry-run testing of routing policies.
Statistics Management
Switchyard aggregates routing metrics accessible through two endpoints:
GET /v1/stats: Theget_statshandler returns aggregated routing statistics, including request counts, latency percentiles, and cache hit rates.POST /v1/stats/reset: Thereset_statshandler clears all accumulated statistics counters.
Session Statistics (GET /v1/routing/session-stats)
When routing-log configuration is enabled, the get_session_stats handler exposes GET /v1/routing/session-stats for retrieving per-session routing logs. This endpoint is conditionally available and only served when the server is started with appropriate observability flags.
Observability and Health Endpoints
Production deployments rely on dedicated monitoring routes outside the standard API paths.
Prometheus Metrics (GET /metrics)
The prometheus_metrics handler at GET /metrics exposes server metrics in Prometheus exposition format, compatible with standard monitoring stacks like Grafana and Datadog.
curl http://localhost:4000/metrics
Health Check (GET /health)
The health handler provides a lightweight liveness probe at GET /health, returning {"status":"ok"} for load balancers and container orchestrators.
curl http://localhost:4000/health
Fallback Handler
Any unmatched paths trigger the not_found handler, returning a 404-style error response with appropriate JSON error formatting.
Programmatic Client Examples
Python Async Client
For Python applications, use httpx for non-blocking requests to the Switchyard HTTP API:
import httpx
import asyncio
async def chat():
async with httpx.AsyncClient(base_url="http://localhost:4000") as client:
resp = await client.post(
"/v1/chat/completions",
json={
"model": "switchyard/general",
"messages": [{"role": "user", "content": "Hello"}],
},
)
print(resp.json())
asyncio.run(chat())
Key Implementation Files
The HTTP surface is implemented across the following source files in the crates/switchyard-server directory:
src/lib.rs: Defines the Axum router, all endpoint handlers includingbuild_switchyard_router, and middleware configuration (lines 1484-1502).src/main.rs: Entry point that parses CLI arguments, createsServerState, and launches the HTTP server.src/cli.rs: Implements the command-line interface with flags for bind address, TLS configuration, and TOML config paths.src/config.rs: Parses the TOML configuration describing routes, target models, and client credentials.src/response.rs: Translates internal response objects into HTTP responses respecting requested wire formats.src/metrics.rs: Exposes Prometheus metrics collected from request handling and routing statistics.src/observability.rs: Provides tracing spans and request-level observability hooks used by all handlers.
Summary
- The Switchyard HTTP server exposes 12 routes through the
build_switchyard_routerfunction incrates/switchyard-server/src/lib.rs. - OpenAI-compatible endpoints include
POST /v1/chat/completions,POST /v1/responses, andGET /v1/models. - Anthropic-compatible endpoints include
POST /v1/messagesandPOST /v1/messages/count_tokens. - Administrative routes provide routing decisions (
POST /v1/decision), statistics management (GET /v1/stats,POST /v1/stats/reset), and conditional session logging. - Observability is supported through Prometheus metrics (
GET /metrics) and health checks (GET /health). - All endpoints enforce a 32 MiB request body limit and automatic latency tracking via Axum middleware.
Frequently Asked Questions
What base URL path prefix do Switchyard endpoints use?
All Switchyard API endpoints use the /v1/ prefix for versioned routes (e.g., /v1/chat/completions), with the exception of system-level endpoints like /health and /metrics which mount at the root path. Configure your client’s base_url to http://localhost:4000 (or your configured host) without the /v1 suffix, as the endpoints include this prefix natively.
Can I use the standard OpenAI Python SDK with Switchyard?
Yes. Because Switchyard implements the OpenAI-compatible POST /v1/chat/completions and GET /v1/models endpoints, you can point the OpenAI SDK to your Switchyard server by setting the base_url parameter to http://localhost:4000/v1. Standard chat completion calls will function without modifying your application code, as the request and response schemas match the OpenAI specification exactly.
What is the maximum payload size for requests?
The Switchyard server enforces a 32 MiB maximum request body size through the DEFAULT_MAX_REQUEST_BODY_BYTES middleware constant defined in crates/switchyard-server/src/lib.rs (lines 1499-1502). Requests exceeding this limit receive an HTTP 413 Payload Too Large response before reaching any handler logic.
How do I access per-session routing decisions for debugging?
When the server is configured with routing-log enabled, the GET /v1/routing/session-stats endpoint exposes detailed per-session routing logs through the get_session_stats handler. This endpoint is conditionally available based on your TOML configuration in crates/switchyard-server/src/config.rs, providing visibility into model selection logic and latency predictions for individual requests without exposing sensitive request content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →