HTTP Endpoints Exposed by the Standalone Switchyard Proxy: Complete API Reference
The standalone Switchyard proxy exposes fifteen HTTP routes—including OpenAI-compatible chat completions, Anthropic message APIs, diagnostic utilities, and Prometheus metrics—all defined in crates/switchyard-server/src/lib.rs and registered via the build_switchyard_router function.
The NVIDIA-NeMo/Switchyard repository provides a high-performance LLM routing proxy designed to sit between client applications and backend language models. The standalone switchyard-server binary implements a comprehensive HTTP endpoint surface that supports drop-in compatibility with existing OpenAI and Anthropic SDKs while providing Switchyard-specific introspection and monitoring capabilities. These endpoints enable request routing, token counting, runtime statistics, and transparent proxy fallback for unsupported paths.
OpenAI-Compatible Endpoints
The proxy implements the standard OpenAI API surface to support existing clients without modification.
/v1/chat/completions
POST — Handles OpenAI-compatible chat completion requests. Located at line 492 in crates/switchyard-server/src/lib.rs, this endpoint receives standard OpenAI chat request payloads, forwards them to the selected LLM backend via Switchyard's routing algorithm, and streams responses back in OpenAI Chat format.
/v1/responses
POST — Provides the legacy OpenAI "responses" endpoint at line 494. It forwards requests to the chosen model and returns responses formatted according to the older OpenAI response schema for backward compatibility with legacy clients.
/v1/responses/input_tokens
POST — Returns the number of input tokens for an OpenAI-style response payload without performing inference (lines 74–76). This endpoint aids in usage-based billing calculations and pre-flight cost estimation.
/v1/responses/compact
POST — Produces a compacted version of an OpenAI response by removing unnecessary fields (line 77). This reduces payload size for downstream efficiency in bandwidth-constrained environments.
/v1/models
GET — Lists the model IDs that the router is capable of serving based on currently configured algorithms (line 78). Clients use this endpoint to discover available models and their identifiers.
Anthropic-Compatible Endpoints
Switchyard provides native support for Anthropic's message-based API format.
/v1/messages
POST — Implements the Anthropic-compatible messages endpoint at line 493. It receives Anthropic-style request payloads, routes them to the appropriate model backend, and returns responses in Anthropic message format.
/v1/messages/count_tokens
POST — Calculates the number of tokens a message would consume on the target model without performing actual inference (line 72). This Anthropic-specific helper enables clients to estimate costs and context window usage before sending full requests.
Switchyard-Specific Diagnostic Endpoints
These endpoints provide introspection into the routing layer and runtime behavior.
/v1/decision
POST — Returns the routing decision for a given request without invoking the model (line 71). This endpoint reveals which backend Switchyard would select for a specific payload, enabling debugging of routing algorithms and configuration validation.
/v1/stats
GET — Exposes aggregated runtime statistics including request counts, latency histograms, and error rates (line 79). The data is collected by the usage metrics subsystem implemented in crates/switchyard-server/src/usage_metrics.rs.
/v1/stats/reset
POST — Resets the in-memory statistics counters back to zero (line 80). Use this endpoint to start fresh measurement windows for benchmark testing or periodic metric rotation.
/v1/routing/session-stats
GET — Returns per-session routing statistics when the server is started with a routing log enabled (line 84). This optional endpoint traces routing decisions over the lifetime of a specific session, providing detailed audit trails for debugging complex routing scenarios.
Observability and Health Endpoints
/metrics
GET — Exposes Prometheus-compatible metrics for the entire server instance (line 81). The metrics collection logic resides in crates/switchyard-server/src/metrics.rs, providing standardized time-series data for monitoring systems like Grafana or Datadog.
/health
GET — Simple health-check endpoint returning HTTP 200 OK when the server is operational (line 82). Orchestration tools like Kubernetes use this for liveness and readiness probes.
Fallback Proxy Behavior
Any request that does not match the defined routes is handled by the fallback proxy handler (lines 50–73 in crates/switchyard-server/src/lib.rs). This ANY method catch-all forwards requests unchanged to the configured fallback_url, enabling Switchyard to act as a transparent proxy for unsupported API paths or custom extensions.
Implementation Details
All routes are constructed in crates/switchyard-server/src/lib.rs within the build_switchyard_router function. The server initialization occurs in crates/switchyard-server/src/main.rs, which starts the Axum server using the constructed router. Response formatting for both OpenAI and Anthropic protocols is handled in crates/switchyard-server/src/response.rs, while usage tracking relies on crates/switchyard-server/src/usage_metrics.rs.
Example: Calling Chat Completions
import requests
import json
url = "http://localhost:8000/v1/chat/completions"
payload = {
"model": "gpt-4",
"messages": [{"role": "user", "content": "Hello!"}]
}
resp = requests.post(url, json=payload)
print(resp.json())
Example: Listing Available Models
curl http://localhost:8000/v1/models
Example: Fetching Routing Statistics
curl http://localhost:8000/v1/stats | jq .
Summary
- Standard Compatibility: The standalone proxy exposes
/v1/chat/completionsand/v1/messagesfor OpenAI and Anthropic compatibility, alongside legacy response formats. - Diagnostic Tools: Endpoints like
/v1/decision,/v1/stats, and/v1/routing/session-statsprovide visibility into routing logic and runtime performance. - Utility Functions: Token counting endpoints (
/v1/messages/count_tokens,/v1/responses/input_tokens) enable pre-flight cost estimation. - Observability: Prometheus metrics via
/metricsand health checks via/healthsupport production monitoring. - Transparent Proxying: Unmatched routes fall back to the configured upstream API root, ensuring compatibility with non-standard paths.
Frequently Asked Questions
How do I check which backend Switchyard selected for my request?
Send a POST request to /v1/decision with your request payload. As implemented at line 71 of crates/switchyard-server/src/lib.rs, this endpoint returns the routing decision without invoking the actual model, showing exactly which backend would handle the request under current routing rules.
Can I use standard OpenAI client libraries with Switchyard?
Yes. The /v1/chat/completions endpoint (line 492) implements the full OpenAI chat completions protocol, allowing you to point existing OpenAI SDKs to http://localhost:8000/v1 as the base URL without code changes.
What metrics format does the /metrics endpoint provide?
The /metrics endpoint (line 81) exposes Prometheus-compatible text format. The underlying implementation in crates/switchyard-server/src/metrics.rs tracks request latency, throughput, and backend health status for scraping by Prometheus or compatible monitoring systems.
How do I reset statistics to start a fresh measurement window?
POST to /v1/stats/reset (line 80) to clear all in-memory counters. This is useful when running benchmark tests or when you need to isolate metrics for specific time periods without restarting the server.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →