How to Use NVIDIA Switchyard with Self-Hosted Models Like vLLM and Ollama
Yes, Switchyard natively supports self-hosted LLMs including vLLM and Ollama by routing requests to any OpenAI-compatible endpoint through a simple TOML configuration.
NVIDIA Switchyard is an open-source LLM routing proxy that sits between AI clients (Claude Code, Codex CLI, etc.) and backend model servers. According to the official README, it explicitly forwards traffic to vLLM, NVIDIA NIM, Ollama, or any OpenAI-compatible endpoint without requiring code modifications.
How Switchyard Routes to Self-Hosted Backends
The architecture follows a three-step flow:
- Client → Switchyard — The client sends requests in OpenAI or Anthropic API format
- Switchyard → Backend — A
routes.tomlconfiguration selects the target and translates the request - Backend → Switchyard → Client — Responses are translated back to the client's original format
The routing decision is entirely configuration-driven. You define backends in a TOML file, and Switchyard handles protocol translation automatically.
Configuring Switchyard for vLLM
vLLM exposes an OpenAI-compatible server via its built-in API. To route Switchyard traffic to a local vLLM instance:
Start the vLLM server:
python -m vllm.entrypoints.openai.server \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--port 8000
Create routes.toml:
[[routes]]
name = "my-vllm-route"
type = "passthrough"
[[routes.targets]]
model_id = "vllm-8b"
backend = "openai"
url = "http://localhost:8000/v1"
Launch Switchyard with your config:
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
Or use the CLI launcher:
switchyard launch claude --model vllm-8b --config routes.toml
Configuring Switchyard for Ollama
Ollama's API server exposes an OpenAI-compatible endpoint at /v1 when running locally.
Start Ollama:
ollama serve
Create routes.toml:
[[routes]]
name = "my-ollama-route"
type = "passthrough"
[[routes.targets]]
model_id = "ollama-phi"
backend = "openai"
url = "http://localhost:11434/v1"
The backend = "openai" setting tells Switchyard to use OpenAI-style request/response translation, which Ollama's endpoint accepts natively.
Complete Python Example
Once Switchyard is running on port 4000, send requests through it:
import os
import httpx
os.environ["OPENAI_API_KEY"] = "dummy" # Required by client, ignored by self-hosted backends
payload = {
"model": "vllm-8b",
"messages": [{"role": "user", "content": "Explain quantum entanglement in 2 sentences"}],
"max_tokens": 256,
}
response = httpx.post(
"http://localhost:4000/v1/chat/completions",
json=payload,
timeout=30
)
print(response.json())
Switchyard receives the request, routes it to localhost:8000, and returns the vLLM response in standard OpenAI Chat format.
Key Source Files for Self-Hosted Routing
| File | Purpose |
|---|---|
switchyard/cli/launch_command.py |
CLI entry point that parses --config and forwards to the appropriate launcher |
switchyard/cli/defaults/openrouter.toml |
Default packaged config; replace with your self-hosted routes |
docs/routing_algorithms/overview.md |
Documents the "passthrough" route type for direct backend mapping |
crates/switchyard-server/README.md |
Server documentation for running switchyard-server with custom TOML |
Passthrough Route Type Explained
The type = "passthrough" configuration used for vLLM and Ollama is the simplest routing mode. As documented in [docs/routing_algorithms/overview.md](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md), it registers a model ID that maps directly to a backend URL with minimal overhead. No load balancing, fallback logic, or token-based routing is applied—requests flow straight through to the configured endpoint.
Summary
- Switchyard supports self-hosted models including vLLM and Ollama via OpenAI-compatible endpoints
- Configuration-only setup: define backends in
routes.toml, no code changes required - Passthrough routes provide direct mapping from model IDs to backend URLs
- CLI and server modes both accept custom configs via
--config - Protocol translation happens automatically between client and backend formats
Frequently Asked Questions
Does Switchyard require code changes to support new self-hosted models?
No. Switchyard's routing is entirely configuration-driven. You add new backends by defining routes in a TOML file. The switchyard launch command and switchyard-server binary both accept --config to load custom routes without recompilation.
What backend types does Switchyard recognize for self-hosted models?
The backend field in routes supports values like "openai" for OpenAI-compatible APIs. This covers vLLM, Ollama, NVIDIA NIM, and any other server implementing the OpenAI chat completions specification. The backend value determines how Switchyard translates request and response formats.
Can Switchyard route to multiple self-hosted models simultaneously?
Yes. A single routes.toml can define multiple [[routes]] sections, each with its own targets. You can route model-a to a vLLM instance on port 8000 and model-b to Ollama on port 11434. Switchyard selects the appropriate backend based on the model parameter in incoming requests.
How do I debug routing problems with self-hosted backends?
Run switchyard-server with verbose logging and verify your routes.toml syntax against the examples in switchyard/cli/defaults/. Confirm the backend's OpenAI-compatible endpoint is accessible at the configured URL before starting Switchyard.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →