How to Integrate Switchyard with NVIDIA NIM or Ollama Backends
Switchyard integrates with NVIDIA NIM and Ollama by defining OpenAI-compatible targets in a routes.toml file, setting the NVIDIA_API_KEY environment variable for NIM authentication, and routing requests through a passthrough route that forwards OpenAI-formatted calls to either backend.
Switchyard (NVIDIA-NeMo/Switchyard) is a proxy and translation library that converts OpenAI Chat, OpenAI Responses, and Anthropic Messages API calls into native LLM backend formats. Because both NVIDIA NIM and Ollama expose OpenAI-compatible HTTP endpoints, integrating them requires only configuration changes rather than custom code. This guide walks through the exact routes.toml syntax, authentication requirements, and verification steps needed to route traffic to either backend.
Configure Backend Targets in routes.toml
Switchyard reads routes.toml at startup to build its internal routing table, as described in crates/switchyard-server/README.md. Each backend is defined as a target with type = "openai" because both NIM and Ollama implement the OpenAI API schema.
NVIDIA NIM Target
Define a target pointing to the NIM OpenAI-compatible endpoint:
[[targets]]
name = "nim"
type = "openai"
base_url = "https://integrate.api.nvidia.com/v1"
Ollama Target
For local Ollama instances, use the default port with the /v1 path:
[[targets]]
name = "ollama"
type = "openai"
base_url = "http://127.0.0.1:11434/v1"
Handle Authentication
Authentication differs between the two backends. According to AGENTS.md at line 244, Switchyard checks for environment variables to inject Authorization headers automatically.
- NVIDIA NIM: Requires the
NVIDIA_API_KEYenvironment variable.
export NVIDIA_API_KEY="sk-nim-your-key-here"
- Ollama: Typically runs locally without authentication, so no environment variable is required.
The proxy forwards the Bearer token automatically when the variable is present, matching NIM’s required security scheme.
Define Passthrough Routes
Routes map incoming model IDs to specific targets. The passthrough route type is the simplest built-in strategy and performs no algorithmic decision making, making it ideal for direct backend mapping (see docs/routing_algorithms/overview.md).
[[routes]]
type = "passthrough"
id = "nim-route"
target = "nim"
model_id = "mixtral-8x7b-instruct"
[[routes]]
type = "passthrough"
id = "ollama-route"
target = "ollama"
model_id = "llama2-7b-chat"
When a client requests the model_id specified in the route, Switchyard forwards the call to the corresponding target.
Launch Options
You can run Switchyard either as a wrapped agent launcher or as a standalone server.
Launcher for Coding Agents
Use the launcher to integrate with tools like Claude Code:
# Route to NVIDIA NIM
switchyard launch claude --model nim-route --config routes.toml
# Route to Ollama
switchyard launch claude --model ollama-route --config routes.toml
Standalone Server
For general HTTP clients, run the Rust server directly:
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
This exposes http://localhost:4000/v1/chat/completions, which accepts standard OpenAI Chat API requests and proxies them to the configured NIM or Ollama endpoints.
Verify the Integration
Test the setup using a Python client with httpx:
import httpx
base = "http://127.0.0.1:4000/v1"
# Test NIM backend
resp = httpx.post(
f"{base}/chat/completions",
json={
"model": "mixtral-8x7b-instruct",
"messages": [{"role": "user", "content": "Hello, world!"}],
"max_tokens": 50,
},
timeout=30,
)
print(resp.json())
Switching between backends only requires changing the model field in the request to match the model_id defined in your routes.toml.
Protocol Translation Architecture
Switchyard performs bidirectional translation via the switchyard-translation crate documented in crates/switchyard-translation/README.md. When receiving an OpenAI Chat API request:
- The server parses the incoming JSON and identifies the route via
model_id. - The translation layer converts the payload if necessary (for NIM and Ollama, the OpenAI schema is passed through directly).
- The HTTP client forwards the request to the backend’s
base_url. - The response is translated back to the client’s expected format before returning.
This architecture allows any OpenAI-compatible client—including Claude Code, Codex CLI, or custom applications—to interact with NIM or Ollama without code changes.
Summary
- Create a
routes.tomlwith OpenAI-compatible targets pointing to NIM (https://integrate.api.nvidia.com/v1) or Ollama (http://127.0.0.1:11434/v1) base URLs. - Export
NVIDIA_API_KEYfor NIM authentication; Ollama requires no key. - Use
passthroughroutes to map specificmodel_idvalues directly to backend targets. - Run Switchyard via the launcher (
switchyard launch) for agent integration or the standalone server (switchyard-server) for general HTTP proxying. - Switchyard automatically translates requests and responses using the
switchyard-translationcrate, requiring no custom client code.
Frequently Asked Questions
Does Switchyard require custom code to support NIM or Ollama?
No. Both NIM and Ollama implement the OpenAI API schema, so Switchyard's existing OpenAI target type works without modification. The proxy handles request forwarding and response translation automatically via the switchyard-translation crate, as noted in the crate's README.
What is the difference between using the launcher and the standalone server?
The launcher (switchyard launch) wraps the server and is optimized for coding agents like Claude Code, automatically injecting configuration and handling process lifecycle. The standalone server (switchyard-server) runs as a persistent HTTP proxy on a configurable host and port, suitable for any OpenAI-compatible client or integration.
How does Switchyard handle authentication headers for NVIDIA NIM?
When the NVIDIA_API_KEY environment variable is present, Switchyard automatically adds an Authorization: Bearer header to requests sent to the NIM target, as implemented in the server logic referenced in AGENTS.md at line 244. Ollama local instances typically operate without authentication headers.
Can I route different models to different backends simultaneously?
Yes. Define multiple targets (one per backend) and multiple passthrough routes, each mapping a specific model_id to its respective target. Clients select the backend by specifying the corresponding model ID in their API calls, allowing you to serve NIM-hosted models and local Ollama models from a single Switchyard instance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →