How to Integrate Switchyard with NVIDIA NIM or Ollama Backends

Switchyard integrates with NVIDIA NIM and Ollama by defining OpenAI-compatible targets in a routes.toml file, setting the NVIDIA_API_KEY environment variable for NIM authentication, and routing requests through a passthrough route that forwards OpenAI-formatted calls to either backend.

Switchyard (NVIDIA-NeMo/Switchyard) is a proxy and translation library that converts OpenAI Chat, OpenAI Responses, and Anthropic Messages API calls into native LLM backend formats. Because both NVIDIA NIM and Ollama expose OpenAI-compatible HTTP endpoints, integrating them requires only configuration changes rather than custom code. This guide walks through the exact routes.toml syntax, authentication requirements, and verification steps needed to route traffic to either backend.

Configure Backend Targets in routes.toml

Switchyard reads routes.toml at startup to build its internal routing table, as described in crates/switchyard-server/README.md. Each backend is defined as a target with type = "openai" because both NIM and Ollama implement the OpenAI API schema.

NVIDIA NIM Target

Define a target pointing to the NIM OpenAI-compatible endpoint:

[[targets]]
name = "nim"
type = "openai"
base_url = "https://integrate.api.nvidia.com/v1"

Ollama Target

For local Ollama instances, use the default port with the /v1 path:

[[targets]]
name = "ollama"
type = "openai"
base_url = "http://127.0.0.1:11434/v1"

Handle Authentication

Authentication differs between the two backends. According to AGENTS.md at line 244, Switchyard checks for environment variables to inject Authorization headers automatically.

  • NVIDIA NIM: Requires the NVIDIA_API_KEY environment variable.
export NVIDIA_API_KEY="sk-nim-your-key-here"
  • Ollama: Typically runs locally without authentication, so no environment variable is required.

The proxy forwards the Bearer token automatically when the variable is present, matching NIM’s required security scheme.

Define Passthrough Routes

Routes map incoming model IDs to specific targets. The passthrough route type is the simplest built-in strategy and performs no algorithmic decision making, making it ideal for direct backend mapping (see docs/routing_algorithms/overview.md).

[[routes]]
type = "passthrough"
id = "nim-route"
target = "nim"
model_id = "mixtral-8x7b-instruct"

[[routes]]
type = "passthrough"
id = "ollama-route"
target = "ollama"
model_id = "llama2-7b-chat"

When a client requests the model_id specified in the route, Switchyard forwards the call to the corresponding target.

Launch Options

You can run Switchyard either as a wrapped agent launcher or as a standalone server.

Launcher for Coding Agents

Use the launcher to integrate with tools like Claude Code:


# Route to NVIDIA NIM

switchyard launch claude --model nim-route --config routes.toml

# Route to Ollama

switchyard launch claude --model ollama-route --config routes.toml

Standalone Server

For general HTTP clients, run the Rust server directly:

switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

This exposes http://localhost:4000/v1/chat/completions, which accepts standard OpenAI Chat API requests and proxies them to the configured NIM or Ollama endpoints.

Verify the Integration

Test the setup using a Python client with httpx:

import httpx

base = "http://127.0.0.1:4000/v1"

# Test NIM backend

resp = httpx.post(
    f"{base}/chat/completions",
    json={
        "model": "mixtral-8x7b-instruct",
        "messages": [{"role": "user", "content": "Hello, world!"}],
        "max_tokens": 50,
    },
    timeout=30,
)
print(resp.json())

Switching between backends only requires changing the model field in the request to match the model_id defined in your routes.toml.

Protocol Translation Architecture

Switchyard performs bidirectional translation via the switchyard-translation crate documented in crates/switchyard-translation/README.md. When receiving an OpenAI Chat API request:

  1. The server parses the incoming JSON and identifies the route via model_id.
  2. The translation layer converts the payload if necessary (for NIM and Ollama, the OpenAI schema is passed through directly).
  3. The HTTP client forwards the request to the backend’s base_url.
  4. The response is translated back to the client’s expected format before returning.

This architecture allows any OpenAI-compatible client—including Claude Code, Codex CLI, or custom applications—to interact with NIM or Ollama without code changes.

Summary

  • Create a routes.toml with OpenAI-compatible targets pointing to NIM (https://integrate.api.nvidia.com/v1) or Ollama (http://127.0.0.1:11434/v1) base URLs.
  • Export NVIDIA_API_KEY for NIM authentication; Ollama requires no key.
  • Use passthrough routes to map specific model_id values directly to backend targets.
  • Run Switchyard via the launcher (switchyard launch) for agent integration or the standalone server (switchyard-server) for general HTTP proxying.
  • Switchyard automatically translates requests and responses using the switchyard-translation crate, requiring no custom client code.

Frequently Asked Questions

Does Switchyard require custom code to support NIM or Ollama?

No. Both NIM and Ollama implement the OpenAI API schema, so Switchyard's existing OpenAI target type works without modification. The proxy handles request forwarding and response translation automatically via the switchyard-translation crate, as noted in the crate's README.

What is the difference between using the launcher and the standalone server?

The launcher (switchyard launch) wraps the server and is optimized for coding agents like Claude Code, automatically injecting configuration and handling process lifecycle. The standalone server (switchyard-server) runs as a persistent HTTP proxy on a configurable host and port, suitable for any OpenAI-compatible client or integration.

How does Switchyard handle authentication headers for NVIDIA NIM?

When the NVIDIA_API_KEY environment variable is present, Switchyard automatically adds an Authorization: Bearer header to requests sent to the NIM target, as implemented in the server logic referenced in AGENTS.md at line 244. Ollama local instances typically operate without authentication headers.

Can I route different models to different backends simultaneously?

Yes. Define multiple targets (one per backend) and multiple passthrough routes, each mapping a specific model_id to its respective target. Clients select the backend by specifying the corresponding model ID in their API calls, allowing you to serve NIM-hosted models and local Ollama models from a single Switchyard instance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →