# How to Integrate Switchyard with NVIDIA NIM or Ollama Backends

> Integrate Switchyard with NVIDIA NIM or Ollama backends easily. Learn how to configure routes.toml and environment variables for seamless OpenAI-compatible API calls.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-17

---

**Switchyard integrates with NVIDIA NIM and Ollama by defining OpenAI-compatible targets in a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file, setting the `NVIDIA_API_KEY` environment variable for NIM authentication, and routing requests through a passthrough route that forwards OpenAI-formatted calls to either backend.**

Switchyard (NVIDIA-NeMo/Switchyard) is a proxy and translation library that converts OpenAI Chat, OpenAI Responses, and Anthropic Messages API calls into native LLM backend formats. Because both NVIDIA NIM and Ollama expose OpenAI-compatible HTTP endpoints, integrating them requires only configuration changes rather than custom code. This guide walks through the exact [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) syntax, authentication requirements, and verification steps needed to route traffic to either backend.

## Configure Backend Targets in routes.toml

Switchyard reads [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) at startup to build its internal routing table, as described in [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md). Each backend is defined as a target with `type = "openai"` because both NIM and Ollama implement the OpenAI API schema.

### NVIDIA NIM Target

Define a target pointing to the NIM OpenAI-compatible endpoint:

```toml
[[targets]]
name = "nim"
type = "openai"
base_url = "https://integrate.api.nvidia.com/v1"

```

### Ollama Target

For local Ollama instances, use the default port with the `/v1` path:

```toml
[[targets]]
name = "ollama"
type = "openai"
base_url = "http://127.0.0.1:11434/v1"

```

## Handle Authentication

Authentication differs between the two backends. According to [`AGENTS.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/AGENTS.md) at line 244, Switchyard checks for environment variables to inject `Authorization` headers automatically.

- **NVIDIA NIM:** Requires the `NVIDIA_API_KEY` environment variable.

```bash
export NVIDIA_API_KEY="sk-nim-your-key-here"

```

- **Ollama:** Typically runs locally without authentication, so no environment variable is required.

The proxy forwards the Bearer token automatically when the variable is present, matching NIM’s required security scheme.

## Define Passthrough Routes

Routes map incoming model IDs to specific targets. The `passthrough` route type is the simplest built-in strategy and performs no algorithmic decision making, making it ideal for direct backend mapping (see [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md)).

```toml
[[routes]]
type = "passthrough"
id = "nim-route"
target = "nim"
model_id = "mixtral-8x7b-instruct"

[[routes]]
type = "passthrough"
id = "ollama-route"
target = "ollama"
model_id = "llama2-7b-chat"

```

When a client requests the `model_id` specified in the route, Switchyard forwards the call to the corresponding target.

## Launch Options

You can run Switchyard either as a wrapped agent launcher or as a standalone server.

### Launcher for Coding Agents

Use the launcher to integrate with tools like Claude Code:

```bash

# Route to NVIDIA NIM

switchyard launch claude --model nim-route --config routes.toml

# Route to Ollama

switchyard launch claude --model ollama-route --config routes.toml

```

### Standalone Server

For general HTTP clients, run the Rust server directly:

```bash
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

This exposes `http://localhost:4000/v1/chat/completions`, which accepts standard OpenAI Chat API requests and proxies them to the configured NIM or Ollama endpoints.

## Verify the Integration

Test the setup using a Python client with `httpx`:

```python
import httpx

base = "http://127.0.0.1:4000/v1"

# Test NIM backend

resp = httpx.post(
    f"{base}/chat/completions",
    json={
        "model": "mixtral-8x7b-instruct",
        "messages": [{"role": "user", "content": "Hello, world!"}],
        "max_tokens": 50,
    },
    timeout=30,
)
print(resp.json())

```

Switching between backends only requires changing the `model` field in the request to match the `model_id` defined in your [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml).

## Protocol Translation Architecture

Switchyard performs bidirectional translation via the `switchyard-translation` crate documented in [`crates/switchyard-translation/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md). When receiving an OpenAI Chat API request:

1. The server parses the incoming JSON and identifies the route via `model_id`.
2. The translation layer converts the payload if necessary (for NIM and Ollama, the OpenAI schema is passed through directly).
3. The HTTP client forwards the request to the backend’s `base_url`.
4. The response is translated back to the client’s expected format before returning.

This architecture allows any OpenAI-compatible client—including Claude Code, Codex CLI, or custom applications—to interact with NIM or Ollama without code changes.

## Summary

- Create a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) with OpenAI-compatible targets pointing to NIM (`https://integrate.api.nvidia.com/v1`) or Ollama (`http://127.0.0.1:11434/v1`) base URLs.
- Export `NVIDIA_API_KEY` for NIM authentication; Ollama requires no key.
- Use `passthrough` routes to map specific `model_id` values directly to backend targets.
- Run Switchyard via the launcher (`switchyard launch`) for agent integration or the standalone server (`switchyard-server`) for general HTTP proxying.
- Switchyard automatically translates requests and responses using the `switchyard-translation` crate, requiring no custom client code.

## Frequently Asked Questions

### Does Switchyard require custom code to support NIM or Ollama?

No. Both NIM and Ollama implement the OpenAI API schema, so Switchyard's existing OpenAI target type works without modification. The proxy handles request forwarding and response translation automatically via the `switchyard-translation` crate, as noted in the crate's README.

### What is the difference between using the launcher and the standalone server?

The launcher (`switchyard launch`) wraps the server and is optimized for coding agents like Claude Code, automatically injecting configuration and handling process lifecycle. The standalone server (`switchyard-server`) runs as a persistent HTTP proxy on a configurable host and port, suitable for any OpenAI-compatible client or integration.

### How does Switchyard handle authentication headers for NVIDIA NIM?

When the `NVIDIA_API_KEY` environment variable is present, Switchyard automatically adds an `Authorization: Bearer` header to requests sent to the NIM target, as implemented in the server logic referenced in [`AGENTS.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/AGENTS.md) at line 244. Ollama local instances typically operate without authentication headers.

### Can I route different models to different backends simultaneously?

Yes. Define multiple targets (one per backend) and multiple passthrough routes, each mapping a specific `model_id` to its respective target. Clients select the backend by specifying the corresponding model ID in their API calls, allowing you to serve NIM-hosted models and local Ollama models from a single Switchyard instance.