# How to Integrate Switchyard with NVIDIA AI Tools: 3 Integration Methods Explained

> Discover how to integrate Switchyard with NVIDIA AI tools using three methods: NeMo Relay plugin, direct library embedding, or an OpenAI-compatible proxy. Optimize your AI workflows.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Switchyard integrates with NVIDIA AI tools through the NeMo Relay plugin for zero-code deployment, direct library embedding for custom Python or Rust harnesses, or a standalone OpenAI-compatible proxy server.**

Switchyard is a Rust-based LLM routing engine with Python bindings that intelligently selects models for each request. According to the NVIDIA-NeMo/Switchyard source code, it plugs directly into the NeMo Relay framework, LiteLLM, or custom NVIDIA AI harnesses to optimize cost and performance across model providers.

## Architecture Overview for NVIDIA AI Integration

Switchyard operates as a set of Rust crates that route LLM requests to specific targets based on configurable algorithms. The architecture separates concerns across six primary components that process requests from reception to provider execution.

### Core Routing and Decision Engine

The **libsy** crate in [`crates/libsy/src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms.rs) implements the decision-making logic through the `Algorithm::run_stream` method. This component evaluates requests against configured routes and selects the appropriate model target based on efficiency, capability, or custom business logic.

### Provider Communication Layer

The **libsy-llm-client** crate located at [`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs) handles HTTP communication with LLM providers and records telemetry observations. It executes the selected target's request and returns usage metrics for routing feedback loops.

### Configuration and Route Management

Deployment configurations load through the **switchyard-runner** crate in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs). The `Runner::load` function parses version-1 TOML deployments, while `Runner::route` resolves model identifiers to their configured targets.

### Protocol Translation

The **switchyard-translation** crate at [`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs) bridges provider-specific formats (OpenAI, Anthropic) with the neutral **switchyard-protocol** IR defined in [`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs). This enables Switchyard to accept OpenAI-compatible requests while routing to any supported backend.

### NeMo Relay Integration

The **switchyard-nemo-relay-plugin** in [`crates/switchyard-nemo-relay-plugin/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/lib.rs) registers two interceptors—`register_buffered` and `register_stream`—that forward Relay calls through Switchyard's `SwitchyardRuntime`.

## Method 1: NeMo Relay Plugin Integration

The NeMo Relay plugin provides the tightest integration with NVIDIA's agent framework, requiring no changes to existing agent code. When deployed, Relay automatically routes any request specifying `model: "switchyard"` through the Switchyard algorithm, wiring translation, routing, execution, and telemetry into Relay's request-intercept API.

Configure your routes and targets in a TOML file:

```toml

# /etc/switchyard/routes.toml

schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5

```

Register the plugin with NeMo Relay:

```toml
[[plugins.dynamic]]
manifest = "./plugins/switchyard/relay-plugin.toml"

[plugins.dynamic.config]
priority = 0
switchyard_config_path = "/etc/switchyard/routes.toml"

[plugins.policy.overrides."nvidia.switchyard"]
attestation = "integrity_only"

```

After building and registering the plugin with `nemo-relay plugins add`, Relay handles all Switchyard routing transparently.

## Method 2: Embedded Library Integration

For custom NVIDIA AI tools or LiteLLM deployments, embed Switchyard directly as a library. Install the Python bindings from the repository:

```bash
pip install git+https://github.com/NVIDIA-NeMo/Switchyard.git

```

The following Python example demonstrates embedding the `stage_router` algorithm in an async harness:

```python
from switchyard.libsy import Step
from switchyard.libsy.algorithms import stage_router
from switchyard.libsy import LlmResponse
from switchyard.protocol import Request
import aiohttp

# Initialize the routing algorithm

algorithm = stage_router(
    capable_target="capable",
    efficient_target="efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

# Create a normalized protocol request

request = Request(
    llm_request={
        "model": "switchyard",
        "messages": [{"role": "user", "content": "What is the capital of France?"}],
    },
    metadata=None,
)

# Provider call implementation

async def call_provider(model, payload):
    async with aiohttp.ClientSession() as sess:
        async with sess.post(
            f"https://openrouter.ai/v1/{model}", 
            json=payload
        ) as resp:
            return await resp.json()

# Routing execution loop

async def route(request):
    async for step in algorithm.run_stream(request):
        if isinstance(step, Step.CallModel):
            raw = await call_provider(
                step.request["model"], 
                step.request
            )
            step.respond(LlmResponse.Agg(raw))
        elif isinstance(step, Step.Done):
            return step.outcome.response

# Execute

response = await route(request)

```

The same pattern applies in Rust using `Algorithm::run_stream` from `switchyard-libsy`.

## Method 3: Standalone Proxy Server

For tools that require an OpenAI-compatible endpoint without code modification, run Switchyard as a standalone HTTP proxy. This approach works with any NVIDIA AI tool that supports custom base URLs.

Install and launch the server:

```bash
cargo install --locked switchyard-server
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

Point any OpenAI-compatible client at the proxy:

```bash
export OPENAI_BASE_URL="http://localhost:4000/v1"
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"Hello"}]}'

```

The proxy uses the same `Runner` and `Algorithm` stack as the embedded library, ensuring consistent routing decisions across all integration methods.

## Key Source Files for Custom Integration

When building custom integrations with NVIDIA AI tools, reference these specific source locations:

- **[`crates/libsy/src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms.rs)**: Implements `Algorithm::run_stream` and routing logic including `stage_router`
- **[`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs)**: HTTP client implementation for provider communication and `RunObservation` reporting
- **[`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs)**: TOML deployment loader (`Runner::load`) and route resolution (`Runner::route`)
- **[`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs)**: Request/response encoding for OpenAI and Anthropic formats
- **[`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs)**: Neutral IR types including `Request`, `Response`, and `Usage`
- **[`crates/switchyard-nemo-relay-plugin/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/lib.rs)**: Relay plugin entry point with `register_buffered` and `register_stream` interceptors
- **[`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs)**: Standalone proxy implementation

## Summary

- **Switchyard** routes LLM requests through Rust-based algorithms with Python bindings for NVIDIA AI integration
- **Three integration paths** exist: NeMo Relay plugin (zero-code), embedded library (Python/Rust), and standalone proxy (OpenAI-compatible)
- **Core flow** involves loading TOML deployments via `Runner::load`, translating requests to neutral IR, running `Algorithm::run_stream`, executing via `libsy-llm-client`, and translating responses back
- **NeMo Relay plugin** in [`crates/switchyard-nemo-relay-plugin/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/lib.rs) provides seamless integration for existing Relay deployments through `register_buffered` and `register_stream` interceptors
- **Protocol translation** in [`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs) enables support for OpenAI and Anthropic formats without vendor lock-in

## Frequently Asked Questions

### Can Switchyard integrate with existing NeMo Relay deployments without code changes?

Yes. The **switchyard-nemo-relay-plugin** registers interceptors that hook into Relay's request-processing pipeline automatically. By adding the plugin configuration to your Relay deployment and building the plugin artifact, existing agents can route through Switchyard by simply requesting `model: "switchyard"` without any code modifications.

### What algorithm does Switchyard use to select between capable and efficient models?

Switchyard implements multiple algorithms in [`crates/libsy/src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms.rs), including **stage_router** which uses a configurable picker strategy such as `efficient_first`. This attempts the efficient target first and falls back to the capable target based on the configured `confidence_threshold`, balancing cost and performance according to your TOML configuration.

### Does Switchyard support async streaming responses for NVIDIA AI applications?

Yes. The `Algorithm::run_stream` method returns an async stream of `Step` variants including `Step.CallModel` and `Step.Done`. The NeMo Relay plugin provides both `register_buffered` and `register_stream` interceptors to handle streaming and non-streaming request patterns, making it compatible with real-time NVIDIA AI inference pipelines.

### Can I use Switchyard with providers other than NVIDIA's models?

Absolutely. Switchyard's provider-neutral architecture in [`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs) supports any OpenAI-compatible or Anthropic-compatible endpoint. The `switchyard-translation` crate handles format conversion, while `libsy-llm-client` executes HTTP requests to arbitrary base URLs defined in your TOML configuration, enabling routing across third-party providers like OpenRouter, Anthropic, or custom endpoints.