How to Integrate Switchyard with NVIDIA AI Tools: 3 Integration Methods Explained
Switchyard integrates with NVIDIA AI tools through the NeMo Relay plugin for zero-code deployment, direct library embedding for custom Python or Rust harnesses, or a standalone OpenAI-compatible proxy server.
Switchyard is a Rust-based LLM routing engine with Python bindings that intelligently selects models for each request. According to the NVIDIA-NeMo/Switchyard source code, it plugs directly into the NeMo Relay framework, LiteLLM, or custom NVIDIA AI harnesses to optimize cost and performance across model providers.
Architecture Overview for NVIDIA AI Integration
Switchyard operates as a set of Rust crates that route LLM requests to specific targets based on configurable algorithms. The architecture separates concerns across six primary components that process requests from reception to provider execution.
Core Routing and Decision Engine
The libsy crate in crates/libsy/src/algorithms.rs implements the decision-making logic through the Algorithm::run_stream method. This component evaluates requests against configured routes and selects the appropriate model target based on efficiency, capability, or custom business logic.
Provider Communication Layer
The libsy-llm-client crate located at crates/libsy-llm-client/src/client.rs handles HTTP communication with LLM providers and records telemetry observations. It executes the selected target's request and returns usage metrics for routing feedback loops.
Configuration and Route Management
Deployment configurations load through the switchyard-runner crate in crates/switchyard-runner/src/runner.rs. The Runner::load function parses version-1 TOML deployments, while Runner::route resolves model identifiers to their configured targets.
Protocol Translation
The switchyard-translation crate at crates/switchyard-translation/src/engine.rs bridges provider-specific formats (OpenAI, Anthropic) with the neutral switchyard-protocol IR defined in crates/protocol/src/lib.rs. This enables Switchyard to accept OpenAI-compatible requests while routing to any supported backend.
NeMo Relay Integration
The switchyard-nemo-relay-plugin in crates/switchyard-nemo-relay-plugin/src/lib.rs registers two interceptors—register_buffered and register_stream—that forward Relay calls through Switchyard's SwitchyardRuntime.
Method 1: NeMo Relay Plugin Integration
The NeMo Relay plugin provides the tightest integration with NVIDIA's agent framework, requiring no changes to existing agent code. When deployed, Relay automatically routes any request specifying model: "switchyard" through the Switchyard algorithm, wiring translation, routing, execution, and telemetry into Relay's request-intercept API.
Configure your routes and targets in a TOML file:
# /etc/switchyard/routes.toml
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"
[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"
[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
Register the plugin with NeMo Relay:
[[plugins.dynamic]]
manifest = "./plugins/switchyard/relay-plugin.toml"
[plugins.dynamic.config]
priority = 0
switchyard_config_path = "/etc/switchyard/routes.toml"
[plugins.policy.overrides."nvidia.switchyard"]
attestation = "integrity_only"
After building and registering the plugin with nemo-relay plugins add, Relay handles all Switchyard routing transparently.
Method 2: Embedded Library Integration
For custom NVIDIA AI tools or LiteLLM deployments, embed Switchyard directly as a library. Install the Python bindings from the repository:
pip install git+https://github.com/NVIDIA-NeMo/Switchyard.git
The following Python example demonstrates embedding the stage_router algorithm in an async harness:
from switchyard.libsy import Step
from switchyard.libsy.algorithms import stage_router
from switchyard.libsy import LlmResponse
from switchyard.protocol import Request
import aiohttp
# Initialize the routing algorithm
algorithm = stage_router(
capable_target="capable",
efficient_target="efficient",
picker="efficient_first",
confidence_threshold=0.5,
)
# Create a normalized protocol request
request = Request(
llm_request={
"model": "switchyard",
"messages": [{"role": "user", "content": "What is the capital of France?"}],
},
metadata=None,
)
# Provider call implementation
async def call_provider(model, payload):
async with aiohttp.ClientSession() as sess:
async with sess.post(
f"https://openrouter.ai/v1/{model}",
json=payload
) as resp:
return await resp.json()
# Routing execution loop
async def route(request):
async for step in algorithm.run_stream(request):
if isinstance(step, Step.CallModel):
raw = await call_provider(
step.request["model"],
step.request
)
step.respond(LlmResponse.Agg(raw))
elif isinstance(step, Step.Done):
return step.outcome.response
# Execute
response = await route(request)
The same pattern applies in Rust using Algorithm::run_stream from switchyard-libsy.
Method 3: Standalone Proxy Server
For tools that require an OpenAI-compatible endpoint without code modification, run Switchyard as a standalone HTTP proxy. This approach works with any NVIDIA AI tool that supports custom base URLs.
Install and launch the server:
cargo install --locked switchyard-server
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
Point any OpenAI-compatible client at the proxy:
export OPENAI_BASE_URL="http://localhost:4000/v1"
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"switchyard","messages":[{"role":"user","content":"Hello"}]}'
The proxy uses the same Runner and Algorithm stack as the embedded library, ensuring consistent routing decisions across all integration methods.
Key Source Files for Custom Integration
When building custom integrations with NVIDIA AI tools, reference these specific source locations:
crates/libsy/src/algorithms.rs: ImplementsAlgorithm::run_streamand routing logic includingstage_routercrates/libsy-llm-client/src/client.rs: HTTP client implementation for provider communication andRunObservationreportingcrates/switchyard-runner/src/runner.rs: TOML deployment loader (Runner::load) and route resolution (Runner::route)crates/switchyard-translation/src/engine.rs: Request/response encoding for OpenAI and Anthropic formatscrates/protocol/src/lib.rs: Neutral IR types includingRequest,Response, andUsagecrates/switchyard-nemo-relay-plugin/src/lib.rs: Relay plugin entry point withregister_bufferedandregister_streaminterceptorscrates/switchyard-server/src/main.rs: Standalone proxy implementation
Summary
- Switchyard routes LLM requests through Rust-based algorithms with Python bindings for NVIDIA AI integration
- Three integration paths exist: NeMo Relay plugin (zero-code), embedded library (Python/Rust), and standalone proxy (OpenAI-compatible)
- Core flow involves loading TOML deployments via
Runner::load, translating requests to neutral IR, runningAlgorithm::run_stream, executing vialibsy-llm-client, and translating responses back - NeMo Relay plugin in
crates/switchyard-nemo-relay-plugin/src/lib.rsprovides seamless integration for existing Relay deployments throughregister_bufferedandregister_streaminterceptors - Protocol translation in
crates/switchyard-translation/src/engine.rsenables support for OpenAI and Anthropic formats without vendor lock-in
Frequently Asked Questions
Can Switchyard integrate with existing NeMo Relay deployments without code changes?
Yes. The switchyard-nemo-relay-plugin registers interceptors that hook into Relay's request-processing pipeline automatically. By adding the plugin configuration to your Relay deployment and building the plugin artifact, existing agents can route through Switchyard by simply requesting model: "switchyard" without any code modifications.
What algorithm does Switchyard use to select between capable and efficient models?
Switchyard implements multiple algorithms in crates/libsy/src/algorithms.rs, including stage_router which uses a configurable picker strategy such as efficient_first. This attempts the efficient target first and falls back to the capable target based on the configured confidence_threshold, balancing cost and performance according to your TOML configuration.
Does Switchyard support async streaming responses for NVIDIA AI applications?
Yes. The Algorithm::run_stream method returns an async stream of Step variants including Step.CallModel and Step.Done. The NeMo Relay plugin provides both register_buffered and register_stream interceptors to handle streaming and non-streaming request patterns, making it compatible with real-time NVIDIA AI inference pipelines.
Can I use Switchyard with providers other than NVIDIA's models?
Absolutely. Switchyard's provider-neutral architecture in crates/protocol/src/lib.rs supports any OpenAI-compatible or Anthropic-compatible endpoint. The switchyard-translation crate handles format conversion, while libsy-llm-client executes HTTP requests to arbitrary base URLs defined in your TOML configuration, enabling routing across third-party providers like OpenRouter, Anthropic, or custom endpoints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →