Where to Find Switchyard Documentation: Complete Guide to NVIDIA-NeMo's LLM Router

Switchyard documentation is distributed across the docs/ and crates/ directories of the NVIDIA-NeMo/Switchyard repository, with primary entry points including docs/getting_started.md for installation and docs/reference/toml_schema.md for configuration schemas.

Switchyard is NVIDIA's open-source framework for intelligent LLM routing and request optimization. The Switchyard documentation is maintained entirely within the repository as modular markdown files, covering everything from high-level architecture to specific implementation paths. Whether you are embedding routing logic in Python, deploying a standalone proxy server, or integrating with NeMo Relay, the documentation provides file-specific references and runnable code examples.

Core Documentation Structure

The repository organizes documentation into two primary locations. The docs/ directory contains conceptual guides and reference materials, while the crates/ directory houses component-specific README files for the Rust-based implementation.

Key documentation files in the repository:

Component-Specific Documentation

Each major component of Switchyard maintains its own documentation within the crate structure. These files contain API references and implementation details specific to their use case.

Python Library

The crates/libsy/README.md documents how to embed Switchyard routing logic directly into Python applications. This crate provides bindings for Rust algorithms, allowing you to instantiate routers like stage_router and handle LlmResponse objects natively.

The repository includes a complete working example in examples/libsy.py, which demonstrates async request handling:

from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# Build a stage-router that prefers the "efficient" model first

algorithm = stage_router(
    capable_target="capable",
    efficient_target="efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

# Helper to call a concrete model client

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception:
            continue
    raise RuntimeError("All candidates failed")

# Drive the algorithm

async def route(request: dict, clients: dict) -> LlmResponse.Agg:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                call.respond(await call_with_fallback(call.request, call.models, clients))
            case Step.Done(outcome):
                return outcome.response or await call_with_fallback(
                    outcome.request, outcome.selected_model_ids, clients
                )
    raise RuntimeError("Algorithm finished without a decision")

Standalone Server

For deploying Switchyard as an HTTP proxy, reference crates/switchyard-server/README.md. This documentation covers binary installation via Cargo, TOML configuration syntax, and server startup parameters.

A minimal server configuration requires a routes.toml file:

schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5

Start the server using the CLI documented in docs/cli_reference.md:

cargo install --locked switchyard-server
export OPENROUTER_API_KEY="your-key"
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

NeMo Relay Plugin

Integration with NVIDIA NeMo Relay is documented in crates/switchyard-nemo-relay-plugin/README.md. This plugin allows Relay to use Switchyard as a native routing backend.

To enable the integration:

  1. Build the plugin according to the crate README instructions
  2. Create a routes.toml configuration file on the Relay host
  3. Add the plugin configuration to Relay's relay-plugin.toml:
[[plugins.dynamic]]
manifest = "./plugins/switchyard/relay-plugin.toml"

[plugins.dynamic.config]
priority = 0
switchyard_config_path = "/etc/switchyard/routes.toml"

[plugins.policy.overrides."nvidia.switchyard"]
attestation = "integrity_only"
  1. Enable via CLI:
nemo-relay plugins enable nvidia.switchyard
nemo-relay plugins validate nvidia.switchyard

Routing Algorithms and Configuration

The docs/routing_algorithms/overview.md file provides detailed explanations of each algorithm implemented in Switchyard, including the stage_router used in the examples above. This documentation explains parameters like confidence_threshold, capable_target, and efficient_target, and how the picker option determines model selection strategy.

For complete configuration syntax, docs/reference/toml_schema.md defines every valid key for routes.toml, including:

  • schema_version (currently 1)
  • llm_clients sections for provider configuration
  • targets definitions mapping model IDs to clients
  • routes sections binding algorithms to target pairs

Protocol and Type Definitions

Low-level request and response types are documented in crates/protocol/README.md. This file defines the provider-neutral structures used throughout the Switchyard stack, ensuring consistent data handling between the Python library, standalone server, and NeMo Relay plugin.

Summary

Frequently Asked Questions

Where is the official Switchyard documentation hosted?

The official Switchyard documentation is maintained entirely within the NVIDIA-NeMo/Switchyard GitHub repository. Unlike external documentation sites, all guides, API references, and schemas live as markdown files in the docs/ directory and crate-specific README files, ensuring version accuracy with the source code.

How do I configure the TOML schema for Switchyard routing?

Configuration uses a routes.toml file with a schema_version key set to 1, followed by [llm_clients], [targets], and [routes] sections. The complete schema including all supported keys, value types, and validation rules is documented in docs/reference/toml_schema.md, while algorithm-specific parameters like confidence_threshold and picker are explained in docs/routing_algorithms/overview.md.

Can I use Switchyard as a Python library instead of a server?

Yes. The libsy crate documented in crates/libsy/README.md provides Python bindings that allow you to embed routing logic directly into applications. You can instantiate algorithms like stage_router, process LlmResponse objects, and handle model fallback logic without running a separate proxy server, as demonstrated in examples/libsy.py.

How do I integrate Switchyard with NeMo Relay?

Integration requires building the plugin from crates/switchyard-nemo-relay-plugin according to that README, placing a routes.toml file on the Relay host, and updating relay-plugin.toml to point switchyard_config_path to that file. After validation with nemo-relay plugins validate nvidia.switchyard, Relay will route all LLM requests through Switchyard's algorithms.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →