# Where to Find Switchyard Documentation: Complete Guide to NVIDIA-NeMo's LLM Router

> Find NVIDIA NeMo Switchyard documentation easily. Explore the official repository for installation guides and configuration schemas to get started with your LLM router.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: getting-started
- Published: 2026-09-11

---

**Switchyard documentation is distributed across the `docs/` and `crates/` directories of the NVIDIA-NeMo/Switchyard repository, with primary entry points including [`docs/getting_started.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/getting_started.md) for installation and [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) for configuration schemas.**

Switchyard is NVIDIA's open-source framework for intelligent LLM routing and request optimization. The **Switchyard documentation** is maintained entirely within the repository as modular markdown files, covering everything from high-level architecture to specific implementation paths. Whether you are embedding routing logic in Python, deploying a standalone proxy server, or integrating with NeMo Relay, the documentation provides file-specific references and runnable code examples.

## Core Documentation Structure

The repository organizes documentation into two primary locations. The `docs/` directory contains conceptual guides and reference materials, while the `crates/` directory houses component-specific README files for the Rust-based implementation.

**Key documentation files in the repository:**

- **[`docs/getting_started.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/getting_started.md)** – Installation instructions for Python, Rust, and the standalone server
- **[`docs/core_concepts.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/core_concepts.md)** – Definitions of LLM clients, targets, routes, and model ID formats
- **[`docs/architecture.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/architecture.md)** – Visual explanation of the proxy flow and request lifecycle
- **[`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md)** – Detailed specifications for each routing algorithm
- **[`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md)** – Complete schema reference for [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration
- **[`docs/cli_reference.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/cli_reference.md)** – Command-line options for the standalone server

## Component-Specific Documentation

Each major component of Switchyard maintains its own documentation within the crate structure. These files contain API references and implementation details specific to their use case.

### Python Library

The **[`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md)** documents how to embed Switchyard routing logic directly into Python applications. This crate provides bindings for Rust algorithms, allowing you to instantiate routers like `stage_router` and handle `LlmResponse` objects natively.

The repository includes a complete working example in **[`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py)**, which demonstrates async request handling:

```python
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# Build a stage-router that prefers the "efficient" model first

algorithm = stage_router(
    capable_target="capable",
    efficient_target="efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

# Helper to call a concrete model client

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception:
            continue
    raise RuntimeError("All candidates failed")

# Drive the algorithm

async def route(request: dict, clients: dict) -> LlmResponse.Agg:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                call.respond(await call_with_fallback(call.request, call.models, clients))
            case Step.Done(outcome):
                return outcome.response or await call_with_fallback(
                    outcome.request, outcome.selected_model_ids, clients
                )
    raise RuntimeError("Algorithm finished without a decision")

```

### Standalone Server

For deploying Switchyard as an HTTP proxy, reference **[`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md)**. This documentation covers binary installation via Cargo, TOML configuration syntax, and server startup parameters.

A minimal server configuration requires a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file:

```toml
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5

```

Start the server using the CLI documented in **[`docs/cli_reference.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/cli_reference.md)**:

```bash
cargo install --locked switchyard-server
export OPENROUTER_API_KEY="your-key"
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

### NeMo Relay Plugin

Integration with NVIDIA NeMo Relay is documented in **[`crates/switchyard-nemo-relay-plugin/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/README.md)**. This plugin allows Relay to use Switchyard as a native routing backend.

To enable the integration:

1. Build the plugin according to the crate README instructions
2. Create a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration file on the Relay host
3. Add the plugin configuration to Relay's [`relay-plugin.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/relay-plugin.toml):

```toml
[[plugins.dynamic]]
manifest = "./plugins/switchyard/relay-plugin.toml"

[plugins.dynamic.config]
priority = 0
switchyard_config_path = "/etc/switchyard/routes.toml"

[plugins.policy.overrides."nvidia.switchyard"]
attestation = "integrity_only"

```

4. Enable via CLI:

```bash
nemo-relay plugins enable nvidia.switchyard
nemo-relay plugins validate nvidia.switchyard

```

## Routing Algorithms and Configuration

The **[`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md)** file provides detailed explanations of each algorithm implemented in Switchyard, including the `stage_router` used in the examples above. This documentation explains parameters like `confidence_threshold`, `capable_target`, and `efficient_target`, and how the `picker` option determines model selection strategy.

For complete configuration syntax, **[`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md)** defines every valid key for [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml), including:
- `schema_version` (currently `1`)
- `llm_clients` sections for provider configuration
- `targets` definitions mapping model IDs to clients
- `routes` sections binding algorithms to target pairs

## Protocol and Type Definitions

Low-level request and response types are documented in **[`crates/protocol/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/README.md)**. This file defines the provider-neutral structures used throughout the Switchyard stack, ensuring consistent data handling between the Python library, standalone server, and NeMo Relay plugin.

## Summary

- **Primary documentation** resides in `docs/` for guides and `crates/*/README.md` for API references
- **Installation** instructions are in [`docs/getting_started.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/getting_started.md)
- **Configuration** schemas are documented in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md)
- **Python embedding** uses [`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) and [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py)
- **Standalone deployment** references [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md) and [`docs/cli_reference.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/cli_reference.md)
- **NeMo Relay integration** is covered in [`crates/switchyard-nemo-relay-plugin/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/README.md)

## Frequently Asked Questions

### Where is the official Switchyard documentation hosted?

The official **Switchyard documentation** is maintained entirely within the NVIDIA-NeMo/Switchyard GitHub repository. Unlike external documentation sites, all guides, API references, and schemas live as markdown files in the `docs/` directory and crate-specific README files, ensuring version accuracy with the source code.

### How do I configure the TOML schema for Switchyard routing?

Configuration uses a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file with a `schema_version` key set to `1`, followed by `[llm_clients]`, `[targets]`, and `[routes]` sections. The complete schema including all supported keys, value types, and validation rules is documented in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md), while algorithm-specific parameters like `confidence_threshold` and `picker` are explained in [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md).

### Can I use Switchyard as a Python library instead of a server?

Yes. The `libsy` crate documented in [`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) provides Python bindings that allow you to embed routing logic directly into applications. You can instantiate algorithms like `stage_router`, process `LlmResponse` objects, and handle model fallback logic without running a separate proxy server, as demonstrated in [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py).

### How do I integrate Switchyard with NeMo Relay?

Integration requires building the plugin from `crates/switchyard-nemo-relay-plugin` according to that README, placing a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file on the Relay host, and updating [`relay-plugin.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/relay-plugin.toml) to point `switchyard_config_path` to that file. After validation with `nemo-relay plugins validate nvidia.switchyard`, Relay will route all LLM requests through Switchyard's algorithms.