Switchyard Components: The Complete Architecture of NVIDIA’s LLM Router

Switchyard is a modular LLM router written in Rust with Python bindings, built around seven core crates—libsy, llm-client, runner, server, Python bindings, protocol, and translation—that together enable intelligent routing across multiple LLM providers while maintaining provider-agnostic data structures.

Developed under the NVIDIA-NeMo organization, Switchyard separates routing logic from transport concerns through a strict crate-based architecture. Each component in the crates/ directory handles a single responsibility, from core algorithmic decision-making to HTTP proxying and language interoperability.

Core Routing Engine: switchyard-libsy

The switchyard-libsy crate forms the algorithmic heart of the system. Located at crates/libsy/src/lib.rs, this crate exposes the Algorithm trait and its primary entry point Algorithm::run_stream, which drives the routing decision process.

According to the Switchyard source code, the core algorithm loop is implemented in crates/libsy/src/core/algorithm.rs. This crate operates entirely on normalized data structures, remaining agnostic to specific LLM providers or transport protocols.

Provider Communication: switchyard-llm-client

The switchyard-llm-client crate handles HTTP communication with upstream LLM providers. Its main entry point is the run function defined in crates/libsy-llm-client/src/run.rs.

This component feeds normalized responses back into the routing algorithms. It abstracts away provider-specific networking, retries, and credential management, allowing the core libsy algorithms to focus purely on routing logic.

Configuration and Orchestration: switchyard-runner

The switchyard-runner crate serves as the orchestration layer, parsing TOML routing configurations and wiring the client to the algorithm. The primary implementation resides in crates/switchyard-runner/src/runner.rs.

The runner drives the routing process until a model is selected, managing the lifecycle of algorithm execution and coordinating between the configuration file and runtime components.

HTTP Interface: switchyard-server

The switchyard-server crate provides a thin HTTP proxy that exposes OpenAI and Anthropic-compatible endpoints. The server implementation is located in crates/switchyard-server/src/lib.rs.

This component delegates incoming requests to the runner, effectively wrapping the entire routing pipeline behind a standard HTTP interface. Any OpenAI-compatible client can connect without modification.

Python Interoperability: switchyard-py

The switchyard-py crate provides Python bindings that expose the same routing primitives available in Rust. The Foreign Function Interface (FFI) layer is implemented in crates/switchyard-py/src/lib.rs.

These bindings allow Python applications to embed Switchyard directly, utilizing the stage_router algorithm and other components without leaving the Python runtime environment.

Shared Infrastructure

Protocol Crate

The protocol crate defines provider-neutral data structures including requests, responses, and streaming events. Defined in crates/protocol/src/lib.rs, these structures serve as the internal intermediate representation (IR) used across all Switchyard components.

Translation Layer

The switchyard-translation crate handles bidirectional conversion between vendor-specific JSON formats (OpenAI, Anthropic) and the internal protocol IR. The translation logic resides in crates/switchyard-translation/src/lib.rs.

This isolation ensures that vendor API changes affect only the translation layer, leaving core algorithms and routing logic unchanged.

Integration Ecosystem

Beyond the core seven crates, Switchyard provides optional integration plugins:

  • NeMo Relay Plugin: Embeds Switchyard directly into NeMo Relay deployments (crates/switchyard-nemo-relay-plugin/)
  • LiteLLM Routing Plugin: Allows LiteLLM to use Switchyard as its decision engine (examples/litellm/)

Component Interaction Flow

The Switchyard components operate in a specific pipeline:

  1. Request Normalization: switchyard-py or direct Rust code builds a protocol::Request using shared data structures from the protocol crate.

  2. Configuration Loading: The switchyard-runner parses the TOML configuration and instantiates the appropriate algorithm from libsy.

  3. Algorithm Execution: The algorithm in switchyard-libsy may invoke the LLM client (libsy-llm-client) multiple times, applying judges, stages, or escalation logic.

  4. Response Translation: switchyard-translation converts provider-specific JSON into the internal IR before returning it to the algorithm.

  5. HTTP Serving: When using switchyard-server, the entire pipeline is exposed behind OpenAI-compatible endpoints.

Implementation Examples

Embedding Switchyard in Python

from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# Build an algorithm (stage router)

algorithm = stage_router(
    capable="capable",
    efficient="efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

# Helper to call a real model

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception:
            continue
    raise RuntimeError("All candidates failed")

# Drive the algorithm

async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                try:
                    call.respond(await call_with_fallback(call.request, call.models, clients))
                except Exception as e:
                    call.fail(e)
            case Step.Done(outcome):
                if outcome.response:
                    return outcome.response
                return await call_with_fallback(
                    outcome.request, 
                    outcome.selected_model_ids, 
                    clients
                )
    raise RuntimeError("Algorithm ended without a decision")

Running the Standalone Proxy


# Install the server

cargo install --locked switchyard-server

# Create a TOML configuration

cat > routes.toml <<'TOML'
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
TOML

# Start the proxy

export OPENROUTER_API_KEY="your_key_here"
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

Summary

Frequently Asked Questions

What is the primary purpose of Switchyard?

Switchyard serves as a modular LLM router that intelligently directs requests across multiple AI providers based on configurable algorithms. It separates routing logic from transport concerns, allowing teams to implement complex fallback strategies, cost optimization, and capability-based routing without modifying application code.

How does Switchyard handle different LLM provider APIs?

The switchyard-translation crate handles all provider-specific format conversions. It translates vendor-specific JSON from OpenAI, Anthropic, and other providers into a provider-neutral intermediate representation defined in the protocol crate. This ensures that core routing algorithms remain agnostic to specific API formats.

Can Switchyard be used without the HTTP server component?

Yes. While switchyard-server provides a convenient OpenAI-compatible HTTP interface, you can embed Switchyard directly in Rust or Python applications using switchyard-libsy and switchyard-py. The switchyard-runner can be invoked programmatically without the HTTP layer, making it suitable for internal service integration.

Which files contain the most critical routing logic?

The algorithmic core resides in crates/libsy/src/core/algorithm.rs, while the public API is exposed through crates/libsy/src/lib.rs. Request processing and provider communication are handled in crates/libsy-llm-client/src/run.rs, and the orchestration layer that binds these together is located in crates/switchyard-runner/src/runner.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →