Typical Workflow for Using Switchyard: Three Integration Methods Explained

Switchyard integrates into LLM pipelines via the NeMo Relay plugin, embedded Rust/Python libraries, or a stand-alone OpenAI-compatible proxy server, with all paths using a common routing-algorithm abstraction defined in crates/libsy/src/lib.rs.

The NVIDIA-NeMo/Switchyard repository provides intelligent routing for large language model requests, allowing systems to dynamically select between efficient and capable models. Understanding the typical workflow for using Switchyard requires choosing from three distinct integration paths that share a unified algorithmic core. Each path leverages the same routing-algorithm abstraction—such as stage_router, llm_classifier, or advisor—to decide whether to route requests to an efficient model, a capable model, or a fallback target.

Path 1: Deploy the NeMo Relay Plugin

The first integration method involves building the native switchyard-nemo-relay-plugin and registering it with a running NeMo Relay deployment. This plugin intercepts every request that Relay receives and routes it through Switchyard’s algorithms before forwarding to the target LLM.

To implement this path, build the plugin crate located at crates/switchyard-nemo-relay-plugin/ and configure it to point at a routes.toml file. The plugin’s implementation in crates/switchyard-nemo-relay-plugin/src/lib.rs handles the native interface with Relay, while the routing logic itself is delegated to the shared library. This approach is ideal if you already operate a NeMo Relay infrastructure and want to add intelligent routing without modifying your client code.

Path 2: Embed the Library in Your Application

The second path embeds Switchyard directly into your own harness using the switchyard-libsy crate (or its Python bindings). Your code remains in charge of making actual model calls; Switchyard only tells you which model to invoke and when to finalize the response. The core API surface lives in crates/libsy/src/lib.rs, exposing the Algorithm trait and Step enum.

Python Integration

Add the Switchyard Python bindings to your project and construct a routing algorithm. Drive the logic using algorithm.run_stream, handling Step.CallModel to invoke your own client code and Step.Done to return the final response.

from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# 1️⃣ Build the routing algorithm

algorithm = stage_router(
    "capable",      # name of the capable target

    "efficient",   # name of the efficient target

    picker="efficient_first",
    confidence_threshold=0.5,
)

# 2️⃣ Helper that calls a real model client (you provide the client dict)

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception:
            continue
    raise RuntimeError("All candidate models failed")

# 3️⃣ Drive the algorithm

async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                try:
                    call.respond(await call_with_fallback(call.request, call.models, clients))
                except Exception as err:
                    call.fail(err)
            case Step.Done(outcome):
                # If the algorithm already produced an answer, return it

                if outcome.response is not None:
                    return outcome.response
                # Otherwise make the final model call ourselves

                return await call_with_fallback(
                    outcome.request, outcome.selected_model_ids, clients
                )
    raise RuntimeError("Algorithm terminated without a decision")

Rust Integration

For Rust applications, import the switchyard-libsy crate and use Algorithm::run_stream to process requests. The StageRouterConfig struct allows you to specify target models, picker strategies, and confidence thresholds.

use switchyard_libsy::{
    algorithms::stage::StageRouterConfig,
    libsy::{Algorithm, Step, StepStream},
    libsy::Outcome,
};

// 1️⃣ Build the algorithm
let config = StageRouterConfig {
    capable_target: "capable".into(),
    efficient_target: "efficient".into(),
    picker: switchyard_libsy::algorithms::util::stage::PickerMode::EfficientFirst,
    confidence_threshold: 0.5,
    ..Default::default()
};
let algorithm = StageRouter::new(config);

// 2️⃣ Normalized request (see `switchyard-protocol` for the shape)
let request = serde_json::json!({
    "messages": [{ "role": "user", "content": [{ "type": "text", "text": "hello" }] }],
    "model": "switchyard"
});

// 3️⃣ Drive the algorithm
let mut stream = algorithm.run_stream(request);
while let Some(step) = stream.next().await {
    match step {
        Step::CallModel(call) => {
            // Your code calls the appropriate model client and returns a normalized response
            let resp = my_client.call(call.request.clone()).await?;
            call.respond(resp);
        }
        Step::Done(outcome) => {
            // `outcome.response` may already contain the final answer
            if let Some(resp) = outcome.response {
                return Ok(resp);
            }
        }
    }
}

The crates/libsy-llm-client/src/lib.rs file provides a convenience run helper that automates driving an Algorithm with real model calls, though embedded harnesses typically implement their own client logic as shown above.

Path 3: Run a Stand-alone Proxy Server

The third path treats Switchyard as a drop-in OpenAI- or Anthropic-compatible proxy. Install the switchyard-server binary from crates/switchyard-server/, define your targets and routing logic in a routes.toml file, and start the server. Any compliant client can then point at the proxy endpoint (e.g., http://localhost:4000/v1/chat/completions).

The entry point for the server is crates/switchyard-server/src/main.rs, which handles HTTP transport and request normalization. The crates/switchyard-translation/src/lib.rs crate translates between vendor-specific payloads (OpenAI, Anthropic) and the internal protocol used by Switchyard.


# Install the server binary

cargo install --locked switchyard-server

# Write a minimal routes.toml (see docs for full schema)

cat > routes.toml <<'TOML'
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
TOML

# Start the proxy (dry‑run first to validate)

export OPENROUTER_API_KEY="your-key"
switchyard-server --config routes.toml --dry-run
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

# Query the proxy as an OpenAI client

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"hello"}]}'

Core Routing Abstraction and Protocol

All three integration paths rely on a common decision flow implemented in the core library. The Algorithm trait defines a run_stream method that yields Step variants: Step::CallModel requests that the harness invoke a classifier or judge model, while Step::Done signals that the algorithm has either produced a final response or determined which target model should generate it.

The normalized request and response schema used throughout the system is defined in crates/protocol/src/lib.rs. This protocol ensures that routing decisions are based on a consistent data structure regardless of whether the request originated from an OpenAI-compatible client, a NeMo Relay instance, or a direct library call. The decision flow typically follows this pattern:

  1. Ingest a normalized request.
  2. Loop through Algorithm::run_stream, handling Step::CallModel by invoking the requested model(s).
  3. Return the final Outcome when Step::Done is reached, either using the pre-computed response or delegating to the selected efficient/capable model.

Summary

Frequently Asked Questions

What configuration file format does Switchyard use?

Switchyard uses a TOML configuration file, conventionally named routes.toml, to define LLM clients, model targets, and routing rules. The schema supports defining provider-specific settings (such as base URLs and API key environment variables) under [llm_clients], model identifiers under [targets], and algorithm configurations (including stage_router parameters) under [routes]. Full schema details are documented in docs/reference/toml_schema.md.

How do I choose between the three integration paths?

Choose the NeMo Relay plugin (Path 1) if you already run NVIDIA's NeMo Relay and want transparent routing without client changes. Choose embedded library (Path 2) if you need fine-grained control over model execution, custom authentication, or integration into an existing Rust or Python application. Choose the stand-alone proxy (Path 3) if you need a drop-in replacement for OpenAI-compatible endpoints that works with any HTTP client.

How does Switchyard decide which model to route a request to?

Switchyard uses routing algorithms such as stage_router, llm_classifier, or advisor defined in the configuration. These algorithms analyze the request (potentially calling intermediate "judge" models) to determine whether an efficient model is sufficient or if a more capable model is required. The decision is based on confidence thresholds and picker strategies (e.g., efficient_first) specified in the routes.toml file.

Is there a Python API available for Switchyard?

Yes, Switchyard provides Python bindings for the core library. You can import switchyard.libsy to construct algorithms like stage_router and drive them using algorithm.run_stream. The Python API mirrors the Rust API, yielding Step objects that your code handles to make actual model calls and return normalized responses.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →