# Typical Workflow for Using Switchyard: Three Integration Methods Explained

> Discover the typical Switchyard workflow integrating LLM pipelines through NeMo Relay, Rust/Python libraries, or a proxy server. Learn common routing algorithms for seamless integration.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Switchyard integrates into LLM pipelines via the NeMo Relay plugin, embedded Rust/Python libraries, or a stand-alone OpenAI-compatible proxy server, with all paths using a common routing-algorithm abstraction defined in [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs).**

The NVIDIA-NeMo/Switchyard repository provides intelligent routing for large language model requests, allowing systems to dynamically select between efficient and capable models. Understanding the typical workflow for using Switchyard requires choosing from three distinct integration paths that share a unified algorithmic core. Each path leverages the same **routing-algorithm** abstraction—such as `stage_router`, `llm_classifier`, or `advisor`—to decide whether to route requests to an efficient model, a capable model, or a fallback target.


## Path 1: Deploy the NeMo Relay Plugin

The first integration method involves building the native **switchyard-nemo-relay-plugin** and registering it with a running NeMo Relay deployment. This plugin intercepts every request that Relay receives and routes it through Switchyard’s algorithms before forwarding to the target LLM.

To implement this path, build the plugin crate located at `crates/switchyard-nemo-relay-plugin/` and configure it to point at a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file. The plugin’s implementation in [`crates/switchyard-nemo-relay-plugin/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/lib.rs) handles the native interface with Relay, while the routing logic itself is delegated to the shared library. This approach is ideal if you already operate a NeMo Relay infrastructure and want to add intelligent routing without modifying your client code.


## Path 2: Embed the Library in Your Application

The second path embeds Switchyard directly into your own harness using the **switchyard-libsy** crate (or its Python bindings). Your code remains in charge of making actual model calls; Switchyard only tells you *which* model to invoke and when to finalize the response. The core API surface lives in [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs), exposing the `Algorithm` trait and `Step` enum.

### Python Integration

Add the Switchyard Python bindings to your project and construct a routing algorithm. Drive the logic using `algorithm.run_stream`, handling `Step.CallModel` to invoke your own client code and `Step.Done` to return the final response.

```python
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# 1️⃣ Build the routing algorithm

algorithm = stage_router(
    "capable",      # name of the capable target

    "efficient",   # name of the efficient target

    picker="efficient_first",
    confidence_threshold=0.5,
)

# 2️⃣ Helper that calls a real model client (you provide the client dict)

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception:
            continue
    raise RuntimeError("All candidate models failed")

# 3️⃣ Drive the algorithm

async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                try:
                    call.respond(await call_with_fallback(call.request, call.models, clients))
                except Exception as err:
                    call.fail(err)
            case Step.Done(outcome):
                # If the algorithm already produced an answer, return it

                if outcome.response is not None:
                    return outcome.response
                # Otherwise make the final model call ourselves

                return await call_with_fallback(
                    outcome.request, outcome.selected_model_ids, clients
                )
    raise RuntimeError("Algorithm terminated without a decision")

```

### Rust Integration

For Rust applications, import the `switchyard-libsy` crate and use `Algorithm::run_stream` to process requests. The `StageRouterConfig` struct allows you to specify target models, picker strategies, and confidence thresholds.

```rust
use switchyard_libsy::{
    algorithms::stage::StageRouterConfig,
    libsy::{Algorithm, Step, StepStream},
    libsy::Outcome,
};

// 1️⃣ Build the algorithm
let config = StageRouterConfig {
    capable_target: "capable".into(),
    efficient_target: "efficient".into(),
    picker: switchyard_libsy::algorithms::util::stage::PickerMode::EfficientFirst,
    confidence_threshold: 0.5,
    ..Default::default()
};
let algorithm = StageRouter::new(config);

// 2️⃣ Normalized request (see `switchyard-protocol` for the shape)
let request = serde_json::json!({
    "messages": [{ "role": "user", "content": [{ "type": "text", "text": "hello" }] }],
    "model": "switchyard"
});

// 3️⃣ Drive the algorithm
let mut stream = algorithm.run_stream(request);
while let Some(step) = stream.next().await {
    match step {
        Step::CallModel(call) => {
            // Your code calls the appropriate model client and returns a normalized response
            let resp = my_client.call(call.request.clone()).await?;
            call.respond(resp);
        }
        Step::Done(outcome) => {
            // `outcome.response` may already contain the final answer
            if let Some(resp) = outcome.response {
                return Ok(resp);
            }
        }
    }
}

```

The [`crates/libsy-llm-client/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/lib.rs) file provides a convenience `run` helper that automates driving an `Algorithm` with real model calls, though embedded harnesses typically implement their own client logic as shown above.


## Path 3: Run a Stand-alone Proxy Server

The third path treats Switchyard as a drop-in OpenAI- or Anthropic-compatible proxy. Install the **switchyard-server** binary from `crates/switchyard-server/`, define your targets and routing logic in a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file, and start the server. Any compliant client can then point at the proxy endpoint (e.g., `http://localhost:4000/v1/chat/completions`).

The entry point for the server is [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs), which handles HTTP transport and request normalization. The [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs) crate translates between vendor-specific payloads (OpenAI, Anthropic) and the internal protocol used by Switchyard.

```bash

# Install the server binary

cargo install --locked switchyard-server

# Write a minimal routes.toml (see docs for full schema)

cat > routes.toml <<'TOML'
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
TOML

# Start the proxy (dry‑run first to validate)

export OPENROUTER_API_KEY="your-key"
switchyard-server --config routes.toml --dry-run
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

# Query the proxy as an OpenAI client

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"hello"}]}'

```


## Core Routing Abstraction and Protocol

All three integration paths rely on a common decision flow implemented in the core library. The **`Algorithm`** trait defines a `run_stream` method that yields **`Step`** variants: `Step::CallModel` requests that the harness invoke a classifier or judge model, while `Step::Done` signals that the algorithm has either produced a final response or determined which target model should generate it.

The normalized request and response schema used throughout the system is defined in **[`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs)**. This protocol ensures that routing decisions are based on a consistent data structure regardless of whether the request originated from an OpenAI-compatible client, a NeMo Relay instance, or a direct library call. The decision flow typically follows this pattern:

1. Ingest a normalized request.
2. Loop through `Algorithm::run_stream`, handling `Step::CallModel` by invoking the requested model(s).
3. Return the final `Outcome` when `Step::Done` is reached, either using the pre-computed response or delegating to the selected efficient/capable model.


## Summary

- **Three integration paths** are available: the NeMo Relay plugin ([`crates/switchyard-nemo-relay-plugin/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/lib.rs)), embedded library usage ([`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs)), and the stand-alone proxy server ([`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs)).
- **Unified configuration** uses a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file to define targets, LLM clients, and routing algorithms like `stage_router`, with full schema documentation available in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md).
- **Core abstraction** centers on the `Algorithm` trait and `Step` enum, allowing consistent routing logic across Rust, Python, and HTTP interfaces.
- **Protocol normalization** occurs in [`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs), ensuring that requests from OpenAI, Anthropic, or other providers are handled uniformly before routing decisions are applied.


## Frequently Asked Questions

### What configuration file format does Switchyard use?

Switchyard uses a **TOML** configuration file, conventionally named [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml), to define LLM clients, model targets, and routing rules. The schema supports defining provider-specific settings (such as base URLs and API key environment variables) under `[llm_clients]`, model identifiers under `[targets]`, and algorithm configurations (including `stage_router` parameters) under `[routes]`. Full schema details are documented in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md).

### How do I choose between the three integration paths?

Choose the **NeMo Relay plugin** (Path 1) if you already run NVIDIA's NeMo Relay and want transparent routing without client changes. Choose **embedded library** (Path 2) if you need fine-grained control over model execution, custom authentication, or integration into an existing Rust or Python application. Choose the **stand-alone proxy** (Path 3) if you need a drop-in replacement for OpenAI-compatible endpoints that works with any HTTP client.

### How does Switchyard decide which model to route a request to?

Switchyard uses **routing algorithms** such as `stage_router`, `llm_classifier`, or `advisor` defined in the configuration. These algorithms analyze the request (potentially calling intermediate "judge" models) to determine whether an efficient model is sufficient or if a more capable model is required. The decision is based on confidence thresholds and picker strategies (e.g., `efficient_first`) specified in the [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file.

### Is there a Python API available for Switchyard?

Yes, Switchyard provides Python bindings for the core library. You can import `switchyard.libsy` to construct algorithms like `stage_router` and drive them using `algorithm.run_stream`. The Python API mirrors the Rust API, yielding `Step` objects that your code handles to make actual model calls and return normalized responses.