# How to Use Switchyard for Custom Model Development: A Complete Guide

> Learn custom model development with Switchyard. This guide shows how Switchyard routes LLM requests to targets without client code changes, enabling flexible and efficient development.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Switchyard enables custom model development by acting as a transparent routing layer that intercepts LLM requests and delegates them to specific targets based on configurable algorithms, requiring zero changes to existing client code.**

Switchyard is an open-source routing framework from NVIDIA-NeMo that determines which model (or *target*) should handle every LLM request. By sitting between the caller and model providers, it allows developers to implement sophisticated routing logic—such as cascading from efficient to capable models or using LLM-based judges—without modifying downstream applications. This guide covers how to configure, extend, and deploy Switchyard for your custom model development workflows.

## Understanding Switchyard's Routing Architecture

The core routing engine lives in the **`switchyard-libsy`** crate. According to the source code, the entry point for any routing decision is the `Algorithm::run_stream` method, which yields a stream of `Step` enums including `Step::CallModel` and `Step::Done`【crates/libsy/README.md】.

Before routing occurs, every request is normalized into a **protocol-neutral** `Request` shape defined in the `switchyard-protocol` crate【crates/protocol/README.md】. This abstraction allows Switchyard to route traffic from OpenAI, Anthropic, or other providers interchangeably.

The routing decision itself is made by a *routing algorithm*—such as the stage router, LLM classifier, or escalation policies—which selects a target name from your configuration.

## Step 1: Expose Your Models as Switchyard Targets

To integrate custom models, you must declare them as **targets** in a TOML configuration file (conventionally named [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml)). Each target requires:

- An `id` field representing the model identifier understood by the provider
- A reference to an LLM client defined in the same file

Switchyard reads this configuration at start-up to establish available endpoints【docs/reference/toml_schema.md】.

The TOML structure distinguishes between three concepts:

1. **`[llm_clients]`** – Connection parameters for providers (base URLs, API keys)
2. **`[targets]`** – Model identifiers mapped to specific clients
3. **`[routes]`** – Routing rules that select among targets based on the active algorithm

## Step 2: Choose a Routing Algorithm

Switchyard provides multiple algorithms for selecting targets, each suited to different latency, cost, and quality requirements.

### Stage Router

The **stage router** (default for benchmarks) routes requests based on tool results and optional classifier calls. It implements an efficient-first strategy, attempting cheaper or faster models before escalating to more capable ones only when necessary【crates/libsy/src/algorithms/stage_router.rs】.

### LLM Classifier with Custom Multi-Target Routing

The **LLM classifier** algorithm uses a judge model to predict whether a request can be solved by a weaker (cheaper) model or requires escalation. Switchyard supports three built-in modes—*capability*, *escalation*, and *custom*—allowing you to route among any number of targets based on the judge's verdict【docs/routing_algorithms/llm_classifier_routing.md#custom-multi-target-routing】.

In **custom** mode, you define a JSON Schema that constrains the judge's output, plus a policy that extracts the target selection from that structured response.

## Step 3: Configure Custom Classifier Policies (Optional)

When the built-in binary routing (weak vs. strong) is insufficient, custom classifier mode enables routing across four or more targets. You must supply:

- A `prompt` instructing the judge how to select among targets
- A `response_schema` defining valid JSON output
- A `[routes.your_route.policy]` section specifying how to parse the verdict

The policy uses JSONPath-like selectors (e.g., `/decision/target`) to extract the target name from the judge's structured output【docs/routing_algorithms/llm_classifier_routing.md】.

## Integration Paths for Custom Development

Switchyard supports three deployment patterns that share the same TOML schema, allowing seamless migration from prototyping to production.

### Path 1: NeMo Relay Plugin

Load the Switchyard plugin into an existing NeMo Relay deployment by pointing the relay at your [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file. This path is ideal if you already use NVIDIA's inference serving infrastructure【crates/switchyard-nemo-relay-plugin/README.md】.

### Path 2: Embed the Library (Python/Rust)

For custom harnesses, instantiate an algorithm directly in Python or Rust and drive it with `run_stream`. This path provides maximum flexibility for integrating with existing model clients or evaluation frameworks【README.md#path-2-embed-the-library】.

The following Python example demonstrates embedding the stage router:

```python
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# 1️⃣ Build the algorithm (stage router picks “efficient” first)

algorithm = stage_router(
    capable="capable",      # name of the strong target in routes.toml

    efficient="efficient",  # name of the weak target

    picker="efficient_first",
    confidence_threshold=0.5,
)

# 2️⃣ Helper that calls your existing model clients

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception:
            continue
    raise RuntimeError("All candidates failed")

# 3️⃣ Drive the algorithm

async def route(request: dict, clients: dict) -> LlmResponse.Agg:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                # Forward the call to the appropriate client(s)

                try:
                    call.respond(
                        await call_with_fallback(call.request, call.models, clients)
                    )
                except Exception as err:
                    call.fail(err)

            case Step.Done(outcome):
                # If the algorithm already produced an answer, return it

                if outcome.response is not None:
                    return outcome.response
                # Otherwise make the final model call

                return await call_with_fallback(
                    outcome.request, outcome.selected_model_ids, clients
                )
    raise RuntimeError("Algorithm terminated without a decision")

```

### Path 3: Standalone Proxy Server

Run the `switchyard-server` binary to expose OpenAI/Anthropic-compatible endpoints. Clients address the router as if it were a model name, enabling drop-in replacement for existing API calls【crates/switchyard-server/README.md】.

## Complete Configuration Example

The following [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configures four distinct targets (fast, balanced, reasoning, premium) and a custom LLM classifier route that selects among them:

```toml
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.classifier]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[targets.fast]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.balanced]
id = "anthropic/claude-3-opus-1.0"
llm_client = "openrouter"

[targets.reasoning]
id = "google/gemini-pro"
llm_client = "openrouter"

[targets.premium]
id = "deepseek/deepseek-v2"
llm_client = "openrouter"

[routes.my_route]
id = "my_route"
type = "llm_classifier"
mode = "custom"
classifier_target = "classifier"
targets = ["fast", "balanced", "reasoning", "premium"]
default_target = "premium"
prompt = """
Choose the best configured target for this request.
Return JSON matching the response schema supplied with the request.
"""
response_schema = '''
{
  "type":"object",
  "properties":{
    "decision":{
      "type":"object",
      "properties":{"target":{"type":"string","enum":["fast","balanced","reasoning","premium"]}},
      "required":["target"],
      "additionalProperties":false
    }
  },
  "required":["decision"],
  "additionalProperties":false
}
'''

[routes.my_route.policy]
type = "target_selector"
selector = "/decision/target"

```

## Running the Standalone Server

To deploy the configuration above as a standalone proxy:

```bash

# 1️⃣ Install the server (requires Rust)

cargo install --locked switchyard-server

# 2️⃣ Start the proxy with the TOML configuration

switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

Any OpenAI-compatible client can now send requests to `http://localhost:4000/v1/chat/completions` with the model name `"my_route"`. Switchyard will invoke the custom classifier, select one of the four configured targets, and forward the request to the chosen provider.

## Summary

- **Switchyard** acts as a transparent routing layer between clients and LLM providers, enabling dynamic target selection without client code changes.
- Configuration centers on [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml), which defines **targets** (models), **clients** (provider connections), and **routes** (selection algorithms).
- Three integration paths—**NeMo Relay plugin**, **embedded library**, and **standalone proxy**—share the same configuration schema for flexible deployment.
- The **LLM classifier** algorithm supports custom multi-target routing via JSON Schema constraints and extraction policies.
- Core implementation resides in the `switchyard-libsy` crate, with `Algorithm::run_stream` serving as the primary entry point for programmatic usage.

## Frequently Asked Questions

### How does Switchyard normalize requests from different providers?

Switchyard converts incoming requests from OpenAI, Anthropic, or other formats into a **protocol-neutral** `Request` shape defined in the `switchyard-protocol` crate. This normalization allows routing algorithms to operate agnostically of the original provider format【crates/protocol/README.md】.

### Can I use Switchyard with providers not explicitly supported by NVIDIA?

Yes. By defining custom entries in the `[llm_clients]` section of your [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml), you can route to any provider that exposes an OpenAI-compatible chat completions endpoint. The `format` and `base_url` fields allow you to specify the exact API contract and endpoint address【docs/reference/toml_schema.md】.

### What happens if the classifier or routing algorithm fails?

Switchyard provides a `default_target` configuration option in each route definition. If the classifier fails to produce a valid verdict, or if all preferred targets are unavailable, the system falls back to this default target to ensure request completion【docs/routing_algorithms/llm_classifier_routing.md】.

### Is it possible to implement entirely custom routing logic beyond the provided algorithms?

Yes. By embedding `switchyard-libsy` directly in a Rust or Python application, you can implement the `Algorithm` trait (or wrap the library) to define completely bespoke routing logic. The `run_stream` method yields `Step::CallModel` and `Step::Done` events that your code handles, allowing you to intercept and modify routing decisions at every step【crates/libsy/README.md】.