How to Use Switchyard for Custom Model Development: A Complete Guide
Switchyard enables custom model development by acting as a transparent routing layer that intercepts LLM requests and delegates them to specific targets based on configurable algorithms, requiring zero changes to existing client code.
Switchyard is an open-source routing framework from NVIDIA-NeMo that determines which model (or target) should handle every LLM request. By sitting between the caller and model providers, it allows developers to implement sophisticated routing logic—such as cascading from efficient to capable models or using LLM-based judges—without modifying downstream applications. This guide covers how to configure, extend, and deploy Switchyard for your custom model development workflows.
Understanding Switchyard's Routing Architecture
The core routing engine lives in the switchyard-libsy crate. According to the source code, the entry point for any routing decision is the Algorithm::run_stream method, which yields a stream of Step enums including Step::CallModel and Step::Done【crates/libsy/README.md】.
Before routing occurs, every request is normalized into a protocol-neutral Request shape defined in the switchyard-protocol crate【crates/protocol/README.md】. This abstraction allows Switchyard to route traffic from OpenAI, Anthropic, or other providers interchangeably.
The routing decision itself is made by a routing algorithm—such as the stage router, LLM classifier, or escalation policies—which selects a target name from your configuration.
Step 1: Expose Your Models as Switchyard Targets
To integrate custom models, you must declare them as targets in a TOML configuration file (conventionally named routes.toml). Each target requires:
- An
idfield representing the model identifier understood by the provider - A reference to an LLM client defined in the same file
Switchyard reads this configuration at start-up to establish available endpoints【docs/reference/toml_schema.md】.
The TOML structure distinguishes between three concepts:
[llm_clients]– Connection parameters for providers (base URLs, API keys)[targets]– Model identifiers mapped to specific clients[routes]– Routing rules that select among targets based on the active algorithm
Step 2: Choose a Routing Algorithm
Switchyard provides multiple algorithms for selecting targets, each suited to different latency, cost, and quality requirements.
Stage Router
The stage router (default for benchmarks) routes requests based on tool results and optional classifier calls. It implements an efficient-first strategy, attempting cheaper or faster models before escalating to more capable ones only when necessary【crates/libsy/src/algorithms/stage_router.rs】.
LLM Classifier with Custom Multi-Target Routing
The LLM classifier algorithm uses a judge model to predict whether a request can be solved by a weaker (cheaper) model or requires escalation. Switchyard supports three built-in modes—capability, escalation, and custom—allowing you to route among any number of targets based on the judge's verdict【docs/routing_algorithms/llm_classifier_routing.md#custom-multi-target-routing】.
In custom mode, you define a JSON Schema that constrains the judge's output, plus a policy that extracts the target selection from that structured response.
Step 3: Configure Custom Classifier Policies (Optional)
When the built-in binary routing (weak vs. strong) is insufficient, custom classifier mode enables routing across four or more targets. You must supply:
- A
promptinstructing the judge how to select among targets - A
response_schemadefining valid JSON output - A
[routes.your_route.policy]section specifying how to parse the verdict
The policy uses JSONPath-like selectors (e.g., /decision/target) to extract the target name from the judge's structured output【docs/routing_algorithms/llm_classifier_routing.md】.
Integration Paths for Custom Development
Switchyard supports three deployment patterns that share the same TOML schema, allowing seamless migration from prototyping to production.
Path 1: NeMo Relay Plugin
Load the Switchyard plugin into an existing NeMo Relay deployment by pointing the relay at your routes.toml file. This path is ideal if you already use NVIDIA's inference serving infrastructure【crates/switchyard-nemo-relay-plugin/README.md】.
Path 2: Embed the Library (Python/Rust)
For custom harnesses, instantiate an algorithm directly in Python or Rust and drive it with run_stream. This path provides maximum flexibility for integrating with existing model clients or evaluation frameworks【README.md#path-2-embed-the-library】.
The following Python example demonstrates embedding the stage router:
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router
# 1️⃣ Build the algorithm (stage router picks “efficient” first)
algorithm = stage_router(
capable="capable", # name of the strong target in routes.toml
efficient="efficient", # name of the weak target
picker="efficient_first",
confidence_threshold=0.5,
)
# 2️⃣ Helper that calls your existing model clients
async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
for model in models:
try:
return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
except Exception:
continue
raise RuntimeError("All candidates failed")
# 3️⃣ Drive the algorithm
async def route(request: dict, clients: dict) -> LlmResponse.Agg:
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
# Forward the call to the appropriate client(s)
try:
call.respond(
await call_with_fallback(call.request, call.models, clients)
)
except Exception as err:
call.fail(err)
case Step.Done(outcome):
# If the algorithm already produced an answer, return it
if outcome.response is not None:
return outcome.response
# Otherwise make the final model call
return await call_with_fallback(
outcome.request, outcome.selected_model_ids, clients
)
raise RuntimeError("Algorithm terminated without a decision")
Path 3: Standalone Proxy Server
Run the switchyard-server binary to expose OpenAI/Anthropic-compatible endpoints. Clients address the router as if it were a model name, enabling drop-in replacement for existing API calls【crates/switchyard-server/README.md】.
Complete Configuration Example
The following routes.toml configures four distinct targets (fast, balanced, reasoning, premium) and a custom LLM classifier route that selects among them:
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.classifier]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
[targets.fast]
id = "openai/gpt-4o"
llm_client = "openrouter"
[targets.balanced]
id = "anthropic/claude-3-opus-1.0"
llm_client = "openrouter"
[targets.reasoning]
id = "google/gemini-pro"
llm_client = "openrouter"
[targets.premium]
id = "deepseek/deepseek-v2"
llm_client = "openrouter"
[routes.my_route]
id = "my_route"
type = "llm_classifier"
mode = "custom"
classifier_target = "classifier"
targets = ["fast", "balanced", "reasoning", "premium"]
default_target = "premium"
prompt = """
Choose the best configured target for this request.
Return JSON matching the response schema supplied with the request.
"""
response_schema = '''
{
"type":"object",
"properties":{
"decision":{
"type":"object",
"properties":{"target":{"type":"string","enum":["fast","balanced","reasoning","premium"]}},
"required":["target"],
"additionalProperties":false
}
},
"required":["decision"],
"additionalProperties":false
}
'''
[routes.my_route.policy]
type = "target_selector"
selector = "/decision/target"
Running the Standalone Server
To deploy the configuration above as a standalone proxy:
# 1️⃣ Install the server (requires Rust)
cargo install --locked switchyard-server
# 2️⃣ Start the proxy with the TOML configuration
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
Any OpenAI-compatible client can now send requests to http://localhost:4000/v1/chat/completions with the model name "my_route". Switchyard will invoke the custom classifier, select one of the four configured targets, and forward the request to the chosen provider.
Summary
- Switchyard acts as a transparent routing layer between clients and LLM providers, enabling dynamic target selection without client code changes.
- Configuration centers on
routes.toml, which defines targets (models), clients (provider connections), and routes (selection algorithms). - Three integration paths—NeMo Relay plugin, embedded library, and standalone proxy—share the same configuration schema for flexible deployment.
- The LLM classifier algorithm supports custom multi-target routing via JSON Schema constraints and extraction policies.
- Core implementation resides in the
switchyard-libsycrate, withAlgorithm::run_streamserving as the primary entry point for programmatic usage.
Frequently Asked Questions
How does Switchyard normalize requests from different providers?
Switchyard converts incoming requests from OpenAI, Anthropic, or other formats into a protocol-neutral Request shape defined in the switchyard-protocol crate. This normalization allows routing algorithms to operate agnostically of the original provider format【crates/protocol/README.md】.
Can I use Switchyard with providers not explicitly supported by NVIDIA?
Yes. By defining custom entries in the [llm_clients] section of your routes.toml, you can route to any provider that exposes an OpenAI-compatible chat completions endpoint. The format and base_url fields allow you to specify the exact API contract and endpoint address【docs/reference/toml_schema.md】.
What happens if the classifier or routing algorithm fails?
Switchyard provides a default_target configuration option in each route definition. If the classifier fails to produce a valid verdict, or if all preferred targets are unavailable, the system falls back to this default target to ensure request completion【docs/routing_algorithms/llm_classifier_routing.md】.
Is it possible to implement entirely custom routing logic beyond the provided algorithms?
Yes. By embedding switchyard-libsy directly in a Rust or Python application, you can implement the Algorithm trait (or wrap the library) to define completely bespoke routing logic. The run_stream method yields Step::CallModel and Step::Done events that your code handles, allowing you to intercept and modify routing decisions at every step【crates/libsy/README.md】.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →