Switchyard Route Types and Their Basic Configurations

Switchyard supports seven distinct route types—llm_classifier, stage_router, random, passthrough, noop, advisor, and subagent—each defined by a type field in a TOML configuration file and mapped to a specific RouteConfig enum variant in the Rust source.

Switchyard is an open-source LLM routing engine developed by NVIDIA that directs incoming requests to appropriate model backends based on configurable strategies. Understanding the available Switchyard route types and their basic configurations allows you to optimize for cost, latency, and accuracy by selecting the right routing algorithm for your use case.

Built-in Route Types in Switchyard

Switchyard implements routing logic through a structured TOML schema where each route entry specifies a type field. The Rust parser in crates/switchyard-server/src/config.rs maps these strings to the RouteConfig enum variants, determining how requests are processed at runtime.

LLM Classifier Routes

LLM Classifier routes analyze request content to determine whether a turn should be handled by a weak or strong model tier. This is ideal when you want a lightweight model to screen requests and only escalate complex queries to expensive frontier models.

Basic configuration requires the type and model fields, with an optional classifier specification:

[routes.classifier]
type = "llm_classifier"
model = "gpt-4o-mini"       # weak tier target

classifier = "gpt-4o-mini"  # model that performs classification

According to the source code, the classifier evaluates the incoming prompt and routes to the appropriate tier without requiring manual traffic splitting.

Stage Router Routes

Stage Router enables signal-driven routing without consuming additional model tokens. Instead of calling a classifier LLM, it examines existing conversation signals—such as tool results, errors, or metadata—to determine the appropriate backend.

Configuration requires picker_mode and confidence_threshold:

[routes.stage]
type = "stage_router"
picker_mode = "efficient_first"
confidence_threshold = 0.5  # route to strong tier if confidence ≥ 0.5

As implemented in crates/switchyard-server/src/config.rs, this route type is more efficient than classifier-based approaches because it leverages already-available conversation state rather than generating new inference calls.

Random Routes

Random routes implement traffic splitting for A/B testing, cost experiments, or baseline comparisons. The router assigns requests to backends based on fixed probability distributions defined in the configuration.

You must specify a distribution mapping model IDs to percentage weights:

[routes.random]
type = "random"
distribution = { "gpt-4o-mini" = 0.5, "gpt-4o" = 0.5 }

The distribution values must sum to 1.0 (100%), and Switchyard samples from this distribution for each incoming request.

Passthrough Routes

Passthrough provides the simplest routing strategy: forwarding every request to a single specified backend without any decision logic. This is useful when you need a direct proxy to a specific model or when integrating Switchyard into existing pipelines that handle routing upstream.

Configuration requires only the target field:

[routes.passthrough]
type = "passthrough"
target = "gpt-4o-mini"

The target value must correspond to a valid model ID configured in your Switchyard server instance.

No-Op Routes

No-Op routes serve as placeholders that return canned responses without invoking any backend model. This is invaluable for integration testing, load testing the routing layer itself, or simulating downstream failures.

Configuration is minimal:

[routes.noop]
type = "noop"

When selected, this route immediately returns a static response defined in the Switchyard server configuration, consuming no inference budget.

Advisor Routes

Advisor routes implement a gating pattern where a secondary "advisor" model vets or approves requests before they reach the primary target. This allows for safety filtering, quality gates, or multi-stage approval workflows.

Configuration specifies the advisor model and optional gating policy:

[routes.advisor]
type = "advisor"
advisor = "gpt-4o-mini"
policy = { selector = "/route" }

The advisor evaluates the request against the defined policy and either allows it to proceed to the primary target or rejects it with a predefined response.

Subagent Routes

Subagent routes delegate request processing to custom Python or Rust implementations via the SubagentRouter interface. This extends Switchyard's capabilities beyond built-in algorithms, allowing domain-specific routing logic.

Basic configuration points to a custom module:


# Conceptual example - points to custom implementation

type = "subagent"

As noted in the routing documentation at docs/routing_algorithms/, this type requires implementing the subagent interface in your chosen language and registering it with the Switchyard runtime.

Complete TOML Configuration Example

A production routes.toml file typically defines multiple route strategies that clients can select via the router field in HTTP requests. Below is a comprehensive example demonstrating all supported route types:


# 1️⃣ Passthrough – simple one-target route

[routes.passthrough]
type = "passthrough"
target = "gpt-4o-mini"

# 2️⃣ Random – split traffic between two models (50% each)

[routes.random]
type = "random"
distribution = { "gpt-4o-mini" = 0.5, "gpt-4o" = 0.5 }

# 3️⃣ LLM Classifier – let a classifier decide weak vs. strong tier

[routes.classifier]
type = "llm_classifier"
model = "gpt-4o-mini"
classifier = "gpt-4o-mini"

# 4️⃣ Stage Router – signal-driven routing without extra model calls

[routes.stage]
type = "stage_router"
picker_mode = "efficient_first"
confidence_threshold = 0.5

# 5️⃣ No-Op – returns a canned response

[routes.noop]
type = "noop"

# 6️⃣ Advisor – gate requests through an advisory model

[routes.advisor]
type = "advisor"
advisor = "gpt-4o-mini"
policy = { selector = "/route" }

Validate your configuration without starting the full server using the dry-run flag:

switchyard-server --config routes.toml --dry-run

Route Selection Runtime Behavior

When Switchyard receives a request, the client specifies which router to use through the HTTP payload (e.g., "router": "stage"). The server matches this identifier against the table keys in routes.toml (e.g., [routes.stage]) and instantiates the corresponding RouteConfig variant defined in crates/switchyard-server/src/config.rs.

The TOML schema reference at docs/reference/toml_schema.md provides the complete specification of required and optional fields for each route type, while detailed algorithm documentation resides in docs/routing_algorithms/ (including llm_classifier_routing.md, stage_router_routing.md, and random_routing.md).

Summary

  • LLM Classifier routes use a classification model to select between weak and strong model tiers based on request content.
  • Stage Router examines existing conversation signals to route efficiently without additional inference costs.
  • Random routes enable A/B testing through configurable probability distributions across multiple backends.
  • Passthrough provides direct forwarding to a single target without routing logic.
  • No-Op routes return static responses for testing and development workflows.
  • Advisor routes implement two-stage processing with a gating model that vets requests before primary processing.
  • Subagent routes delegate to custom implementations for specialized routing logic beyond built-in strategies.

Frequently Asked Questions

What is the difference between LLM Classifier and Stage Router routes?

LLM Classifier routes invoke a language model to analyze request content and determine routing, consuming tokens for the classification step. Stage Router routes instead examine existing signals in the conversation—such as previous tool outputs or error states—to make routing decisions without additional model calls, making them more token-efficient when conversation metadata is already available.

How can I test Switchyard route configurations without invoking actual model backends?

Use the No-Op route type to return canned responses without hitting backends, or start the server with the --dry-run flag (e.g., switchyard-server --config routes.toml --dry-run) to validate TOML syntax and route definitions. The dev-server/config.toml file in the repository provides a ready-to-run example containing all route types for local testing.

Where is the formal schema for route configuration documented?

The complete TOML schema specification resides in docs/reference/toml_schema.md, which lists all required and optional fields for each route type. Additionally, the Rust struct definitions in crates/switchyard-server/src/config.rs define the exact RouteConfig enum variants that enforce schema validation at compile time.

When should I use a Random route versus an LLM Classifier route?

Use Random routes when you need statistical A/B testing, cost baselining, or gradual rollouts where traffic splits are fixed percentages independent of request content. Use LLM Classifier routes when request complexity varies significantly and you want intelligent routing that sends simple queries to cheap models and complex queries to powerful models based on content analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →