How Switchyard Optimizes LLM Costs Using Tiered Models
Switchyard reduces LLM inference costs by dynamically routing requests between a capable (high-quality, expensive) tier and an efficient (lower-cost) tier, using the stage-router algorithm to make per-turn escalation decisions based on tool-result signals and confidence thresholds.
NVIDIA NeMo Switchyard tackles the high cost of large language model (LLM) inference through intelligent request routing. By analyzing conversation context in real-time, Switchyard automatically selects the most cost-effective model tier capable of handling each specific turn, ensuring premium model costs are incurred only when justified by the complexity of the task.
Understanding the Two-Tier Architecture
Switchyard implements a dual-tier routing system defined by two distinct model targets in your configuration. The capable tier represents your high-quality, higher-priced model (such as GPT-4o), while the efficient tier uses a lower-cost alternative (such as GPT-4o-mini).
In the TOML schema, these tiers are explicitly declared using the capable_target and efficient_target fields within a stage-router route definition. These correspond to target definitions elsewhere in your configuration file that specify model IDs, API endpoints, and authentication details. This architecture allows Switchyard to maintain a single conversation while potentially switching between models of different capabilities and costs on a per-turn basis.
How the Stage-Router Algorithm Makes Routing Decisions
The core optimization logic resides in the stage-router algorithm, implemented in crates/libsy/src/algorithms and exposed through switchyard/libsy/algorithms.py. This algorithm evaluates each conversation turn independently to determine whether the current tier is sufficient or if escalation to the capable tier is warranted.
Picker Strategies: Starting Points for Cost Control
The routing behavior begins with a picker setting that establishes the default tier for each new turn:
- efficient_first: This production-ready default starts every turn on the efficient (cheaper) model. The algorithm only escalates to the capable tier when accumulated signals exceed the
confidence_threshold. This approach maximizes cost savings by defaulting to the lowest-cost option. - capable_first: This experimental strategy starts on the capable tier and drops to the efficient tier only after receiving strong signals that the cheaper model can handle subsequent turns. This prioritizes quality over initial cost savings.
You configure the picker in your route definition's TOML configuration, alongside the threshold and target specifications.
Signal-Based Scoring Mechanism
For each turn, Switchyard calculates routing confidence using two opposing signal axes:
WRONG signals (push toward capable tier):
severity– Error severity indicators from tool executionsspinning– Detected repetitive or non-productive generation patternsexploring– High uncertainty or exploratory behavior indicators
PROGRESS signals (push toward efficient tier):
recent_production_intensity– Measures of successful code generation or task completion
These raw scores are combined and processed through a tanh squashing function to normalize the confidence value between -1 and 1. The algorithm then compares this confidence score against your configured confidence_threshold (recommended default: 0.5). When confidence exceeds the threshold, the router switches to the capable tier; otherwise, it remains on the efficient tier, avoiding unnecessary high-cost inference calls.
Optional LLM Classifier Fallback
For borderline cases where confidence scores fall near the threshold, Switchyard supports an optional LLM classifier consultation. When enabled, this additional model call provides fine-grained analysis before making the final tier decision, allowing operators to trade a small additional latency cost for improved routing accuracy in ambiguous scenarios.
Observability and Cost Tracking
Switchyard exposes detailed Prometheus metrics that include a tier label, enabling precise tracking of costs, token usage, and latency split between your capable and efficient models. You can inspect these metrics through the server's metrics endpoint to quantify actual cost savings.
Additionally, every response includes the header x-model-router-selected-model, allowing client applications to log which specific tier handled each request for billing reconciliation and performance analysis.
Configuring Tiered Routing in Practice
To implement cost optimization, define your tiered route in a TOML configuration file:
[schema]
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"
[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3
Start the server with your configuration:
switchyard-server --config routes.toml --port 4000
Use the Python client to leverage automatic routing:
import switchyard as sy
client = sy.SwitchyardClient(
base_url="http://localhost:4000",
api_key=None,
)
resp = client.chat_completion(
model="switchyard/stage",
messages=[{"role": "user", "content": "Refactor the foo function."}]
)
print(resp["choices"][0]["message"]["content"])
print("Routed model:", resp["headers"]["x-model-router-selected-model"])
Monitor cost metrics via the Prometheus endpoint:
curl http://localhost:4000/metrics | grep switchyard_route_cost
Core Implementation Files
The tiered routing system spans several critical components in the Switchyard repository:
switchyard/libsy/algorithms.py– Exports the Python-level factory functions that instantiate the stage-router algorithm for client use.crates/libsy/src/algorithms– Contains the Rust implementation of the signal scoring logic, tier decision engine, and optional classifier integration.crates/switchyard-server/src/config.rs– Parses TOML route definitions and wires the stage-router configuration into the server runtime.docs/routing_algorithms/stage_router_routing.md– Comprehensive documentation covering picker options, signal definitions, and threshold tuning strategies.tests/test_libsy_minimal_bindings.py– Unit tests verifying the stage-router's tier selection behavior across different confidence scenarios.
Summary
- Switchyard optimizes LLM costs using a two-tier architecture with capable (expensive) and efficient (cheap) model targets configured via
capable_targetandefficient_target. - The stage-router algorithm makes per-turn routing decisions using signal-based scoring (WRONG vs. PROGRESS axes) and tanh-normalized confidence thresholds.
- Picker strategies (
efficient_firstvs.capable_first) determine whether conversations start cheap and escalate, or start premium and downgrade. - An optional LLM classifier provides additional routing precision for marginal confidence cases.
- Prometheus metrics with
tierlabels enable precise cost tracking and ROI measurement of the tiered approach.
Frequently Asked Questions
How does Switchyard decide which model tier to use?
Switchyard evaluates tool-result signals for each conversation turn, calculating a confidence score based on factors like error severity, code generation activity, and conversational progress. If the confidence exceeds the configured confidence_threshold (default 0.5), the request routes to the capable tier; otherwise, it remains on the efficient tier. This decision occurs in the Rust implementation within crates/libsy/src/algorithms and is exposed to Python via switchyard/libsy/algorithms.py.
What is the difference between efficient_first and capable_first pickers?
The efficient_first picker starts every turn on the lower-cost efficient model and escalates to the capable tier only when signals warrant it, making it the production-ready default for cost optimization. The capable_first picker takes the opposite approach, starting on the capable tier and downgrading to the efficient tier only when strong progress signals indicate the cheaper model can handle the workload. The former minimizes costs while the latter prioritizes initial quality.
Can I use an LLM to improve routing decisions?
Yes, Switchyard supports an optional LLM classifier fallback that engages when confidence scores fall near the threshold boundary. This classifier makes a fine-grained assessment of whether the current turn requires the capable tier before the final routing decision, trading a small additional inference cost for improved routing accuracy in ambiguous situations.
How do I monitor cost savings from tiered routing?
Switchyard emits Prometheus metrics including the switchyard_route_cost metric with a tier label, allowing you to aggregate expenses by capable and efficient tiers separately. You can also inspect the x-model-router-selected-model header in each response to log which tier handled specific requests. These observability features, documented in docs/internal/metrics_reference.md, enable precise calculation of infrastructure savings versus quality trade-offs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →