# How Switchyard Optimizes LLM Costs Using Tiered Models

> Reduce LLM costs with Switchyard. Dynamically route requests between tiers using tiered models and the stage-router algorithm for efficient, cost-effective inference.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-21

---

**Switchyard reduces LLM inference costs by dynamically routing requests between a capable (high-quality, expensive) tier and an efficient (lower-cost) tier, using the stage-router algorithm to make per-turn escalation decisions based on tool-result signals and confidence thresholds.**

NVIDIA NeMo Switchyard tackles the high cost of large language model (LLM) inference through intelligent request routing. By analyzing conversation context in real-time, Switchyard automatically selects the most cost-effective model tier capable of handling each specific turn, ensuring premium model costs are incurred only when justified by the complexity of the task.

## Understanding the Two-Tier Architecture

Switchyard implements a **dual-tier routing system** defined by two distinct model targets in your configuration. The **capable tier** represents your high-quality, higher-priced model (such as GPT-4o), while the **efficient tier** uses a lower-cost alternative (such as GPT-4o-mini).

In the TOML schema, these tiers are explicitly declared using the `capable_target` and `efficient_target` fields within a stage-router route definition. These correspond to target definitions elsewhere in your configuration file that specify model IDs, API endpoints, and authentication details. This architecture allows Switchyard to maintain a single conversation while potentially switching between models of different capabilities and costs on a per-turn basis.

## How the Stage-Router Algorithm Makes Routing Decisions

The core optimization logic resides in the **stage-router algorithm**, implemented in `crates/libsy/src/algorithms` and exposed through [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py). This algorithm evaluates each conversation turn independently to determine whether the current tier is sufficient or if escalation to the capable tier is warranted.

### Picker Strategies: Starting Points for Cost Control

The routing behavior begins with a **picker** setting that establishes the default tier for each new turn:

- **efficient_first**: This production-ready default starts every turn on the efficient (cheaper) model. The algorithm only escalates to the capable tier when accumulated signals exceed the `confidence_threshold`. This approach maximizes cost savings by defaulting to the lowest-cost option.
- **capable_first**: This experimental strategy starts on the capable tier and drops to the efficient tier only after receiving strong signals that the cheaper model can handle subsequent turns. This prioritizes quality over initial cost savings.

You configure the picker in your route definition's TOML configuration, alongside the threshold and target specifications.

### Signal-Based Scoring Mechanism

For each turn, Switchyard calculates routing confidence using two opposing signal axes:

**WRONG signals (push toward capable tier):**
- `severity` – Error severity indicators from tool executions
- `spinning` – Detected repetitive or non-productive generation patterns
- `exploring` – High uncertainty or exploratory behavior indicators

**PROGRESS signals (push toward efficient tier):**
- `recent_production_intensity` – Measures of successful code generation or task completion

These raw scores are combined and processed through a **tanh squashing function** to normalize the confidence value between -1 and 1. The algorithm then compares this confidence score against your configured `confidence_threshold` (recommended default: `0.5`). When confidence exceeds the threshold, the router switches to the capable tier; otherwise, it remains on the efficient tier, avoiding unnecessary high-cost inference calls.

### Optional LLM Classifier Fallback

For borderline cases where confidence scores fall near the threshold, Switchyard supports an **optional LLM classifier** consultation. When enabled, this additional model call provides fine-grained analysis before making the final tier decision, allowing operators to trade a small additional latency cost for improved routing accuracy in ambiguous scenarios.

## Observability and Cost Tracking

Switchyard exposes detailed **Prometheus metrics** that include a `tier` label, enabling precise tracking of costs, token usage, and latency split between your capable and efficient models. You can inspect these metrics through the server's metrics endpoint to quantify actual cost savings.

Additionally, every response includes the header `x-model-router-selected-model`, allowing client applications to log which specific tier handled each request for billing reconciliation and performance analysis.

## Configuring Tiered Routing in Practice

To implement cost optimization, define your tiered route in a TOML configuration file:

```toml
[schema]
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3

```

Start the server with your configuration:

```bash
switchyard-server --config routes.toml --port 4000

```

Use the Python client to leverage automatic routing:

```python
import switchyard as sy

client = sy.SwitchyardClient(
    base_url="http://localhost:4000",
    api_key=None,
)

resp = client.chat_completion(
    model="switchyard/stage",
    messages=[{"role": "user", "content": "Refactor the foo function."}]
)

print(resp["choices"][0]["message"]["content"])
print("Routed model:", resp["headers"]["x-model-router-selected-model"])

```

Monitor cost metrics via the Prometheus endpoint:

```bash
curl http://localhost:4000/metrics | grep switchyard_route_cost

```

## Core Implementation Files

The tiered routing system spans several critical components in the Switchyard repository:

- **[`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)** – Exports the Python-level factory functions that instantiate the stage-router algorithm for client use.
- **`crates/libsy/src/algorithms`** – Contains the Rust implementation of the signal scoring logic, tier decision engine, and optional classifier integration.
- **[`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs)** – Parses TOML route definitions and wires the stage-router configuration into the server runtime.
- **[`docs/routing_algorithms/stage_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/stage_router_routing.md)** – Comprehensive documentation covering picker options, signal definitions, and threshold tuning strategies.
- **[`tests/test_libsy_minimal_bindings.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_libsy_minimal_bindings.py)** – Unit tests verifying the stage-router's tier selection behavior across different confidence scenarios.

## Summary

- Switchyard optimizes LLM costs using a **two-tier architecture** with capable (expensive) and efficient (cheap) model targets configured via `capable_target` and `efficient_target`.
- The **stage-router algorithm** makes per-turn routing decisions using signal-based scoring (WRONG vs. PROGRESS axes) and tanh-normalized confidence thresholds.
- **Picker strategies** (`efficient_first` vs. `capable_first`) determine whether conversations start cheap and escalate, or start premium and downgrade.
- An **optional LLM classifier** provides additional routing precision for marginal confidence cases.
- **Prometheus metrics** with `tier` labels enable precise cost tracking and ROI measurement of the tiered approach.

## Frequently Asked Questions

### How does Switchyard decide which model tier to use?

Switchyard evaluates tool-result signals for each conversation turn, calculating a confidence score based on factors like error severity, code generation activity, and conversational progress. If the confidence exceeds the configured `confidence_threshold` (default 0.5), the request routes to the capable tier; otherwise, it remains on the efficient tier. This decision occurs in the Rust implementation within `crates/libsy/src/algorithms` and is exposed to Python via [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py).

### What is the difference between efficient_first and capable_first pickers?

The `efficient_first` picker starts every turn on the lower-cost efficient model and escalates to the capable tier only when signals warrant it, making it the production-ready default for cost optimization. The `capable_first` picker takes the opposite approach, starting on the capable tier and downgrading to the efficient tier only when strong progress signals indicate the cheaper model can handle the workload. The former minimizes costs while the latter prioritizes initial quality.

### Can I use an LLM to improve routing decisions?

Yes, Switchyard supports an **optional LLM classifier fallback** that engages when confidence scores fall near the threshold boundary. This classifier makes a fine-grained assessment of whether the current turn requires the capable tier before the final routing decision, trading a small additional inference cost for improved routing accuracy in ambiguous situations.

### How do I monitor cost savings from tiered routing?

Switchyard emits Prometheus metrics including the `switchyard_route_cost` metric with a `tier` label, allowing you to aggregate expenses by capable and efficient tiers separately. You can also inspect the `x-model-router-selected-model` header in each response to log which tier handled specific requests. These observability features, documented in [`docs/internal/metrics_reference.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/internal/metrics_reference.md), enable precise calculation of infrastructure savings versus quality trade-offs.