# How the LLM Classifier Routing Escalation Mode Works in Switchyard

> Understand Switchyard's LLM Classifier Routing escalation mode. Learn how it uses an efficient LLM, judge LLM, and configurable criteria for intelligent response routing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**Switchyard's escalation mode implements a response-based routing workflow that calls an efficient LLM first, uses a judge LLM to evaluate response quality, and escalates to a capable LLM only when the judge verdict meets configurable confirmation criteria.**

Switchyard, NVIDIA's open-source LLM routing framework, provides three built-in routing strategies through the **LlmClassifierConfig** type. The **LLM Classifier Routing escalation mode** enables intelligent cost optimization by attempting fast, inexpensive inference first while maintaining quality guarantees through automated escalation triggers. This pattern reduces infrastructure costs by routing only complex queries to expensive, high-capability models.

## Understanding the Escalation Workflow

The escalation classifier implements a three-stage decision pipeline:

1. **Initial inference** – The request routes to the `efficient_target`, a cheap or fast model that handles straightforward queries.

2. **Quality judgment** – A specialized `judge_target` evaluates the efficient model's response against quality or safety criteria.

3. **Conditional escalation** – If the judge returns consecutive affirmative verdicts meeting the `confirmations` threshold, the system re-routes the request to the `capable_target` and returns the higher-quality response.

### Escalation Annotations

The workflow attaches optional **escalation notes** and **de-escalation notes** to requests, providing observability hooks for downstream logging components to track routing decisions.

## Configuring the Escalation Classifier

The `EscalationClassifierConfig` class controls the escalation behavior. According to the Switchyard source code in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) (lines 81-98), the configuration accepts these parameters:

- `confirmations` – Number of consecutive "yes" judgments required before triggering escalation (default: 2)
- `recent_turn_window` – Conversation history turns examined during judgment (default: 28)
- `window_message_chars` – Minimum character threshold per turn for context inclusion (≥ 50)
- `max_output_tokens` – Token limit for the escalation target's response (default: 4096)
- `prompt` – Optional system prompt prepended to judge requests
- `response_format_type` – Expected judge output format (`json_schema` or `json_object`)

## Implementation in the Source Code

The escalation factory method resides in the Rust-backed Python bindings. The static method `LlmClassifierConfig.escalation` (lines 78-86 of [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py)) constructs a classifier instance binding the three targets together:

- `judge_target` – The lightweight evaluator model
- `efficient_target` – The fast/cheap initial inference model  
- `capable_target` – The high-quality fallback model

Internally, the Rust algorithm:
- Sends the initial request to the **efficient** endpoint
- Feeds the response to the **judge** endpoint for binary verdict classification
- Invokes the **capable** endpoint only when verdicts satisfy the confirmation rule

## Practical Implementation Example

Configure escalation routing using the Python API:

```python
from switchyard.libsy import (
    EscalationClassifierConfig,
    LlmClassifierConfig,
    llm_classifier,
    stage_router,
)

# Configure escalation criteria

escalation_cfg = EscalationClassifierConfig(
    confirmations=3,
    recent_turn_window=20,
    window_message_chars=200,
    max_output_tokens=2048,
    prompt="You are a safety judge. Decide if the answer is sufficient.",
    response_format_type="json_object",
)

# Create classifier binding three targets

classifier = LlmClassifierConfig.escalation(
    judge_target="safety-judge",
    efficient_target="fast-model",
    capable_target="high-quality",
    config=escalation_cfg,
)

# Build routing algorithm with observability hooks

routing_algo = stage_router(
    capable_target="high-quality",
    efficient_target="fast-model",
    picker="confidence",
    confidence_threshold=0.8,
    recent_window=10,
    escalation_note="Escalated after safety check failed",
    deescalation_note="Returned fast answer",
    only_on_wrong_signal_escalation=True,
    classifier=llm_classifier(classifier),
)

```

This configuration requires three consecutive affirmative judgments before escalating from the fast model to the high-quality model, with full audit trails attached to routing decisions.

## Summary

- The **escalation mode** implements a "try-fast-first" pattern using three distinct LLM targets: efficient, judge, and capable
- **EscalationClassifierConfig** controls confirmation thresholds, context windows, and output limits through parameters like `confirmations` and `recent_turn_window`
- The implementation lives in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), specifically the `EscalationClassifierConfig` definition (lines 81-98) and `LlmClassifierConfig.escalation` factory (lines 78-86)
- Optional **escalation notes** provide observability into routing decisions for downstream logging systems

## Frequently Asked Questions

### What triggers an escalation in Switchyard's classifier?

An escalation triggers when the judge LLM returns consecutive affirmative verdicts meeting the `confirmations` threshold defined in `EscalationClassifierConfig`. By default, the system requires two consecutive "yes" judgments, though this is configurable per deployment.

### How does the escalation mode differ from other Switchyard routing strategies?

Unlike static routing or confidence-based selection, the escalation mode specifically implements a sequential evaluation pattern that unconditionally calls the efficient model first, then conditionally escalates based on post-hoc quality assessment rather than pre-request classification.

### Can I configure different confirmation thresholds for different query types?

Yes, you instantiate separate `EscalationClassifierConfig` objects with varying `confirmations` values and create distinct classifiers using `LlmClassifierConfig.escalation()`. Each classifier instance can then be bound to different route patterns or API endpoints within your `stage_router` configuration.

### Where is the escalation logic actually executed?

The core escalation algorithm executes within the compiled Rust library loaded via [`switchyard_rust/_native.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/_native.py), while the configuration and target binding occur in the Python layer through [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py). The Python bindings marshal configuration data to the Rust runtime where the actual three-stage routing logic executes.