How to Use the LLM Classifier Routing Algorithm in Switchyard

Switchyard's LLM Classifier routing algorithm dynamically routes requests between weak and strong language models using a classifier LLM that predicts task difficulty, supporting capability-based, escalation, and custom policy modes.

The LLM Classifier routing algorithm in NVIDIA-NeMo/Switchyard enables intelligent request distribution across heterogeneous model targets. By leveraging a dedicated classifier LLM, the system evaluates incoming prompts and automatically selects the most cost-effective target capable of handling the complexity. This approach optimizes latency and cost while maintaining response quality through configurable thresholds and fail-open safety mechanisms.

How the LLM Classifier Routing Algorithm Works

The algorithm operates by intercepting incoming requests and delegating them to a classifier target before final routing. This classifier analyzes the prompt difficulty and returns a structured verdict that determines the ultimate destination.

The workflow follows four distinct stages:

  1. Configuration: Define routing parameters in a TOML configuration file, specifying classifier, weak, and strong targets.

  2. Classification: Switchyard sends the user request to the classifier LLM, which returns a JSON verdict containing fields such as p_solve, capability_boundary, primary_rule, and crux.

  3. Routing Decision: Based on the verdict and selected mode, Switchyard chooses between the weak target, strong target, or intermediate options.

  4. Fail-Open Handling: If the classifier response is missing, malformed, or judgment fails, the system routes to the strong target to ensure reliability.

Three Routing Modes Explained

The LlmClassifierConfig in crates/libsy/src/algorithms/llm_class.rs supports three distinct operational modes, each suited for different deployment scenarios.

Capability Mode

Capability mode uses a packaged prompt that estimates the probability (p_solve) that the weak model can complete the task. The algorithm applies threshold logic using base_threshold and threshold_step parameters to determine whether to route to the weak target or fall back to the strong target.

This mode is ideal for binary routing decisions between a cost-effective weak model and a premium strong model.

Escalation Mode

Escalation mode extends capability routing by adding trajectory escalation logic. Rather than binary weak/strong decisions, this mode supports promoting requests through intermediate targets based on progressive difficulty assessments.

Use this when your infrastructure includes multiple model tiers (e.g., fast, balanced, reasoning) and you want automatic promotion through the capability stack.

Custom Mode

Custom mode allows you to supply your own JSON schema and policy for target selection. Instead of relying on built-in probability thresholds, you define a custom verdict structure and specify a selector path (e.g., /decision/target) to extract the routing decision.

This mode supports arbitrary target configurations beyond the traditional weak/strong dichotomy.

Configuration Parameters and Verdict Schema

The routing behavior is controlled through several key configuration knobs defined in the TaskClassifierConfig and CustomClassifierConfig structures:

  • base_threshold: The minimum p_solve value required for the classifier to route supported tasks to the weak model.

  • threshold_step: The increment added to the base threshold for each boundary step (such as uncertain or unsupported), creating progressive routing tiers.

  • classify_trigger: Controls when classification occurs—options include every_request, user_turn, or new_session.

  • message_hash_fallback: Enables reuse of first-message affinity when session metadata is absent, ensuring consistent routing for stateless clients.

  • prompt: Optional override of the packaged classifier prompt, allowing customization of the difficulty assessment rubric.

  • response_format_type: Specifies the output format as either json_schema (default) or json_object for providers without strict schema support.

Source Code Architecture

The LLM Classifier routing algorithm spans multiple crates and bindings within the Switchyard repository:

Implementing the LLM Classifier in Python

Python users access the algorithm through switchyard.libsy.algorithms.llm_classifier. Below are practical implementations for common scenarios.

Binary Routing with Capability Mode

Configure a simple weak/strong binary router using the capability mode:

from switchyard.libsy import algorithms, LlmClassifierConfig, TaskClassifierConfig

# Build a capability-mode classifier that routes between "weak" and "strong"

algorithm = algorithms.llm_classifier(
    LlmClassifierConfig.capability(
        classifier_target="classifier",   # LLM that judges the request

        weak_target="weak",               # cheap model

        strong_target="strong",           # premium model

        config=TaskClassifierConfig(
            base_threshold=0.5,
            threshold_step=0.1,
            prompt="Custom capability rubric.",  # optional prompt override

        ),
    )
)

Multi-Target Custom Routing

Define a custom schema to route across four distinct targets:

from switchyard.libsy import algorithms, LlmClassifierConfig, CustomClassifierConfig

custom_schema = """
{
  "type": "object",
  "properties": {
    "decision": {
      "type": "object",
      "properties": {
        "target": {
          "type": "string",
          "enum": ["fast", "balanced", "reasoning", "premium"]
        }
      },
      "required": ["target"]
    }
  },
  "required": ["decision"]
}
"""

algorithm = algorithms.llm_classifier(
    LlmClassifierConfig.custom(
        classifier_target="classifier",
        targets=["fast", "balanced", "reasoning", "premium"],
        default_target="premium",
        config=CustomClassifierConfig(
            prompt="Choose the best configured target for this request.",
            response_schema=custom_schema,
            selector="/decision/target",
        ),
    )
)

Executing the Router

Run the configured algorithm against a set of LLM clients:

import asyncio
from switchyard.libsy import algorithms, LlmClassifierConfig, TaskClassifierConfig

async def demo():
    alg = algorithms.llm_classifier(
        LlmClassifierConfig.capability(
            classifier_target="classifier",
            weak_target="weak",
            strong_target="strong",
            config=TaskClassifierConfig(0.5, threshold_step=0.1),
        )
    )
    # EchoClient is a lightweight mock that just echoes the model name

    client = EchoClient("weak")          # Replace with a real LLM client

    selected, response = await run_algorithm(alg, {"classifier": client, "weak": client, "strong": client})
    print(f"Selected target: {selected}, response model: {response['model']}")

asyncio.run(demo())

Summary

  • The LLM Classifier routing algorithm in Switchyard uses a dedicated classifier LLM to predict request difficulty and route to appropriate targets.
  • Three modes are available: capability (binary threshold-based), escalation (progressive promotion), and custom (user-defined schema).
  • Configuration is managed through TaskClassifierConfig or CustomClassifierConfig, with parameters like base_threshold and threshold_step controlling routing boundaries.
  • The implementation resides in crates/libsy/src/algorithms/llm_class.rs with Python bindings exposed via switchyard.libsy.algorithms.
  • The system implements fail-open behavior, routing to the strong target when classification fails.

Frequently Asked Questions

What happens if the classifier LLM fails or returns malformed JSON?

Switchyard implements a fail-open safety mechanism. If the classifier response is missing, malformed, or the judgment logic encounters an error, the system automatically routes the request to the strong target. This ensures that temporary classifier failures do not degrade service quality.

How do I configure the classifier to run only on new sessions?

Set the classify_trigger parameter in your TaskClassifierConfig to new_session. Other valid options include every_request (classifies all incoming requests) and user_turn (classifies at specific interaction points). This allows you to optimize classification costs by avoiding redundant difficulty assessments within active sessions.

Can I use the LLM Classifier routing algorithm with more than two targets?

Yes, while capability and escalation modes traditionally handle weak/strong binary decisions, the custom mode supports arbitrary target configurations. Supply a JSON schema enumerating your targets (e.g., ["fast", "balanced", "reasoning", "premium"]) and specify the selector path where the classifier writes its routing decision.

Where is the core routing logic implemented in the Switchyard source code?

The core logic resides in crates/libsy/src/algorithms/llm_class.rs, which defines the LlmTaskClassifier struct and LlmClassifierConfig enum. Python bindings are generated in crates/switchyard-py/src/libsy_bindings.rs and exposed through switchyard/libsy/algorithms.py, allowing Python applications to instantiate and execute the classifier.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →