How to Use the LLM Classifier Routing Algorithm in Switchyard
Switchyard's LLM Classifier routing algorithm dynamically routes requests between weak and strong language models using a classifier LLM that predicts task difficulty, supporting capability-based, escalation, and custom policy modes.
The LLM Classifier routing algorithm in NVIDIA-NeMo/Switchyard enables intelligent request distribution across heterogeneous model targets. By leveraging a dedicated classifier LLM, the system evaluates incoming prompts and automatically selects the most cost-effective target capable of handling the complexity. This approach optimizes latency and cost while maintaining response quality through configurable thresholds and fail-open safety mechanisms.
How the LLM Classifier Routing Algorithm Works
The algorithm operates by intercepting incoming requests and delegating them to a classifier target before final routing. This classifier analyzes the prompt difficulty and returns a structured verdict that determines the ultimate destination.
The workflow follows four distinct stages:
-
Configuration: Define routing parameters in a TOML configuration file, specifying classifier, weak, and strong targets.
-
Classification: Switchyard sends the user request to the classifier LLM, which returns a JSON verdict containing fields such as
p_solve,capability_boundary,primary_rule, andcrux. -
Routing Decision: Based on the verdict and selected mode, Switchyard chooses between the weak target, strong target, or intermediate options.
-
Fail-Open Handling: If the classifier response is missing, malformed, or judgment fails, the system routes to the strong target to ensure reliability.
Three Routing Modes Explained
The LlmClassifierConfig in crates/libsy/src/algorithms/llm_class.rs supports three distinct operational modes, each suited for different deployment scenarios.
Capability Mode
Capability mode uses a packaged prompt that estimates the probability (p_solve) that the weak model can complete the task. The algorithm applies threshold logic using base_threshold and threshold_step parameters to determine whether to route to the weak target or fall back to the strong target.
This mode is ideal for binary routing decisions between a cost-effective weak model and a premium strong model.
Escalation Mode
Escalation mode extends capability routing by adding trajectory escalation logic. Rather than binary weak/strong decisions, this mode supports promoting requests through intermediate targets based on progressive difficulty assessments.
Use this when your infrastructure includes multiple model tiers (e.g., fast, balanced, reasoning) and you want automatic promotion through the capability stack.
Custom Mode
Custom mode allows you to supply your own JSON schema and policy for target selection. Instead of relying on built-in probability thresholds, you define a custom verdict structure and specify a selector path (e.g., /decision/target) to extract the routing decision.
This mode supports arbitrary target configurations beyond the traditional weak/strong dichotomy.
Configuration Parameters and Verdict Schema
The routing behavior is controlled through several key configuration knobs defined in the TaskClassifierConfig and CustomClassifierConfig structures:
-
base_threshold: The minimump_solvevalue required for the classifier to route supported tasks to the weak model. -
threshold_step: The increment added to the base threshold for each boundary step (such asuncertainorunsupported), creating progressive routing tiers. -
classify_trigger: Controls when classification occurs—options includeevery_request,user_turn, ornew_session. -
message_hash_fallback: Enables reuse of first-message affinity when session metadata is absent, ensuring consistent routing for stateless clients. -
prompt: Optional override of the packaged classifier prompt, allowing customization of the difficulty assessment rubric. -
response_format_type: Specifies the output format as eitherjson_schema(default) orjson_objectfor providers without strict schema support.
Source Code Architecture
The LLM Classifier routing algorithm spans multiple crates and bindings within the Switchyard repository:
-
crates/libsy/src/algorithms/llm_class.rs: Contains the core Rust implementation ofLlmTaskClassifierand theLlmClassifierConfigenum. -
crates/switchyard-py/src/libsy_bindings.rs: Provides PyO3 bindings that bridge the Rust algorithm to Python, exposing thellm_classifierfunction. -
switchyard/libsy/algorithms.py: Serves as the Python façade that exposesalgorithms.llm_classifierto end users. -
tests/test_libsy_minimal_bindings.py: Contains unit tests demonstrating instantiation and execution patterns for the classifier. -
docs/routing_algorithms/llm_classifier_routing.md: Comprehensive documentation of configuration options and mode details.
Implementing the LLM Classifier in Python
Python users access the algorithm through switchyard.libsy.algorithms.llm_classifier. Below are practical implementations for common scenarios.
Binary Routing with Capability Mode
Configure a simple weak/strong binary router using the capability mode:
from switchyard.libsy import algorithms, LlmClassifierConfig, TaskClassifierConfig
# Build a capability-mode classifier that routes between "weak" and "strong"
algorithm = algorithms.llm_classifier(
LlmClassifierConfig.capability(
classifier_target="classifier", # LLM that judges the request
weak_target="weak", # cheap model
strong_target="strong", # premium model
config=TaskClassifierConfig(
base_threshold=0.5,
threshold_step=0.1,
prompt="Custom capability rubric.", # optional prompt override
),
)
)
Multi-Target Custom Routing
Define a custom schema to route across four distinct targets:
from switchyard.libsy import algorithms, LlmClassifierConfig, CustomClassifierConfig
custom_schema = """
{
"type": "object",
"properties": {
"decision": {
"type": "object",
"properties": {
"target": {
"type": "string",
"enum": ["fast", "balanced", "reasoning", "premium"]
}
},
"required": ["target"]
}
},
"required": ["decision"]
}
"""
algorithm = algorithms.llm_classifier(
LlmClassifierConfig.custom(
classifier_target="classifier",
targets=["fast", "balanced", "reasoning", "premium"],
default_target="premium",
config=CustomClassifierConfig(
prompt="Choose the best configured target for this request.",
response_schema=custom_schema,
selector="/decision/target",
),
)
)
Executing the Router
Run the configured algorithm against a set of LLM clients:
import asyncio
from switchyard.libsy import algorithms, LlmClassifierConfig, TaskClassifierConfig
async def demo():
alg = algorithms.llm_classifier(
LlmClassifierConfig.capability(
classifier_target="classifier",
weak_target="weak",
strong_target="strong",
config=TaskClassifierConfig(0.5, threshold_step=0.1),
)
)
# EchoClient is a lightweight mock that just echoes the model name
client = EchoClient("weak") # Replace with a real LLM client
selected, response = await run_algorithm(alg, {"classifier": client, "weak": client, "strong": client})
print(f"Selected target: {selected}, response model: {response['model']}")
asyncio.run(demo())
Summary
- The LLM Classifier routing algorithm in Switchyard uses a dedicated classifier LLM to predict request difficulty and route to appropriate targets.
- Three modes are available: capability (binary threshold-based), escalation (progressive promotion), and custom (user-defined schema).
- Configuration is managed through
TaskClassifierConfigorCustomClassifierConfig, with parameters likebase_thresholdandthreshold_stepcontrolling routing boundaries. - The implementation resides in
crates/libsy/src/algorithms/llm_class.rswith Python bindings exposed viaswitchyard.libsy.algorithms. - The system implements fail-open behavior, routing to the strong target when classification fails.
Frequently Asked Questions
What happens if the classifier LLM fails or returns malformed JSON?
Switchyard implements a fail-open safety mechanism. If the classifier response is missing, malformed, or the judgment logic encounters an error, the system automatically routes the request to the strong target. This ensures that temporary classifier failures do not degrade service quality.
How do I configure the classifier to run only on new sessions?
Set the classify_trigger parameter in your TaskClassifierConfig to new_session. Other valid options include every_request (classifies all incoming requests) and user_turn (classifies at specific interaction points). This allows you to optimize classification costs by avoiding redundant difficulty assessments within active sessions.
Can I use the LLM Classifier routing algorithm with more than two targets?
Yes, while capability and escalation modes traditionally handle weak/strong binary decisions, the custom mode supports arbitrary target configurations. Supply a JSON schema enumerating your targets (e.g., ["fast", "balanced", "reasoning", "premium"]) and specify the selector path where the classifier writes its routing decision.
Where is the core routing logic implemented in the Switchyard source code?
The core logic resides in crates/libsy/src/algorithms/llm_class.rs, which defines the LlmTaskClassifier struct and LlmClassifierConfig enum. Python bindings are generated in crates/switchyard-py/src/libsy_bindings.rs and exposed through switchyard/libsy/algorithms.py, allowing Python applications to instantiate and execute the classifier.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →