LLM Classifier Routing in Switchyard: 3 Modes Explained
Switchyard provides three distinct LLM classifier routing modes—Capability, Escalation, and Custom—that determine how requests are directed between efficient and capable models or arbitrary named targets via LlmClassifierConfig static factories.
The NVIDIA-NeMo/Switchyard framework implements intelligent request routing through LLM-based classification. Developers configure LLM classifier routing using the LlmClassifierConfig type defined in the Rust-backed Python façade, selecting from three distinct strategies that control how the classifier judge selects target models for incoming requests.
LLM Classifier Routing Configuration
The routing behavior is controlled by the LlmClassifierConfig class defined in switchyard_rust/libsy.py (lines 61-89). This configuration exposes three static factory methods that instantiate different routing strategies based on your performance and cost requirements. The chosen configuration is passed to the llm_classifier constructor exposed in switchyard/libsy/algorithms.py (lines 26-30), which builds the corresponding routing algorithm ready for pipeline integration.
The Three Routing Modes
Capability Mode
Capability mode routes requests based on the predicted capability requirements of the incoming task. The classifier evaluates whether the task requires an "efficient" or "capable" model and directs the request to the appropriate target.
Construct this mode using LlmClassifierConfig.capability():
from switchyard_rust.libsy import LlmClassifierConfig, TaskClassifierConfig
import switchyard.libsy.algorithms as algorithms
cap_cfg = LlmClassifierConfig.capability(
judge_target="switchyard/classifier",
efficient_target="my/efficient-model",
capable_target="my/capable-model",
config=TaskClassifierConfig(),
)
cap_algorithm = algorithms.llm_classifier(cap_cfg)
Escalation Mode
Escalation mode implements a fallback strategy to optimize costs. Requests are initially sent to an efficient target model. If the classifier judges the response insufficient for the task complexity, the request escalates automatically to a more capable target.
Configure escalation routing with LlmClassifierConfig.escalation():
esc_cfg = LlmClassifierConfig.escalation(
judge_target="switchyard/classifier",
efficient_target="my/efficient-model",
capable_target="my/capable-model",
config=EscalationClassifierConfig(),
)
esc_algorithm = algorithms.llm_classifier(esc_cfg)
Custom Mode
Custom mode enables routing among an arbitrary list of named targets using schema-selected labels returned by the classifier. This supports complex routing topologies beyond binary decisions, allowing you to route to specialized models based on content type.
Implement custom routing with LlmClassifierConfig.custom():
custom_cfg = LlmClassifierConfig.custom(
judge_target="switchyard/classifier",
targets=[("news", "my/news-model"), ("code", "my/code-model")],
default_target="my/default-model",
config=CustomClassifierConfig(),
)
custom_algorithm = algorithms.llm_classifier(custom_cfg)
Response Format Configuration
In addition to routing modes, the classifier supports two output formats controlled by the response_format_type parameter defined in the LlmClassifierConfig constructor (line 57 of switchyard_rust/libsy.py). The default value is "json_schema", with "json_object" available as an alternative for downstream parsing compatibility.
json_obj_cfg = LlmClassifierConfig.capability(
judge_target="switchyard/classifier",
efficient_target="my/efficient-model",
capable_target="my/capable-model",
config=TaskClassifierConfig(),
response_format_type="json_object",
)
Implementation Details
The routing mode logic resides in switchyard_rust/libsy.py, where LlmClassifierConfig defines the static factory methods at lines 61-89. Each factory accepts a specific configuration type: TaskClassifierConfig for Capability mode, EscalationClassifierConfig for Escalation mode, and CustomClassifierConfig for Custom mode.
The actual algorithm instantiation occurs in switchyard/libsy/algorithms.py through the llm_classifier() function (lines 26-30), which consumes the configuration and returns an Algorithm instance. This instance can be plugged into a Switchyard pipeline via stage_router or used directly in request handling loops.
Summary
- Capability mode selects between efficient and capable targets based on predicted task requirements using
LlmClassifierConfig.capability(). - Escalation mode routes to efficient targets first, falling back to capable models when judged necessary via
LlmClassifierConfig.escalation(). - Custom mode supports arbitrary target lists with label-based routing through
LlmClassifierConfig.custom(), enabling complex multi-model topologies. - All configurations are defined in
switchyard_rust/libsy.pyand instantiated throughswitchyard/libsy/algorithms.py. - Output parsing can use
"json_schema"(default) or"json_object"formats via theresponse_format_typeparameter.
Frequently Asked Questions
What is the default response format for Switchyard LLM classifiers?
The default response format is "json_schema", as defined at line 57 of switchyard_rust/libsy.py. You can override this to "json_object" by passing response_format_type="json_object" to any LlmClassifierConfig factory method to match your downstream parsing requirements.
How do I route between more than two models?
Use Custom mode via LlmClassifierConfig.custom(). Pass a list of (label, target) tuples to the targets parameter, along with a default_target for fallback routing when the classifier returns labels that do not match your defined targets.
Where is the routing logic implemented?
The routing mode definitions and static factories are implemented in switchyard_rust/libsy.py (lines 61-89). The executable algorithm is constructed in switchyard/libsy/algorithms.py (lines 26-30) through the llm_classifier() function, which produces an Algorithm instance from the configuration.
Can I combine different routing modes?
No, each LlmClassifierConfig instance represents a single routing mode selected exclusively through its specific factory method (capability(), escalation(), or custom()). To implement complex multi-stage routing logic, chain multiple classifier algorithms sequentially in your Switchyard pipeline rather than combining modes within a single configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →