How to Implement Content-Based Routing for LLMs with Switchyard
Switchyard routes LLM requests through the LLM Classifier algorithm—a Rust-native content analyzer that inspects message payloads and selects backend models based on textual patterns, configurable via Python or TOML deployment files.
Content-based routing enables intelligent request distribution by analyzing the actual text content rather than just headers or metadata. In the NVIDIA-NeMo/Switchyard framework, this capability is provided by the LLM Classifier, a high-performance Rust algorithm exposed through a Python façade. This architecture allows you to route requests to specialized models (e.g., code-optimized vs. summarization-optimized) based on the semantic content of the prompt.
Architecture of Content-Based Routing in Switchyard
Switchyard implements content-based routing through a pipeline of algorithms that examine requests and decide which backend model should handle them. The architecture follows a client-server flow where the Python layer provides ergonomic access to Rust-native performance.
The core data flow is:
- Client sends an OpenAI-compatible request
- Switchyard Python façade receives the request via
switchyard/libsy/algorithms.py switchyard_rust.libsy.llm_classifieranalyzes the content asynchronously- A
RoutingDecisionselects the target backend (e.g.,model/fast,model/accurate)
The Python export layer in switchyard/libsy/algorithms.py re-exports Rust-owned routing algorithms so they can be referenced from Python code or TOML deployment files. The actual content analysis happens in the Rust crate under crates/switchyard-py, where the classifier inspects the request's messages field to determine the appropriate backend. Documentation for this algorithm is available in docs/routing_algorithms/llm_classifier_routing.md.
Algorithm Implementation Pattern
The LLM Classifier follows the same implementation pattern as other Switchyard algorithms such as stage_router. The reference implementation in crates/switchyard-server/src/stats/algorithms/stage_router.rs demonstrates how algorithms collect Prometheus metrics, compute deltas, and expose snapshots. The LLM Classifier uses this identical pattern to ensure consistent observability across all routing decisions.
How the LLM Classifier Works
The LLM Classifier operates as an asynchronous Rust algorithm that receives the full request object conforming to the OpenAI API schema. It extracts the messages field and traverses each message's content to handle raw text, structured blocks, or embeddings depending on configuration.
Based on extracted content, the classifier applies a rule-set—such as regex matching, keyword mapping, or lightweight ML models—to map requests to target identifiers. It returns a RoutingDecision that Switchyard uses to forward the request to the chosen backend. Because the algorithm lives in Rust, it runs efficiently at high throughput while the Python façade allows flexible configuration.
Implementing Content-Based Routing
Importing the Algorithm in Python
Access the classifier through the Python export layer:
from switchyard.libsy import llm_classifier
# The class exposed by the Rust crate is called LLMClassifier
Classifier = llm_classifier.LLMClassifier
Configuring Keyword-Based Routing
Create a classifier instance with a mapping of content patterns to backend models:
# Example: provide a JSON mapping of keywords → target model
keyword_map = {
"summarize": "model/fast",
"code": "model/strong",
"legal": "model/precise",
}
classifier = Classifier(keyword_map=keyword_map)
Route requests programmatically:
request = {
"model": "gpt-4",
"messages": [
{"role": "user", "content": [{"type": "text", "text": "Please summarize this article."}]}
],
}
target = classifier.route(request)
print(f"Chosen backend: {target}") # → "model/fast"
Deploying via TOML Configuration
For production deployments, configure the classifier in a routes.toml file that declares which algorithms to run and how they should be chained:
[algorithms]
# Use the built‑in LLM classifier for content routing
classifier = { type = "llm_classifier", config = "config/classifier.json" }
[routing]
# Chain the classifier with any downstream algorithm (e.g., stage router)
pipeline = ["classifier", "stage_router"]
Start the server with your configuration:
switchyard-server --config routes.toml --port 4000
Any request reaching the server will now be examined by the LLM Classifier and routed accordingly before subsequent algorithms in the pipeline process it.
Testing Your Routing Implementation
Verify your routing logic using mock HTTP clients. The following example uses respx to mock downstream model endpoints:
import respx
import httpx
from switchyard.libsy import llm_classifier
@respx.mock
async def test_content_routing():
# Mock the downstream model endpoint
mock = respx.post("https://api.fake/fast").mock(
return_value=httpx.Response(200, json={"answer": "OK"})
)
# Create a classifier that routes “summarize” → fast model
classifier = llm_classifier.LLMClassifier(
keyword_map={"summarize": "fast"}
)
# Simulate a Switchyard client call
client = httpx.AsyncClient(base_url="http://localhost:4000")
request = {
"model": "gpt-4",
"messages": [{"role": "user", "content": [{"type": "text", "text": "Summarize the report"}]}],
}
resp = await client.post("/v1/chat/completions", json=request)
assert resp.json()["answer"] == "OK"
assert mock.called
Summary
- Switchyard provides content-based routing through the LLM Classifier algorithm implemented in Rust and exposed via
switchyard/libsy/algorithms.py - The classifier analyzes OpenAI-compatible message payloads in
crates/switchyard-pyto route requests based on textual content rather than static rules - Configure routing logic through Python dictionaries for programmatic use or TOML files for server deployments
- The algorithm follows the same architectural pattern as
stage_routerincrates/switchyard-server/src/stats/algorithms/stage_router.rs, ensuring consistent metrics collection and snapshot exposure - High-throughput workloads benefit from the Rust-native asynchronous implementation while maintaining Python configurability
Frequently Asked Questions
What algorithm handles content-based routing in Switchyard?
The LLM Classifier algorithm handles content-based routing. Implemented in Rust within the switchyard_rust.libsy.llm_classifier crate and exposed through switchyard/libsy/algorithms.py, it inspects request messages and selects backend models based on configurable content rules. The algorithm supports keyword matching, regex patterns, and embedding-based similarity detection.
How do I configure the LLM Classifier for custom routing rules?
Pass a dictionary mapping keywords or patterns to target model identifiers when instantiating LLMClassifier in Python, or specify a JSON configuration file path in your TOML deployment under [algorithms]. The classifier supports灵活的 rule definitions including regex matching, keyword detection, and embedding-based similarity, allowing you to route "summarize" queries to fast models and "code" queries to strong models.
Can I chain the LLM Classifier with other routing algorithms?
Yes. The TOML configuration supports pipeline definitions where the classifier operates as the first stage, feeding its output into downstream algorithms like stage_router. Define the execution order in the [routing] section's pipeline array. This allows content-based selection to inform subsequent load-balancing or latency-optimization decisions.
Is the LLM Classifier suitable for high-throughput production workloads?
Yes. Because the algorithm runs in Rust with asynchronous execution and follows the production-grade patterns established in crates/switchyard-server/src/stats/algorithms/stage_router.rs, it handles high-throughput scenarios efficiently. The Python façade provides convenient configuration without sacrificing the performance needed for latency-sensitive LLM routing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →