How to Implement Content-Based Routing for LLMs with Switchyard

Switchyard routes LLM requests through the LLM Classifier algorithm—a Rust-native content analyzer that inspects message payloads and selects backend models based on textual patterns, configurable via Python or TOML deployment files.

Content-based routing enables intelligent request distribution by analyzing the actual text content rather than just headers or metadata. In the NVIDIA-NeMo/Switchyard framework, this capability is provided by the LLM Classifier, a high-performance Rust algorithm exposed through a Python façade. This architecture allows you to route requests to specialized models (e.g., code-optimized vs. summarization-optimized) based on the semantic content of the prompt.

Architecture of Content-Based Routing in Switchyard

Switchyard implements content-based routing through a pipeline of algorithms that examine requests and decide which backend model should handle them. The architecture follows a client-server flow where the Python layer provides ergonomic access to Rust-native performance.

The core data flow is:

  • Client sends an OpenAI-compatible request
  • Switchyard Python façade receives the request via switchyard/libsy/algorithms.py
  • switchyard_rust.libsy.llm_classifier analyzes the content asynchronously
  • A RoutingDecision selects the target backend (e.g., model/fast, model/accurate)

The Python export layer in switchyard/libsy/algorithms.py re-exports Rust-owned routing algorithms so they can be referenced from Python code or TOML deployment files. The actual content analysis happens in the Rust crate under crates/switchyard-py, where the classifier inspects the request's messages field to determine the appropriate backend. Documentation for this algorithm is available in docs/routing_algorithms/llm_classifier_routing.md.

Algorithm Implementation Pattern

The LLM Classifier follows the same implementation pattern as other Switchyard algorithms such as stage_router. The reference implementation in crates/switchyard-server/src/stats/algorithms/stage_router.rs demonstrates how algorithms collect Prometheus metrics, compute deltas, and expose snapshots. The LLM Classifier uses this identical pattern to ensure consistent observability across all routing decisions.

How the LLM Classifier Works

The LLM Classifier operates as an asynchronous Rust algorithm that receives the full request object conforming to the OpenAI API schema. It extracts the messages field and traverses each message's content to handle raw text, structured blocks, or embeddings depending on configuration.

Based on extracted content, the classifier applies a rule-set—such as regex matching, keyword mapping, or lightweight ML models—to map requests to target identifiers. It returns a RoutingDecision that Switchyard uses to forward the request to the chosen backend. Because the algorithm lives in Rust, it runs efficiently at high throughput while the Python façade allows flexible configuration.

Implementing Content-Based Routing

Importing the Algorithm in Python

Access the classifier through the Python export layer:

from switchyard.libsy import llm_classifier

# The class exposed by the Rust crate is called LLMClassifier

Classifier = llm_classifier.LLMClassifier

Configuring Keyword-Based Routing

Create a classifier instance with a mapping of content patterns to backend models:


# Example: provide a JSON mapping of keywords → target model

keyword_map = {
    "summarize": "model/fast",
    "code": "model/strong",
    "legal": "model/precise",
}

classifier = Classifier(keyword_map=keyword_map)

Route requests programmatically:

request = {
    "model": "gpt-4",
    "messages": [
        {"role": "user", "content": [{"type": "text", "text": "Please summarize this article."}]}
    ],
}

target = classifier.route(request)
print(f"Chosen backend: {target}")  # → "model/fast"

Deploying via TOML Configuration

For production deployments, configure the classifier in a routes.toml file that declares which algorithms to run and how they should be chained:

[algorithms]

# Use the built‑in LLM classifier for content routing

classifier = { type = "llm_classifier", config = "config/classifier.json" }

[routing]

# Chain the classifier with any downstream algorithm (e.g., stage router)

pipeline = ["classifier", "stage_router"]

Start the server with your configuration:

switchyard-server --config routes.toml --port 4000

Any request reaching the server will now be examined by the LLM Classifier and routed accordingly before subsequent algorithms in the pipeline process it.

Testing Your Routing Implementation

Verify your routing logic using mock HTTP clients. The following example uses respx to mock downstream model endpoints:

import respx
import httpx
from switchyard.libsy import llm_classifier

@respx.mock
async def test_content_routing():
    # Mock the downstream model endpoint

    mock = respx.post("https://api.fake/fast").mock(
        return_value=httpx.Response(200, json={"answer": "OK"})
    )

    # Create a classifier that routes “summarize” → fast model

    classifier = llm_classifier.LLMClassifier(
        keyword_map={"summarize": "fast"}
    )

    # Simulate a Switchyard client call

    client = httpx.AsyncClient(base_url="http://localhost:4000")
    request = {
        "model": "gpt-4",
        "messages": [{"role": "user", "content": [{"type": "text", "text": "Summarize the report"}]}],
    }
    
    resp = await client.post("/v1/chat/completions", json=request)
    assert resp.json()["answer"] == "OK"
    assert mock.called

Summary

  • Switchyard provides content-based routing through the LLM Classifier algorithm implemented in Rust and exposed via switchyard/libsy/algorithms.py
  • The classifier analyzes OpenAI-compatible message payloads in crates/switchyard-py to route requests based on textual content rather than static rules
  • Configure routing logic through Python dictionaries for programmatic use or TOML files for server deployments
  • The algorithm follows the same architectural pattern as stage_router in crates/switchyard-server/src/stats/algorithms/stage_router.rs, ensuring consistent metrics collection and snapshot exposure
  • High-throughput workloads benefit from the Rust-native asynchronous implementation while maintaining Python configurability

Frequently Asked Questions

What algorithm handles content-based routing in Switchyard?

The LLM Classifier algorithm handles content-based routing. Implemented in Rust within the switchyard_rust.libsy.llm_classifier crate and exposed through switchyard/libsy/algorithms.py, it inspects request messages and selects backend models based on configurable content rules. The algorithm supports keyword matching, regex patterns, and embedding-based similarity detection.

How do I configure the LLM Classifier for custom routing rules?

Pass a dictionary mapping keywords or patterns to target model identifiers when instantiating LLMClassifier in Python, or specify a JSON configuration file path in your TOML deployment under [algorithms]. The classifier supports灵活的 rule definitions including regex matching, keyword detection, and embedding-based similarity, allowing you to route "summarize" queries to fast models and "code" queries to strong models.

Can I chain the LLM Classifier with other routing algorithms?

Yes. The TOML configuration supports pipeline definitions where the classifier operates as the first stage, feeding its output into downstream algorithms like stage_router. Define the execution order in the [routing] section's pipeline array. This allows content-based selection to inform subsequent load-balancing or latency-optimization decisions.

Is the LLM Classifier suitable for high-throughput production workloads?

Yes. Because the algorithm runs in Rust with asynchronous execution and follows the production-grade patterns established in crates/switchyard-server/src/stats/algorithms/stage_router.rs, it handles high-throughput scenarios efficiently. The Python façade provides convenient configuration without sacrificing the performance needed for latency-sensitive LLM routing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →