# How to Implement Content-Based Routing for LLMs with Switchyard

> Implement content-based routing for LLMs with Switchyard. Inspect message payloads with the LLM Classifier and select backend models using textual patterns. Configure via Python or TOML.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-21

---

**Switchyard routes LLM requests through the LLM Classifier algorithm—a Rust-native content analyzer that inspects message payloads and selects backend models based on textual patterns, configurable via Python or TOML deployment files.**

Content-based routing enables intelligent request distribution by analyzing the actual text content rather than just headers or metadata. In the NVIDIA-NeMo/Switchyard framework, this capability is provided by the **LLM Classifier**, a high-performance Rust algorithm exposed through a Python façade. This architecture allows you to route requests to specialized models (e.g., code-optimized vs. summarization-optimized) based on the semantic content of the prompt.

## Architecture of Content-Based Routing in Switchyard

Switchyard implements content-based routing through a pipeline of algorithms that examine requests and decide which backend model should handle them. The architecture follows a client-server flow where the Python layer provides ergonomic access to Rust-native performance.

The core data flow is:

- Client sends an OpenAI-compatible request
- Switchyard Python façade receives the request via [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)
- `switchyard_rust.libsy.llm_classifier` analyzes the content asynchronously
- A `RoutingDecision` selects the target backend (e.g., `model/fast`, `model/accurate`)

The Python export layer in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) re-exports Rust-owned routing algorithms so they can be referenced from Python code or TOML deployment files. The actual content analysis happens in the Rust crate under `crates/switchyard-py`, where the classifier inspects the request's `messages` field to determine the appropriate backend. Documentation for this algorithm is available in [`docs/routing_algorithms/llm_classifier_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/llm_classifier_routing.md).

### Algorithm Implementation Pattern

The LLM Classifier follows the same implementation pattern as other Switchyard algorithms such as `stage_router`. The reference implementation in [`crates/switchyard-server/src/stats/algorithms/stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs) demonstrates how algorithms collect Prometheus metrics, compute deltas, and expose snapshots. The LLM Classifier uses this identical pattern to ensure consistent observability across all routing decisions.

## How the LLM Classifier Works

The LLM Classifier operates as an asynchronous Rust algorithm that receives the full request object conforming to the OpenAI API schema. It extracts the `messages` field and traverses each message's `content` to handle raw text, structured blocks, or embeddings depending on configuration.

Based on extracted content, the classifier applies a rule-set—such as regex matching, keyword mapping, or lightweight ML models—to map requests to target identifiers. It returns a `RoutingDecision` that Switchyard uses to forward the request to the chosen backend. Because the algorithm lives in Rust, it runs efficiently at high throughput while the Python façade allows flexible configuration.

## Implementing Content-Based Routing

### Importing the Algorithm in Python

Access the classifier through the Python export layer:

```python
from switchyard.libsy import llm_classifier

# The class exposed by the Rust crate is called LLMClassifier

Classifier = llm_classifier.LLMClassifier

```

### Configuring Keyword-Based Routing

Create a classifier instance with a mapping of content patterns to backend models:

```python

# Example: provide a JSON mapping of keywords → target model

keyword_map = {
    "summarize": "model/fast",
    "code": "model/strong",
    "legal": "model/precise",
}

classifier = Classifier(keyword_map=keyword_map)

```

Route requests programmatically:

```python
request = {
    "model": "gpt-4",
    "messages": [
        {"role": "user", "content": [{"type": "text", "text": "Please summarize this article."}]}
    ],
}

target = classifier.route(request)
print(f"Chosen backend: {target}")  # → "model/fast"

```

### Deploying via TOML Configuration

For production deployments, configure the classifier in a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file that declares which algorithms to run and how they should be chained:

```toml
[algorithms]

# Use the built‑in LLM classifier for content routing

classifier = { type = "llm_classifier", config = "config/classifier.json" }

[routing]

# Chain the classifier with any downstream algorithm (e.g., stage router)

pipeline = ["classifier", "stage_router"]

```

Start the server with your configuration:

```bash
switchyard-server --config routes.toml --port 4000

```

Any request reaching the server will now be examined by the LLM Classifier and routed accordingly before subsequent algorithms in the pipeline process it.

## Testing Your Routing Implementation

Verify your routing logic using mock HTTP clients. The following example uses `respx` to mock downstream model endpoints:

```python
import respx
import httpx
from switchyard.libsy import llm_classifier

@respx.mock
async def test_content_routing():
    # Mock the downstream model endpoint

    mock = respx.post("https://api.fake/fast").mock(
        return_value=httpx.Response(200, json={"answer": "OK"})
    )

    # Create a classifier that routes “summarize” → fast model

    classifier = llm_classifier.LLMClassifier(
        keyword_map={"summarize": "fast"}
    )

    # Simulate a Switchyard client call

    client = httpx.AsyncClient(base_url="http://localhost:4000")
    request = {
        "model": "gpt-4",
        "messages": [{"role": "user", "content": [{"type": "text", "text": "Summarize the report"}]}],
    }
    
    resp = await client.post("/v1/chat/completions", json=request)
    assert resp.json()["answer"] == "OK"
    assert mock.called

```

## Summary

- Switchyard provides **content-based routing** through the LLM Classifier algorithm implemented in Rust and exposed via [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)
- The classifier analyzes OpenAI-compatible message payloads in `crates/switchyard-py` to route requests based on textual content rather than static rules
- Configure routing logic through Python dictionaries for programmatic use or TOML files for server deployments
- The algorithm follows the same architectural pattern as `stage_router` in [`crates/switchyard-server/src/stats/algorithms/stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs), ensuring consistent metrics collection and snapshot exposure
- High-throughput workloads benefit from the Rust-native asynchronous implementation while maintaining Python configurability

## Frequently Asked Questions

### What algorithm handles content-based routing in Switchyard?

The **LLM Classifier** algorithm handles content-based routing. Implemented in Rust within the `switchyard_rust.libsy.llm_classifier` crate and exposed through [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py), it inspects request messages and selects backend models based on configurable content rules. The algorithm supports keyword matching, regex patterns, and embedding-based similarity detection.

### How do I configure the LLM Classifier for custom routing rules?

Pass a dictionary mapping keywords or patterns to target model identifiers when instantiating `LLMClassifier` in Python, or specify a JSON configuration file path in your TOML deployment under `[algorithms]`. The classifier supports灵活的 rule definitions including regex matching, keyword detection, and embedding-based similarity, allowing you to route "summarize" queries to fast models and "code" queries to strong models.

### Can I chain the LLM Classifier with other routing algorithms?

Yes. The TOML configuration supports pipeline definitions where the classifier operates as the first stage, feeding its output into downstream algorithms like `stage_router`. Define the execution order in the `[routing]` section's `pipeline` array. This allows content-based selection to inform subsequent load-balancing or latency-optimization decisions.

### Is the LLM Classifier suitable for high-throughput production workloads?

Yes. Because the algorithm runs in Rust with asynchronous execution and follows the production-grade patterns established in [`crates/switchyard-server/src/stats/algorithms/stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs), it handles high-throughput scenarios efficiently. The Python façade provides convenient configuration without sacrificing the performance needed for latency-sensitive LLM routing.