# How to Develop a LiteLLM Routing Plugin for Switchyard Using switchyard_litellm

> Learn to develop a LiteLLM routing plugin for Switchyard. Implement a CustomLogger, intercept requests with _request, and use build_request_patch to modify them with Switchyard's Python bindings.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-12

---

**You can develop a LiteLLM routing plugin for Switchyard by implementing a `CustomLogger` subclass that intercepts requests via the `_request` hook, processes them through Switchyard's Python bindings, and returns a modified request object using the `build_request_patch` helper.**

The NVIDIA-NeMo/Switchyard repository provides a reference implementation for integrating Switchyard's advanced LLM routing capabilities with LiteLLM through the `switchyard_litellm` library. This integration allows you to leverage Switchyard's deployment selection algorithms while maintaining LiteLLM's unified API interface. By developing a LiteLLM routing plugin for Switchyard, you can dynamically route requests to optimal model deployments based on custom policies, latency requirements, or cost constraints.

## Architecture of the Switchyard LiteLLM Routing Plugin

The integration follows a layered architecture where LiteLLM's request lifecycle hooks into Switchyard's routing core. LiteLLM forwards incoming chat completion requests to the `CustomLogger._request` method, which the `SwitchyardRoutingPlugin` implements to call Switchyard's routing logic. After selecting a deployment, the plugin optionally modifies the request using the `build_request_patch` utility before returning the enriched payload to LiteLLM for final transmission to the chosen model endpoint.

Key files in the `examples/litellm/` directory implement this flow:

- [`examples/litellm/plugins/switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/plugins/switchyard_routing_plugin.py): Contains the `SwitchyardRoutingPlugin` class extending `CustomLogger`.
- [`examples/litellm/plugins/request_rewrite.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/plugins/request_rewrite.py): Provides the `build_request_patch` utility for modifying requests.
- [`examples/litellm/configuration.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/configuration.py): Exposes the `load_routing_plugin` factory function.

## Implementing the Core Plugin Components

### Extending CustomLogger in switchyard_routing_plugin.py

The plugin class must inherit from `litellm.integrations.custom_logger.CustomLogger` and implement the async `_request` method. This method receives the LiteLLM-converted request and is responsible for invoking Switchyard's routing logic.

```python

# examples/litellm/plugins/switchyard_routing_plugin.py

from __future__ import annotations
import asyncio
from typing import Any
from litellm.integrations.custom_logger import CustomLogger

class SwitchyardRoutingPlugin(CustomLogger):
    async def _request(self, litellm_messages: Any) -> Any:
        """
        Intercept the LiteLLM request, route through Switchyard, and return patched result.
        """
        # Invoke Switchyard core routing (implementation-specific call)

        routing_result = await self._select_deployment(litellm_messages)
        
        # Apply request modifications using the rewrite utility

        from .request_rewrite import build_request_patch
        patch = build_request_patch(
            litellm_messages,
            model=routing_result["model"],
            stage="prod",
        )
        
        return {**litellm_messages, **patch}
    
    async def _select_deployment(self, messages: Any) -> dict:
        """Placeholder for Switchyard routing logic."""
        # Integrate with switchyard_py bindings here

        return {"model": "gpt-4o"}

```

### Rewriting Requests with build_request_patch

The `build_request_patch` function in [`examples/litellm/plugins/request_rewrite.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/plugins/request_rewrite.py) constructs partial request dictionaries that LiteLLM merges with the original request. This preserves critical fields like the `messages` list while injecting routing-specific metadata.

```python

# examples/litellm/plugins/request_rewrite.py

from typing import List, Dict, Any

def build_request_patch(
    messages: List[Dict[str, Any]],
    *,
    model: str,
    stage: str,
    **kwargs,
) -> Dict[str, Any]:
    """
    Build a partial request patch preserving fields through LiteLLM's conversion.
    
    Parameters:
        messages: List of chat messages from the original request.
        model: Target model identifier selected by Switchyard.
        stage: Deployment stage (e.g., 'prod', 'dev').
        **kwargs: Additional metadata to include.
        
    Returns:
        Dictionary containing fields to merge into the final request.
    """
    return {
        "model": model,
        "metadata": {"stage": stage, **kwargs},
    }

```

### Plugin Factory in configuration.py

LiteLLM instantiates the plugin through a factory function defined in [`examples/litellm/configuration.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/configuration.py). This pattern allows dynamic configuration loading from YAML files.

```python

# examples/litellm/configuration.py

from typing import Dict, Any
from switchyard_litellm.plugins.switchyard_routing_plugin import SwitchyardRoutingPlugin

def load_routing_plugin(config: Dict[str, Any]) -> SwitchyardRoutingPlugin:
    """
    Factory used by LiteLLM to create the routing plugin instance.
    
    Parameters:
        config: Configuration dictionary from the LiteLLM router config.
        
    Returns:
        Configured SwitchyardRoutingPlugin instance.
    """
    return SwitchyardRoutingPlugin()

```

## Registering and Testing Your Plugin

### Registration with LiteLLM

To activate the plugin, reference the factory in your LiteLLM router configuration:

```yaml

# router_config.yaml

custom_logger: switchyard_litellm.configuration.load_routing_plugin

```

LiteLLM calls this factory during initialization and passes the `_request` hook for every chat completion request.

### Unit Testing the Routing Logic

The [`examples/litellm/tests/unit/test_switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/tests/unit/test_switchyard_routing_plugin.py) file validates that the plugin correctly preserves LiteLLM fields and applies routing decisions.

```python

# examples/litellm/tests/unit/test_switchyard_routing_plugin.py

import pytest
from litellm.types.router import RoutingContext
from switchyard_litellm import SwitchyardRoutingPlugin
from switchyard_litellm.plugins.request_rewrite import build_request_patch

@pytest.mark.asyncio
async def test_request_routing():
    plugin = SwitchyardRoutingPlugin()
    context = RoutingContext(
        messages=[{"role": "user", "content": "Hello"}],
        model="gpt-4"
    )
    
    result = await plugin._request(context)
    
    assert "model" in result
    assert result["metadata"]["stage"] == "prod"

def test_build_request_patch():
    messages = [{"role": "user", "content": "test"}]
    patch = build_request_patch(messages, model="claude-3", stage="dev", custom_key="value")
    
    assert patch["model"] == "claude-3"
    assert patch["metadata"]["stage"] == "dev"
    assert patch["metadata"]["custom_key"] == "value"

```

### Deployment Configuration Profiles

Switchyard uses YAML profiles to define routing strategies. The example includes two reference profiles:

- **[`examples/litellm/deployment/profiles/stage/litellm.yaml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/deployment/profiles/stage/litellm.yaml)**: Production-ready configuration with specific deployment targets and priority weights.
- **[`examples/litellm/deployment/profiles/random/litellm.yaml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/deployment/profiles/random/litellm.yaml)**: Random selection profile for testing load distribution across endpoints.

Place your custom profiles in this directory structure and reference them in your Switchyard runner configuration to control how the plugin selects between model deployments.

## Summary

- **Implement** a `CustomLogger` subclass in [`switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_routing_plugin.py) with an async `_request` method to intercept LiteLLM traffic before it reaches the model provider.
- **Use** `build_request_patch` from [`request_rewrite.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/request_rewrite.py) to safely inject routing decisions and metadata without losing LiteLLM's converted message fields.
- **Expose** a `load_routing_plugin` factory in [`configuration.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/configuration.py) so LiteLLM can instantiate your plugin from YAML configuration using the `custom_logger` key.
- **Test** your implementation using the patterns in [`test_switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/test_switchyard_routing_plugin.py), verifying that `RoutingContext` fields propagate correctly through the `_request` lifecycle.
- **Configure** deployment profiles in `examples/litellm/deployment/profiles/` to define selection algorithms, cost constraints, and endpoint priorities for the routing plugin.

## Frequently Asked Questions

### What is the relationship between Switchyard and LiteLLM in this plugin architecture?

Switchyard provides the core routing algorithms and deployment selection logic written in Rust, while LiteLLM offers a unified API gateway to multiple LLM providers. The plugin acts as a bridge, using LiteLLM's `CustomLogger` hook to call Switchyard's Python bindings for routing decisions before the request reaches the final model endpoint.

### How does the build_request_patch function preserve request fields?

The `build_request_patch` function in [`examples/litellm/plugins/request_rewrite.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/plugins/request_rewrite.py) returns a partial dictionary containing only the fields to update, such as the selected `model` identifier and `stage` metadata. By returning a patch rather than a complete request object, the function ensures that LiteLLM's internal message conversion preserves the original `messages` list and other critical fields while merging in the routing-specific updates.

### Can I use synchronous Switchyard calls within the _request hook?

While the `_request` hook is defined as `async` to comply with LiteLLM's `CustomLogger` interface, Switchyard's Python bindings may expose synchronous Rust functions. Use `asyncio.to_thread()` to wrap synchronous Switchyard calls, as shown in the [`switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_routing_plugin.py) pattern, to prevent blocking the LiteLLM event loop during routing decisions.

### Where should I place custom deployment profiles for the plugin?

Place YAML configuration files in the `examples/litellm/deployment/profiles/` directory, following the structure of the existing `stage` and `random` profiles. The Switchyard runner loads these profiles at initialization to determine selection algorithms, cost constraints, and endpoint priorities for the routing plugin.