How to Develop a LiteLLM Routing Plugin for Switchyard Using switchyard_litellm

You can develop a LiteLLM routing plugin for Switchyard by implementing a CustomLogger subclass that intercepts requests via the _request hook, processes them through Switchyard's Python bindings, and returns a modified request object using the build_request_patch helper.

The NVIDIA-NeMo/Switchyard repository provides a reference implementation for integrating Switchyard's advanced LLM routing capabilities with LiteLLM through the switchyard_litellm library. This integration allows you to leverage Switchyard's deployment selection algorithms while maintaining LiteLLM's unified API interface. By developing a LiteLLM routing plugin for Switchyard, you can dynamically route requests to optimal model deployments based on custom policies, latency requirements, or cost constraints.

Architecture of the Switchyard LiteLLM Routing Plugin

The integration follows a layered architecture where LiteLLM's request lifecycle hooks into Switchyard's routing core. LiteLLM forwards incoming chat completion requests to the CustomLogger._request method, which the SwitchyardRoutingPlugin implements to call Switchyard's routing logic. After selecting a deployment, the plugin optionally modifies the request using the build_request_patch utility before returning the enriched payload to LiteLLM for final transmission to the chosen model endpoint.

Key files in the examples/litellm/ directory implement this flow:

Implementing the Core Plugin Components

Extending CustomLogger in switchyard_routing_plugin.py

The plugin class must inherit from litellm.integrations.custom_logger.CustomLogger and implement the async _request method. This method receives the LiteLLM-converted request and is responsible for invoking Switchyard's routing logic.


# examples/litellm/plugins/switchyard_routing_plugin.py

from __future__ import annotations
import asyncio
from typing import Any
from litellm.integrations.custom_logger import CustomLogger

class SwitchyardRoutingPlugin(CustomLogger):
    async def _request(self, litellm_messages: Any) -> Any:
        """
        Intercept the LiteLLM request, route through Switchyard, and return patched result.
        """
        # Invoke Switchyard core routing (implementation-specific call)

        routing_result = await self._select_deployment(litellm_messages)
        
        # Apply request modifications using the rewrite utility

        from .request_rewrite import build_request_patch
        patch = build_request_patch(
            litellm_messages,
            model=routing_result["model"],
            stage="prod",
        )
        
        return {**litellm_messages, **patch}
    
    async def _select_deployment(self, messages: Any) -> dict:
        """Placeholder for Switchyard routing logic."""
        # Integrate with switchyard_py bindings here

        return {"model": "gpt-4o"}

Rewriting Requests with build_request_patch

The build_request_patch function in examples/litellm/plugins/request_rewrite.py constructs partial request dictionaries that LiteLLM merges with the original request. This preserves critical fields like the messages list while injecting routing-specific metadata.


# examples/litellm/plugins/request_rewrite.py

from typing import List, Dict, Any

def build_request_patch(
    messages: List[Dict[str, Any]],
    *,
    model: str,
    stage: str,
    **kwargs,
) -> Dict[str, Any]:
    """
    Build a partial request patch preserving fields through LiteLLM's conversion.
    
    Parameters:
        messages: List of chat messages from the original request.
        model: Target model identifier selected by Switchyard.
        stage: Deployment stage (e.g., 'prod', 'dev').
        **kwargs: Additional metadata to include.
        
    Returns:
        Dictionary containing fields to merge into the final request.
    """
    return {
        "model": model,
        "metadata": {"stage": stage, **kwargs},
    }

Plugin Factory in configuration.py

LiteLLM instantiates the plugin through a factory function defined in examples/litellm/configuration.py. This pattern allows dynamic configuration loading from YAML files.


# examples/litellm/configuration.py

from typing import Dict, Any
from switchyard_litellm.plugins.switchyard_routing_plugin import SwitchyardRoutingPlugin

def load_routing_plugin(config: Dict[str, Any]) -> SwitchyardRoutingPlugin:
    """
    Factory used by LiteLLM to create the routing plugin instance.
    
    Parameters:
        config: Configuration dictionary from the LiteLLM router config.
        
    Returns:
        Configured SwitchyardRoutingPlugin instance.
    """
    return SwitchyardRoutingPlugin()

Registering and Testing Your Plugin

Registration with LiteLLM

To activate the plugin, reference the factory in your LiteLLM router configuration:


# router_config.yaml

custom_logger: switchyard_litellm.configuration.load_routing_plugin

LiteLLM calls this factory during initialization and passes the _request hook for every chat completion request.

Unit Testing the Routing Logic

The examples/litellm/tests/unit/test_switchyard_routing_plugin.py file validates that the plugin correctly preserves LiteLLM fields and applies routing decisions.


# examples/litellm/tests/unit/test_switchyard_routing_plugin.py

import pytest
from litellm.types.router import RoutingContext
from switchyard_litellm import SwitchyardRoutingPlugin
from switchyard_litellm.plugins.request_rewrite import build_request_patch

@pytest.mark.asyncio
async def test_request_routing():
    plugin = SwitchyardRoutingPlugin()
    context = RoutingContext(
        messages=[{"role": "user", "content": "Hello"}],
        model="gpt-4"
    )
    
    result = await plugin._request(context)
    
    assert "model" in result
    assert result["metadata"]["stage"] == "prod"

def test_build_request_patch():
    messages = [{"role": "user", "content": "test"}]
    patch = build_request_patch(messages, model="claude-3", stage="dev", custom_key="value")
    
    assert patch["model"] == "claude-3"
    assert patch["metadata"]["stage"] == "dev"
    assert patch["metadata"]["custom_key"] == "value"

Deployment Configuration Profiles

Switchyard uses YAML profiles to define routing strategies. The example includes two reference profiles:

Place your custom profiles in this directory structure and reference them in your Switchyard runner configuration to control how the plugin selects between model deployments.

Summary

  • Implement a CustomLogger subclass in switchyard_routing_plugin.py with an async _request method to intercept LiteLLM traffic before it reaches the model provider.
  • Use build_request_patch from request_rewrite.py to safely inject routing decisions and metadata without losing LiteLLM's converted message fields.
  • Expose a load_routing_plugin factory in configuration.py so LiteLLM can instantiate your plugin from YAML configuration using the custom_logger key.
  • Test your implementation using the patterns in test_switchyard_routing_plugin.py, verifying that RoutingContext fields propagate correctly through the _request lifecycle.
  • Configure deployment profiles in examples/litellm/deployment/profiles/ to define selection algorithms, cost constraints, and endpoint priorities for the routing plugin.

Frequently Asked Questions

What is the relationship between Switchyard and LiteLLM in this plugin architecture?

Switchyard provides the core routing algorithms and deployment selection logic written in Rust, while LiteLLM offers a unified API gateway to multiple LLM providers. The plugin acts as a bridge, using LiteLLM's CustomLogger hook to call Switchyard's Python bindings for routing decisions before the request reaches the final model endpoint.

How does the build_request_patch function preserve request fields?

The build_request_patch function in examples/litellm/plugins/request_rewrite.py returns a partial dictionary containing only the fields to update, such as the selected model identifier and stage metadata. By returning a patch rather than a complete request object, the function ensures that LiteLLM's internal message conversion preserves the original messages list and other critical fields while merging in the routing-specific updates.

Can I use synchronous Switchyard calls within the _request hook?

While the _request hook is defined as async to comply with LiteLLM's CustomLogger interface, Switchyard's Python bindings may expose synchronous Rust functions. Use asyncio.to_thread() to wrap synchronous Switchyard calls, as shown in the switchyard_routing_plugin.py pattern, to prevent blocking the LiteLLM event loop during routing decisions.

Where should I place custom deployment profiles for the plugin?

Place YAML configuration files in the examples/litellm/deployment/profiles/ directory, following the structure of the existing stage and random profiles. The Switchyard runner loads these profiles at initialization to determine selection algorithms, cost constraints, and endpoint priorities for the routing plugin.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →