# How Switchyard’s LiteLLM Integration Connects Routing to LiteLLM’s Router/Proxy

> Discover how Switchyard's LiteLLM integration connects routing to LiteLLM's Router/proxy. Learn how plugins select optimal models and rewrite requests for seamless AI routing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-13

---

**The Switchyard LiteLLM integration works by chaining two plugins: a routing plugin that runs Switchyard’s algorithm to select the optimal model, and a request rewriter plugin that applies Switchyard’s decision to LiteLLM’s kwargs via a patch mechanism.**

The **NVIDIA-NeMo/Switchyard** repository provides a reference implementation in the `examples/litellm` directory that bridges Switchyard’s intelligent routing algorithms with LiteLLM’s production proxy infrastructure. This integration allows you to leverage Switchyard’s stage-based routing logic while retaining LiteLLM’s deployment management, load balancing, and observability features.

## Architecture of the Switchyard LiteLLM Integration

The integration relies on a dual-plugin architecture that separates **decision-making** from **request mutation**. This design ensures that Switchyard’s routing logic runs as a first-class citizen within LiteLLM’s request lifecycle without breaking existing proxy functionality.

### The Two-Plugin System

**SwitchyardRoutingPlugin** ([`switchyard_litellm/plugins/switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_litellm/plugins/switchyard_routing_plugin.py)) serves as the brain of the operation. It converts LiteLLM’s structured messages into a canonical Switchyard request, streams the algorithm execution via `Algorithm.run_stream` (line 101), and generates a minimal diff called a *request patch* using `build_request_patch` (lines 20‑33). This patch captures exactly which fields must change to reflect the routing decision.

**LiteLLMRequestRewriter** ([`switchyard_litellm/plugins/lite_llm_request_rewriter.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_litellm/plugins/lite_llm_request_rewriter.py)) implements LiteLLM’s `CustomLogger` interface and specifically hooks into `async_pre_call_deployment_hook` (lines 52‑94). After LiteLLM selects a deployment candidate but before the HTTP call executes, this hook inspects `context.signals["switchyard"]["request_patch"]` and applies the "set" and "remove" operations toLiteLLM’s kwargs.

### Data Flow Through the System

According to the Switchyard source code, a request flows through five distinct stages:

1. **Context Creation**: LiteLLM’s `Router` instantiates a `RoutingContext` containing `candidate_models` and the original message list.
2. **Algorithm Execution**: `StageRoutingPlugin` (or a custom `SwitchyardRoutingPlugin`) invokes the Switchyard algorithm, which evaluates candidates using metrics like efficiency or capability scores.
3. **Patch Generation**: Upon completion, the plugin creates a request patch describing field modifications (e.g., changing the model name or updating message content) and stores it in `context.signals["switchyard"]`.
4. **Hook Execution**: LiteLLM’s logger mechanism triggers `LiteLLMRequestRewriter.async_pre_call_deployment_hook`, which retrieves the patch, validates allowed fields against a whitelist, and mutates the kwargs dictionary.
5. **Final Dispatch**: LiteLLM executes the HTTP call against the selected deployment using the rewritten parameters.

## Core Components in Detail

### StageRoutingPlugin: The Convenience Wrapper

For deployments using the Stage routing algorithm, [`switchyard_litellm/plugins/stage_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_litellm/plugins/stage_routing_plugin.py) provides a high-level abstraction. It maps LiteLLM’s two-candidate model order (typically `any`, `capable`, `efficient`) to Switchyard’s stage router configuration.

The plugin constructs a `SwitchyardRoutingPlugin` internally (lines 55‑64) with a fixed model mapping:

```python

# switchyard_litellm/plugins/stage_routing_plugin.py

plugin = SwitchyardRoutingPlugin(
    algorithms.stage_router(
        picker=self._picker,
        confidence_threshold=self._confidence_threshold,
        recent_window=self._recent_window,
    ),
    models={
        "any": candidates,
        "capable": [candidates[0]],
        "efficient": [candidates[1]],
    },
)
return await plugin.run(context)

```

### SwitchyardRoutingPlugin: Building the Request Patch

This plugin handles the heavy lifting of translating between LiteLLM’s message format and Switchyard’s canonical representation. In [`switchyard_litellm/plugins/switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_litellm/plugins/switchyard_routing_plugin.py), the `_request` helper (lines 37‑51) normalizes the input before streaming the algorithm:

```python

# switchyard_litellm/plugins/switchyard_routing_plugin.py

request = _request(context.structured_messages)

async for step in self._algorithm.run_stream(request, models):
    match step:
        case Step.Done(outcome):
            request_patch = build_request_patch(
                request,
                outcome.request,
                selected_model_id=selected,
            )
            context.signals["switchyard"]["request_patch"] = request_patch

```

The resulting patch is a dictionary with `"set"` and `"remove"` keys that describe minimal transformations.

### LiteLLMRequestRewriter: Applying the Patch

The rewriter operates as a LiteLLM callback. In [`switchyard_litellm/plugins/lite_llm_request_rewriter.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_litellm/plugins/lite_llm_request_rewriter.py), lines 82‑94 demonstrate how the patch merges into the final kwargs:

```python

# switchyard_litellm/plugins/lite_llm_request_rewriter.py

patch = switchyard_signal.get("request_patch")

# Filter disallowed fields

allowed_patch = {
    k: v for k, v in patch.items() 
    if k in self._allowed_fields
}

rewritten = dict(kwargs)
for field in allowed_patch.get("remove", []):
    rewritten.pop(field, None)
rewritten.update(allowed_patch.get("set", {}))

```

This ensures that only Switchyard-authorized modifications (such as model selection or parameter tuning) reach the upstream provider.

## Implementing the Integration

### Setting Up the Router

To use the integration, instantiate a LiteLLM `Router` with the `StageRoutingPlugin` and register it as both a plugin and a callback, as shown in [`examples/litellm/examples/python_router.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/examples/python_router.py):

```python

# examples/litellm/examples/python_router.py

import asyncio
import litellm
from litellm.router import Router
from switchyard_litellm import StageRoutingPlugin

MODEL_GROUP = "switchyard"
MODEL_LIST = [
    {
        "model_name": MODEL_GROUP,
        "litellm_params": {"model": "openrouter/openai/gpt-5.6-sol"},
    },
    {
        "model_name": MODEL_GROUP,
        "litellm_params": {"model": "openrouter/openai/gpt-5.6-terra"},
    },
]

STAGE_ROUTING_PLUGIN = StageRoutingPlugin(
    picker="efficient_first",
    confidence_threshold=0.5,
    recent_window=3,
)

async def main() -> None:
    router = Router(model_list=MODEL_LIST, plugins=[STAGE_ROUTING_PLUGIN])
    litellm.callbacks.append(STAGE_ROUTING_PLUGIN)
    
    response = await router.acompletion(
        model=MODEL_GROUP,
        messages=[{"role": "user", "content": "What is the capital of France?"}],
        max_tokens=64,
    )
    print("Selected model:", response.model)
    print("Response:", response.choices[0].message.content)

if __name__ == "__main__":
    asyncio.run(main())

```

### Configuration Files

Algorithm parameters are defined in TOML configuration files. The Stage router profile at [`examples/litellm/deployment/profiles/stage/switchyard.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/deployment/profiles/stage/switchyard.toml) specifies thresholds and window sizes that the `StageRoutingPlugin` consumes during initialization.

## Summary

- **Dual-plugin architecture**: The Switchyard LiteLLM integration uses `SwitchyardRoutingPlugin` to make routing decisions and `LiteLLMRequestRewriter` to apply them.
- **Patch-based mutation**: Instead of regenerating the entire request, Switchyard creates a minimal patch stored in `context.signals["switchyard"]["request_patch"]` that the rewriter applies just before the HTTP call.
- **Stage algorithm support**: `StageRoutingPlugin` provides a convenient wrapper that maps LiteLLM’s candidate pairs to Switchyard’s stage router with configurable pickers like `efficient_first`.
- **File locations**: Key implementation files reside in `examples/litellm/src/switchyard_litellm/plugins/` with entry-point examples in `examples/litellm/examples/`.

## Frequently Asked Questions

### How does Switchyard modify LiteLLM requests without breaking the proxy chain?

Switchyard uses LiteLLM’s native `CustomLogger` callback interface, specifically the `async_pre_call_deployment_hook`. This hook runs after LiteLLM has selected a deployment but before the network call, allowing Switchyard to rewrite kwargs safely within LiteLLM’s expected lifecycle. The `LiteLLMRequestRewriter` validates changes against an allow-list to prevent corruption.

### What fields can Switchyard modify in a LiteLLM request?

According to the source code in [`lite_llm_request_rewriter.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/lite_llm_request_rewriter.py), the rewriter filters the request patch against `self._allowed_fields` before applying changes. Typically, this includes the model identifier, message content, and sampling parameters like temperature or max_tokens. Disallowed fields are stripped from the patch to maintain security.

### Can I use custom Switchyard algorithms beyond the Stage router?

Yes. While `StageRoutingPlugin` provides a convenient wrapper for the Stage algorithm, you can instantiate `SwitchyardRoutingPlugin` directly with any Switchyard algorithm implementation. Pass your custom algorithm to the plugin’s constructor along with a model mapping dictionary, then register the plugin with the LiteLLM Router.

### Where is the routing decision actually made in the codebase?

The routing decision occurs in [`switchyard_litellm/plugins/switchyard_routing_plugin.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_litellm/plugins/switchyard_routing_plugin.py) at line 101, where `self._algorithm.run_stream(request, models)` executes. The algorithm evaluates candidates based on the configured strategy (e.g., efficiency-first) and returns an outcome that the plugin converts into a request patch.