How Switchyard’s LiteLLM Integration Connects Routing to LiteLLM’s Router/Proxy
The Switchyard LiteLLM integration works by chaining two plugins: a routing plugin that runs Switchyard’s algorithm to select the optimal model, and a request rewriter plugin that applies Switchyard’s decision to LiteLLM’s kwargs via a patch mechanism.
The NVIDIA-NeMo/Switchyard repository provides a reference implementation in the examples/litellm directory that bridges Switchyard’s intelligent routing algorithms with LiteLLM’s production proxy infrastructure. This integration allows you to leverage Switchyard’s stage-based routing logic while retaining LiteLLM’s deployment management, load balancing, and observability features.
Architecture of the Switchyard LiteLLM Integration
The integration relies on a dual-plugin architecture that separates decision-making from request mutation. This design ensures that Switchyard’s routing logic runs as a first-class citizen within LiteLLM’s request lifecycle without breaking existing proxy functionality.
The Two-Plugin System
SwitchyardRoutingPlugin (switchyard_litellm/plugins/switchyard_routing_plugin.py) serves as the brain of the operation. It converts LiteLLM’s structured messages into a canonical Switchyard request, streams the algorithm execution via Algorithm.run_stream (line 101), and generates a minimal diff called a request patch using build_request_patch (lines 20‑33). This patch captures exactly which fields must change to reflect the routing decision.
LiteLLMRequestRewriter (switchyard_litellm/plugins/lite_llm_request_rewriter.py) implements LiteLLM’s CustomLogger interface and specifically hooks into async_pre_call_deployment_hook (lines 52‑94). After LiteLLM selects a deployment candidate but before the HTTP call executes, this hook inspects context.signals["switchyard"]["request_patch"] and applies the "set" and "remove" operations toLiteLLM’s kwargs.
Data Flow Through the System
According to the Switchyard source code, a request flows through five distinct stages:
- Context Creation: LiteLLM’s
Routerinstantiates aRoutingContextcontainingcandidate_modelsand the original message list. - Algorithm Execution:
StageRoutingPlugin(or a customSwitchyardRoutingPlugin) invokes the Switchyard algorithm, which evaluates candidates using metrics like efficiency or capability scores. - Patch Generation: Upon completion, the plugin creates a request patch describing field modifications (e.g., changing the model name or updating message content) and stores it in
context.signals["switchyard"]. - Hook Execution: LiteLLM’s logger mechanism triggers
LiteLLMRequestRewriter.async_pre_call_deployment_hook, which retrieves the patch, validates allowed fields against a whitelist, and mutates the kwargs dictionary. - Final Dispatch: LiteLLM executes the HTTP call against the selected deployment using the rewritten parameters.
Core Components in Detail
StageRoutingPlugin: The Convenience Wrapper
For deployments using the Stage routing algorithm, switchyard_litellm/plugins/stage_routing_plugin.py provides a high-level abstraction. It maps LiteLLM’s two-candidate model order (typically any, capable, efficient) to Switchyard’s stage router configuration.
The plugin constructs a SwitchyardRoutingPlugin internally (lines 55‑64) with a fixed model mapping:
# switchyard_litellm/plugins/stage_routing_plugin.py
plugin = SwitchyardRoutingPlugin(
algorithms.stage_router(
picker=self._picker,
confidence_threshold=self._confidence_threshold,
recent_window=self._recent_window,
),
models={
"any": candidates,
"capable": [candidates[0]],
"efficient": [candidates[1]],
},
)
return await plugin.run(context)
SwitchyardRoutingPlugin: Building the Request Patch
This plugin handles the heavy lifting of translating between LiteLLM’s message format and Switchyard’s canonical representation. In switchyard_litellm/plugins/switchyard_routing_plugin.py, the _request helper (lines 37‑51) normalizes the input before streaming the algorithm:
# switchyard_litellm/plugins/switchyard_routing_plugin.py
request = _request(context.structured_messages)
async for step in self._algorithm.run_stream(request, models):
match step:
case Step.Done(outcome):
request_patch = build_request_patch(
request,
outcome.request,
selected_model_id=selected,
)
context.signals["switchyard"]["request_patch"] = request_patch
The resulting patch is a dictionary with "set" and "remove" keys that describe minimal transformations.
LiteLLMRequestRewriter: Applying the Patch
The rewriter operates as a LiteLLM callback. In switchyard_litellm/plugins/lite_llm_request_rewriter.py, lines 82‑94 demonstrate how the patch merges into the final kwargs:
# switchyard_litellm/plugins/lite_llm_request_rewriter.py
patch = switchyard_signal.get("request_patch")
# Filter disallowed fields
allowed_patch = {
k: v for k, v in patch.items()
if k in self._allowed_fields
}
rewritten = dict(kwargs)
for field in allowed_patch.get("remove", []):
rewritten.pop(field, None)
rewritten.update(allowed_patch.get("set", {}))
This ensures that only Switchyard-authorized modifications (such as model selection or parameter tuning) reach the upstream provider.
Implementing the Integration
Setting Up the Router
To use the integration, instantiate a LiteLLM Router with the StageRoutingPlugin and register it as both a plugin and a callback, as shown in examples/litellm/examples/python_router.py:
# examples/litellm/examples/python_router.py
import asyncio
import litellm
from litellm.router import Router
from switchyard_litellm import StageRoutingPlugin
MODEL_GROUP = "switchyard"
MODEL_LIST = [
{
"model_name": MODEL_GROUP,
"litellm_params": {"model": "openrouter/openai/gpt-5.6-sol"},
},
{
"model_name": MODEL_GROUP,
"litellm_params": {"model": "openrouter/openai/gpt-5.6-terra"},
},
]
STAGE_ROUTING_PLUGIN = StageRoutingPlugin(
picker="efficient_first",
confidence_threshold=0.5,
recent_window=3,
)
async def main() -> None:
router = Router(model_list=MODEL_LIST, plugins=[STAGE_ROUTING_PLUGIN])
litellm.callbacks.append(STAGE_ROUTING_PLUGIN)
response = await router.acompletion(
model=MODEL_GROUP,
messages=[{"role": "user", "content": "What is the capital of France?"}],
max_tokens=64,
)
print("Selected model:", response.model)
print("Response:", response.choices[0].message.content)
if __name__ == "__main__":
asyncio.run(main())
Configuration Files
Algorithm parameters are defined in TOML configuration files. The Stage router profile at examples/litellm/deployment/profiles/stage/switchyard.toml specifies thresholds and window sizes that the StageRoutingPlugin consumes during initialization.
Summary
- Dual-plugin architecture: The Switchyard LiteLLM integration uses
SwitchyardRoutingPluginto make routing decisions andLiteLLMRequestRewriterto apply them. - Patch-based mutation: Instead of regenerating the entire request, Switchyard creates a minimal patch stored in
context.signals["switchyard"]["request_patch"]that the rewriter applies just before the HTTP call. - Stage algorithm support:
StageRoutingPluginprovides a convenient wrapper that maps LiteLLM’s candidate pairs to Switchyard’s stage router with configurable pickers likeefficient_first. - File locations: Key implementation files reside in
examples/litellm/src/switchyard_litellm/plugins/with entry-point examples inexamples/litellm/examples/.
Frequently Asked Questions
How does Switchyard modify LiteLLM requests without breaking the proxy chain?
Switchyard uses LiteLLM’s native CustomLogger callback interface, specifically the async_pre_call_deployment_hook. This hook runs after LiteLLM has selected a deployment but before the network call, allowing Switchyard to rewrite kwargs safely within LiteLLM’s expected lifecycle. The LiteLLMRequestRewriter validates changes against an allow-list to prevent corruption.
What fields can Switchyard modify in a LiteLLM request?
According to the source code in lite_llm_request_rewriter.py, the rewriter filters the request patch against self._allowed_fields before applying changes. Typically, this includes the model identifier, message content, and sampling parameters like temperature or max_tokens. Disallowed fields are stripped from the patch to maintain security.
Can I use custom Switchyard algorithms beyond the Stage router?
Yes. While StageRoutingPlugin provides a convenient wrapper for the Stage algorithm, you can instantiate SwitchyardRoutingPlugin directly with any Switchyard algorithm implementation. Pass your custom algorithm to the plugin’s constructor along with a model mapping dictionary, then register the plugin with the LiteLLM Router.
Where is the routing decision actually made in the codebase?
The routing decision occurs in switchyard_litellm/plugins/switchyard_routing_plugin.py at line 101, where self._algorithm.run_stream(request, models) executes. The algorithm evaluates candidates based on the configured strategy (e.g., efficiency-first) and returns an outcome that the plugin converts into a request patch.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →