How to Migrate from LlmTarget API to Step/LlmResponse API in Switchyard
To migrate from the legacy LlmTarget API to the modern Step/LlmResponse API, update your imports to expose Step and LlmResponse from switchyard.libsy, replace direct response handling with an async loop over run_stream that yields Step objects, and wrap your model outputs in LlmResponse.Agg for aggregated responses or LlmResponse.Stream for streaming chunks.
Switchyard, NVIDIA's LLM routing library, redesigned its Python bindings in version 0.2.0 to support async-first architectures. The monolithic LlmTarget type—which previously doubled as both request descriptor and response carrier—has been deprecated in favor of a cleaner, event-driven model using the Step enum and LlmResponse discriminated union.
Understanding the API Changes
The migration from LlmTarget to Step and LlmResponse fundamentally changes how you interact with the routing algorithm:
- Legacy pattern: The
LlmTargetobject acted as a mutable container that you passed into the router and then inspected for the final model selection and response. - Modern pattern: The
run_streammethod yieldsStepenum variants (such asStep.CallModelandStep.Done) that explicitly signal each phase of the routing lifecycle. You provide responses back to the router by constructingLlmResponseinstances rather than mutating a shared object.
This shift enables proper backpressure handling and native support for streaming responses without blocking the event loop.
Step-by-Step Migration Guide
Update Your Imports
Replace any imports of the deprecated LlmTarget with the new public symbols. In switchyard/libsy/__init__.py, the 0.2.0 release exposes Step and LlmResponse as top-level exports:
# Old (removed in 0.2.0)
# from switchyard.libsy import LlmTarget
# New
from switchyard.libsy import Step, LlmResponse, LlmClassifierConfig
from switchyard.libsy.algorithms import stage_router
Construct Protocol-Compliant Requests
The run_stream entry point expects a normalized request dictionary defined in the protocol crate. Consult crates/protocol/src/llm.rs for the exact schema, which typically includes a messages array and routing metadata:
request = {
"messages": [{"role": "user", "content": [{"type": "text", "text": "Explain quantum computing"}]}],
"model": "any", # Router overrides this based on its selection logic
}
Handle Step Objects in the Event Loop
Algorithms like stage_router now return an async iterator of Step objects rather than a final LlmTarget. Drive the router by iterating over run_stream and pattern-matching on each step. The test suite in tests/test_libsy_minimal_bindings.py demonstrates the canonical handling:
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
# Router requests a model inference
raw = await my_client.call(call.request)
await call.respond(LlmResponse.Agg(raw))
case Step.Done(outcome):
# Final routing decision complete
print(f"Selected: {outcome.model_id}")
Common Step variants include:
Step.CallModel(call)– The router requires a classifier or judge inference before proceeding.Step.Done(outcome)– The algorithm has finalized its model selection and contains the aggregated response.
Wrap Model Responses with LlmResponse
When the router yields Step.CallModel, you must wrap your client's output in an LlmResponse variant before calling respond(). The class definitions reside in switchyard_rust/libsy.py:
LlmResponse.Agg(data)– For complete, non-streaming JSON responses from your model client.LlmResponse.Stream(generator)– For async generators yielding streaming chunks.
The examples/libsy.py file provides a complete reference for implementing both patterns.
Complete Migration Example
Below is a before-and-after comparison showing the transition from the imperative LlmTarget style to the async Step/LlmResponse approach:
# ==========================================
# Legacy API (pre-0.2.0) - DEPRECATED
# ==========================================
# from switchyard.libsy import LlmTarget
#
# target = LlmTarget(messages=[...], model="gpt-4")
# result = router.route(target) # Blocking, mutates target in-place
# print(result.model_id)
# ==========================================
# Modern API (0.2.0+) - Async Step Model
# ==========================================
from switchyard.libsy import Step, LlmResponse
from switchyard.libsy.algorithms import stage_router
async def route_request(user_message: str):
algorithm = stage_router(picker="efficient_first", confidence_threshold=0.5)
request = {
"messages": [{"role": "user", "content": user_message}],
"model": "any",
}
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
# Integrate your own client
response_data = await my_llm_client.complete(call.request)
await call.respond(LlmResponse.Agg(response_data))
case Step.Done(outcome):
return outcome.model_id, outcome.response
Key Source Files and References
Understanding the implementation details in the source code will help debug migration issues:
switchyard/libsy/__init__.py– Defines the public Python API, re-exportingStep,LlmResponse, and configuration classes.switchyard_rust/libsy.py– Contains the Python bindings for theStepenum andLlmResponsediscriminated union (lines 44–52).crates/protocol/src/llm.rs– Specifies the normalized request schema thatrun_streamconsumes.tests/test_libsy_minimal_bindings.py– Demonstrates pattern-matching againstStepvariants in a test context (lines 70–85).examples/libsy.py– Provides a runnable example of the fullStep-driven workflow including bothCallModelandDonehandling.README.md– Documents the high-level migration rationale and checklist (sections around lines 53–56 and 78–86).
Summary
- Import changes: Replace
LlmTargetwithStepandLlmResponsefromswitchyard.libsy. - Execution model: Use
run_streamto drive the router, yieldingStepobjects instead of returning a target directly. - Response handling: Wrap model outputs in
LlmResponse.AggorLlmResponse.Streamwhen responding toStep.CallModel. - Protocol compliance: Ensure request dictionaries match the schema defined in
crates/protocol/src/llm.rs. - Async requirement: The new API is async-first; wrap synchronous clients with
asyncioor use native async HTTP clients.
Frequently Asked Questions
What happened to LlmTarget in Switchyard 0.2.0?
The LlmTarget class was removed in version 0.2.0 because it tightly coupled request description, routing state, and response storage into a single mutable object. According to the Switchyard source code, this design prevented clean async streaming and made it difficult to implement backpressure. The new Step/LlmResponse architecture separates these concerns, allowing the router to yield control back to your code at each decision point.
How do I handle streaming responses with the new API?
When you receive a Step.CallModel, wrap your streaming generator in LlmResponse.Stream instead of LlmResponse.Agg. Pass the generator to call.respond(); the router will consume it asynchronously. This pattern is defined in switchyard_rust/libsy.py and demonstrated in the main repository examples.
Do I need to rewrite my routing algorithms?
No. Algorithms such as stage_router maintain the same construction API and configuration options. They now emit Step objects through their run_stream method rather than returning an LlmTarget directly. Update your calling code to iterate over the stream, but the algorithm initialization logic remains unchanged.
Where is LlmResponse defined in the codebase?
The LlmResponse class is defined in switchyard_rust/libsy.py as a discriminated union with two variants: Agg for aggregated (complete) responses and Stream for streaming iterators. This file bridges the Rust core implementation to Python, exposing the types that you import from switchyard.libsy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →