How to Migrate from LlmTarget API to Step/LlmResponse API in Switchyard

To migrate from the legacy LlmTarget API to the modern Step/LlmResponse API, update your imports to expose Step and LlmResponse from switchyard.libsy, replace direct response handling with an async loop over run_stream that yields Step objects, and wrap your model outputs in LlmResponse.Agg for aggregated responses or LlmResponse.Stream for streaming chunks.

Switchyard, NVIDIA's LLM routing library, redesigned its Python bindings in version 0.2.0 to support async-first architectures. The monolithic LlmTarget type—which previously doubled as both request descriptor and response carrier—has been deprecated in favor of a cleaner, event-driven model using the Step enum and LlmResponse discriminated union.

Understanding the API Changes

The migration from LlmTarget to Step and LlmResponse fundamentally changes how you interact with the routing algorithm:

  • Legacy pattern: The LlmTarget object acted as a mutable container that you passed into the router and then inspected for the final model selection and response.
  • Modern pattern: The run_stream method yields Step enum variants (such as Step.CallModel and Step.Done) that explicitly signal each phase of the routing lifecycle. You provide responses back to the router by constructing LlmResponse instances rather than mutating a shared object.

This shift enables proper backpressure handling and native support for streaming responses without blocking the event loop.

Step-by-Step Migration Guide

Update Your Imports

Replace any imports of the deprecated LlmTarget with the new public symbols. In switchyard/libsy/__init__.py, the 0.2.0 release exposes Step and LlmResponse as top-level exports:


# Old (removed in 0.2.0)

# from switchyard.libsy import LlmTarget

# New

from switchyard.libsy import Step, LlmResponse, LlmClassifierConfig
from switchyard.libsy.algorithms import stage_router

Construct Protocol-Compliant Requests

The run_stream entry point expects a normalized request dictionary defined in the protocol crate. Consult crates/protocol/src/llm.rs for the exact schema, which typically includes a messages array and routing metadata:

request = {
    "messages": [{"role": "user", "content": [{"type": "text", "text": "Explain quantum computing"}]}],
    "model": "any",  # Router overrides this based on its selection logic

}

Handle Step Objects in the Event Loop

Algorithms like stage_router now return an async iterator of Step objects rather than a final LlmTarget. Drive the router by iterating over run_stream and pattern-matching on each step. The test suite in tests/test_libsy_minimal_bindings.py demonstrates the canonical handling:

async for step in algorithm.run_stream(request):
    match step:
        case Step.CallModel(call):
            # Router requests a model inference

            raw = await my_client.call(call.request)
            await call.respond(LlmResponse.Agg(raw))
        case Step.Done(outcome):
            # Final routing decision complete

            print(f"Selected: {outcome.model_id}")

Common Step variants include:

  • Step.CallModel(call) – The router requires a classifier or judge inference before proceeding.
  • Step.Done(outcome) – The algorithm has finalized its model selection and contains the aggregated response.

Wrap Model Responses with LlmResponse

When the router yields Step.CallModel, you must wrap your client's output in an LlmResponse variant before calling respond(). The class definitions reside in switchyard_rust/libsy.py:

  • LlmResponse.Agg(data) – For complete, non-streaming JSON responses from your model client.
  • LlmResponse.Stream(generator) – For async generators yielding streaming chunks.

The examples/libsy.py file provides a complete reference for implementing both patterns.

Complete Migration Example

Below is a before-and-after comparison showing the transition from the imperative LlmTarget style to the async Step/LlmResponse approach:


# ==========================================

# Legacy API (pre-0.2.0) - DEPRECATED

# ==========================================

# from switchyard.libsy import LlmTarget

# 
# target = LlmTarget(messages=[...], model="gpt-4")

# result = router.route(target)  # Blocking, mutates target in-place

# print(result.model_id)

# ==========================================

# Modern API (0.2.0+) - Async Step Model

# ==========================================

from switchyard.libsy import Step, LlmResponse
from switchyard.libsy.algorithms import stage_router

async def route_request(user_message: str):
    algorithm = stage_router(picker="efficient_first", confidence_threshold=0.5)
    request = {
        "messages": [{"role": "user", "content": user_message}],
        "model": "any",
    }
    
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                # Integrate your own client

                response_data = await my_llm_client.complete(call.request)
                await call.respond(LlmResponse.Agg(response_data))
            case Step.Done(outcome):
                return outcome.model_id, outcome.response

Key Source Files and References

Understanding the implementation details in the source code will help debug migration issues:

  • switchyard/libsy/__init__.py – Defines the public Python API, re-exporting Step, LlmResponse, and configuration classes.
  • switchyard_rust/libsy.py – Contains the Python bindings for the Step enum and LlmResponse discriminated union (lines 44–52).
  • crates/protocol/src/llm.rs – Specifies the normalized request schema that run_stream consumes.
  • tests/test_libsy_minimal_bindings.py – Demonstrates pattern-matching against Step variants in a test context (lines 70–85).
  • examples/libsy.py – Provides a runnable example of the full Step-driven workflow including both CallModel and Done handling.
  • README.md – Documents the high-level migration rationale and checklist (sections around lines 53–56 and 78–86).

Summary

  • Import changes: Replace LlmTarget with Step and LlmResponse from switchyard.libsy.
  • Execution model: Use run_stream to drive the router, yielding Step objects instead of returning a target directly.
  • Response handling: Wrap model outputs in LlmResponse.Agg or LlmResponse.Stream when responding to Step.CallModel.
  • Protocol compliance: Ensure request dictionaries match the schema defined in crates/protocol/src/llm.rs.
  • Async requirement: The new API is async-first; wrap synchronous clients with asyncio or use native async HTTP clients.

Frequently Asked Questions

What happened to LlmTarget in Switchyard 0.2.0?

The LlmTarget class was removed in version 0.2.0 because it tightly coupled request description, routing state, and response storage into a single mutable object. According to the Switchyard source code, this design prevented clean async streaming and made it difficult to implement backpressure. The new Step/LlmResponse architecture separates these concerns, allowing the router to yield control back to your code at each decision point.

How do I handle streaming responses with the new API?

When you receive a Step.CallModel, wrap your streaming generator in LlmResponse.Stream instead of LlmResponse.Agg. Pass the generator to call.respond(); the router will consume it asynchronously. This pattern is defined in switchyard_rust/libsy.py and demonstrated in the main repository examples.

Do I need to rewrite my routing algorithms?

No. Algorithms such as stage_router maintain the same construction API and configuration options. They now emit Step objects through their run_stream method rather than returning an LlmTarget directly. Update your calling code to iterate over the stream, but the algorithm initialization logic remains unchanged.

Where is LlmResponse defined in the codebase?

The LlmResponse class is defined in switchyard_rust/libsy.py as a discriminated union with two variants: Agg for aggregated (complete) responses and Stream for streaming iterators. This file bridges the Rust core implementation to Python, exposing the types that you import from switchyard.libsy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →