# How to Migrate from LlmTarget API to Step/LlmResponse API in Switchyard

> Migrate from LlmTarget API to Switchyard's Step/LlmResponse API. Update imports, use an async loop for run_stream, and wrap outputs for aggregated or streamed responses. Learn the upgrade process.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: migration-guide
- Published: 2026-09-12

---

**To migrate from the legacy `LlmTarget` API to the modern `Step`/`LlmResponse` API, update your imports to expose `Step` and `LlmResponse` from `switchyard.libsy`, replace direct response handling with an async loop over `run_stream` that yields `Step` objects, and wrap your model outputs in `LlmResponse.Agg` for aggregated responses or `LlmResponse.Stream` for streaming chunks.**

Switchyard, NVIDIA's LLM routing library, redesigned its Python bindings in version `0.2.0` to support async-first architectures. The monolithic `LlmTarget` type—which previously doubled as both request descriptor and response carrier—has been deprecated in favor of a cleaner, event-driven model using the `Step` enum and `LlmResponse` discriminated union.

## Understanding the API Changes

The migration from `LlmTarget` to `Step` and `LlmResponse` fundamentally changes how you interact with the routing algorithm:

- **Legacy pattern**: The `LlmTarget` object acted as a mutable container that you passed into the router and then inspected for the final model selection and response.
- **Modern pattern**: The `run_stream` method yields `Step` enum variants (such as `Step.CallModel` and `Step.Done`) that explicitly signal each phase of the routing lifecycle. You provide responses back to the router by constructing `LlmResponse` instances rather than mutating a shared object.

This shift enables proper backpressure handling and native support for streaming responses without blocking the event loop.

## Step-by-Step Migration Guide

### Update Your Imports

Replace any imports of the deprecated `LlmTarget` with the new public symbols. In [`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py), the `0.2.0` release exposes `Step` and `LlmResponse` as top-level exports:

```python

# Old (removed in 0.2.0)

# from switchyard.libsy import LlmTarget

# New

from switchyard.libsy import Step, LlmResponse, LlmClassifierConfig
from switchyard.libsy.algorithms import stage_router

```

### Construct Protocol-Compliant Requests

The `run_stream` entry point expects a normalized request dictionary defined in the protocol crate. Consult [`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs) for the exact schema, which typically includes a `messages` array and routing metadata:

```python
request = {
    "messages": [{"role": "user", "content": [{"type": "text", "text": "Explain quantum computing"}]}],
    "model": "any",  # Router overrides this based on its selection logic

}

```

### Handle Step Objects in the Event Loop

Algorithms like `stage_router` now return an async iterator of `Step` objects rather than a final `LlmTarget`. Drive the router by iterating over `run_stream` and pattern-matching on each step. The test suite in [`tests/test_libsy_minimal_bindings.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_libsy_minimal_bindings.py) demonstrates the canonical handling:

```python
async for step in algorithm.run_stream(request):
    match step:
        case Step.CallModel(call):
            # Router requests a model inference

            raw = await my_client.call(call.request)
            await call.respond(LlmResponse.Agg(raw))
        case Step.Done(outcome):
            # Final routing decision complete

            print(f"Selected: {outcome.model_id}")

```

Common `Step` variants include:
- `Step.CallModel(call)` – The router requires a classifier or judge inference before proceeding.
- `Step.Done(outcome)` – The algorithm has finalized its model selection and contains the aggregated response.

### Wrap Model Responses with LlmResponse

When the router yields `Step.CallModel`, you must wrap your client's output in an `LlmResponse` variant before calling `respond()`. The class definitions reside in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py):

- `LlmResponse.Agg(data)` – For complete, non-streaming JSON responses from your model client.
- `LlmResponse.Stream(generator)` – For async generators yielding streaming chunks.

The [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) file provides a complete reference for implementing both patterns.

## Complete Migration Example

Below is a before-and-after comparison showing the transition from the imperative `LlmTarget` style to the async `Step`/`LlmResponse` approach:

```python

# ==========================================

# Legacy API (pre-0.2.0) - DEPRECATED

# ==========================================

# from switchyard.libsy import LlmTarget

# 
# target = LlmTarget(messages=[...], model="gpt-4")

# result = router.route(target)  # Blocking, mutates target in-place

# print(result.model_id)

# ==========================================

# Modern API (0.2.0+) - Async Step Model

# ==========================================

from switchyard.libsy import Step, LlmResponse
from switchyard.libsy.algorithms import stage_router

async def route_request(user_message: str):
    algorithm = stage_router(picker="efficient_first", confidence_threshold=0.5)
    request = {
        "messages": [{"role": "user", "content": user_message}],
        "model": "any",
    }
    
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                # Integrate your own client

                response_data = await my_llm_client.complete(call.request)
                await call.respond(LlmResponse.Agg(response_data))
            case Step.Done(outcome):
                return outcome.model_id, outcome.response

```

## Key Source Files and References

Understanding the implementation details in the source code will help debug migration issues:

- **[`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py)** – Defines the public Python API, re-exporting `Step`, `LlmResponse`, and configuration classes.
- **[`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py)** – Contains the Python bindings for the `Step` enum and `LlmResponse` discriminated union (lines 44–52).
- **[`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs)** – Specifies the normalized request schema that `run_stream` consumes.
- **[`tests/test_libsy_minimal_bindings.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_libsy_minimal_bindings.py)** – Demonstrates pattern-matching against `Step` variants in a test context (lines 70–85).
- **[`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py)** – Provides a runnable example of the full `Step`-driven workflow including both `CallModel` and `Done` handling.
- **[`README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/README.md)** – Documents the high-level migration rationale and checklist (sections around lines 53–56 and 78–86).

## Summary

- **Import changes**: Replace `LlmTarget` with `Step` and `LlmResponse` from `switchyard.libsy`.
- **Execution model**: Use `run_stream` to drive the router, yielding `Step` objects instead of returning a target directly.
- **Response handling**: Wrap model outputs in `LlmResponse.Agg` or `LlmResponse.Stream` when responding to `Step.CallModel`.
- **Protocol compliance**: Ensure request dictionaries match the schema defined in [`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs).
- **Async requirement**: The new API is async-first; wrap synchronous clients with `asyncio` or use native async HTTP clients.

## Frequently Asked Questions

### What happened to LlmTarget in Switchyard 0.2.0?

The `LlmTarget` class was removed in version `0.2.0` because it tightly coupled request description, routing state, and response storage into a single mutable object. According to the Switchyard source code, this design prevented clean async streaming and made it difficult to implement backpressure. The new `Step`/`LlmResponse` architecture separates these concerns, allowing the router to yield control back to your code at each decision point.

### How do I handle streaming responses with the new API?

When you receive a `Step.CallModel`, wrap your streaming generator in `LlmResponse.Stream` instead of `LlmResponse.Agg`. Pass the generator to `call.respond()`; the router will consume it asynchronously. This pattern is defined in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) and demonstrated in the main repository examples.

### Do I need to rewrite my routing algorithms?

No. Algorithms such as `stage_router` maintain the same construction API and configuration options. They now emit `Step` objects through their `run_stream` method rather than returning an `LlmTarget` directly. Update your calling code to iterate over the stream, but the algorithm initialization logic remains unchanged.

### Where is LlmResponse defined in the codebase?

The `LlmResponse` class is defined in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) as a discriminated union with two variants: `Agg` for aggregated (complete) responses and `Stream` for streaming iterators. This file bridges the Rust core implementation to Python, exposing the types that you import from `switchyard.libsy`.