# LiteLLM Streaming Responses Across LLM Providers: Architecture and Implementation

> Master LiteLLM streaming responses from OpenAI, Anthropic, Gemini & more. Discover the unified, provider-agnostic interface and three-layer architecture for normalized real-time LLM data.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: architecture
- Published: 2026-03-26

---

**LiteLLM provides a unified, provider-agnostic streaming interface that normalizes real-time responses from OpenAI, Anthropic, Gemini, Azure, and other LLM providers through a reusable three-layer architecture involving transport handling, chunk transformation, and integrated logging hooks.**

LiteLLM enables developers to stream responses from any supported LLM provider using a single, consistent API. Whether you are consuming Server-Sent Events from OpenAI or handling WebSocket connections for real-time applications, the library abstracts transport differences into a standardized `ResponsesAPIStreamingResponse` model. This article examines the streaming architecture implemented in the BerriAI/litellm repository, detailing the iterator patterns, provider-specific transformers, and observability hooks that power the system.

## The Three-Layer Streaming Architecture

The streaming stack in [`litellm/responses/streaming_iterator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/responses/streaming_iterator.py) is built on three reusable layers that ensure consistent behavior regardless of the underlying provider.

### Layer 1: Transport and Chunk Acquisition

The base transport layer reads line-by-line chunks from an `httpx.Response` and enforces maximum stream durations to prevent runaway connections. The `BaseResponsesAPIStreamingIterator` class manages both async (`ResponsesAPIStreamingIterator`) and sync (`SyncResponsesAPIStreamingIterator`) implementations.

Before yielding each chunk, the iterator calls `_check_max_streaming_duration` to verify the stream has not exceeded `LITELLM_MAX_STREAMING_DURATION_SECONDS` (defined in [`litellm/constants.py`](https://github.com/BerriAI/litellm/blob/main/litellm/constants.py)). This safety mechanism ensures that long-running streams against providers like Azure or Anthropic terminate predictably if they exceed configured timeouts.

### Layer 2: Chunk Parsing and Transformation

The `_process_chunk` method (lines 104–176 in [`streaming_iterator.py`](https://github.com/BerriAI/litellm/blob/main/streaming_iterator.py)) handles the normalization of provider-specific payloads. This layer:

- Strips SSE wrappers using `CustomStreamWrapper._strip_sse_data_from_chunk` to remove the `data:` prefix
- Detects the `[DONE]` marker via `STREAM_SSE_DONE_STRING` to set `self.finished`
- Parses JSON and invokes the provider-specific transformer via `responses_api_provider_config.transform_streaming_response`
- Converts native provider schemas into the canonical `ResponsesAPIStreamingResponse` model
- Injects hidden metadata (model ID, API base) for downstream logging

For encrypted content, the iterator wraps ciphertext with the model-ID so downstream decryption routines know the source context.

### Layer 3: Hook and Logging Pipeline

After transformation, the iterator executes the logging and hook pipeline through methods like `_handle_logging_completed_response`, `_run_post_success_hooks`, and `_handle_failure` (lines 170–258 and 295–336). This layer:

- Creates a deep-copy of completed responses using Pydantic’s `model_dump` / `model_validate` to prevent mutation of user-facing objects
- Fires async and sync success handlers (`async_success_handler`, `success_handler`)
- Captures errors once via `_failure_handled` and routes them through `async_failure_handler` before re-raising

## How the Streaming Iterator Works

Understanding the iteration flow reveals how LiteLLM maintains uniform behavior across HTTP and WebSocket transports.

### Iterator Initialization

When you call `litellm.aresponses()` with `stream=True`, the library creates an `httpx.AsyncClient`, receives the raw response, and wraps it in a `ResponsesAPIStreamingIterator`:

```python
response = httpx.get(url, stream=True)
iterator = ResponsesAPIStreamingIterator(
    response=response,
    model="gpt-4o-mini",
    responses_api_provider_config=provider_cfg,
    logging_obj=lite_logger,
    litellm_metadata=metadata,
)

```

The iterator stores the raw `httpx.Response`, model name, logging object, and any extra metadata needed for the provider-specific transformer.

### Iteration and Duration Enforcement

Each call to `__anext__` (or `__next__` for sync) pulls the next line from `response.aiter_lines()` or `response.iter_lines()`. Before yielding, the iterator verifies the stream duration against `LITELLM_MAX_STREAMING_DURATION_SECONDS`, throwing a timeout error if the provider stalls.

### Chunk Processing and Normalization

Raw lines pass through `_process_chunk`, which detects response completion events (`RESPONSE_COMPLETED`, `RESPONSE_FAILED`) and stores the final response in `self.completed_response` for cost calculation and logging.

### Post-Streaming Hook Execution

The `_call_post_streaming_deployment_hook` method iterates over registered callbacks in `litellm.callbacks`, invoking `async_post_call_streaming_deployment_hook` on each. This allows users to modify chunks before they reach the caller—for example, injecting token-level telemetry or applying custom sanitization rules.

### Mock Streaming for Non-Streaming Models

For models that do not natively support streaming (such as `o1-pro`), LiteLLM provides `MockResponsesAPIStreamingIterator`. This utility slices the full response text into 5-character deltas and yields synthetic `ResponseCompletedEvent` objects, guaranteeing a uniform streaming API even when the underlying provider only supports synchronous responses.

## Provider-Specific Adapters

Each LLM provider implements a transformer that maps its native streaming payload to the canonical `ResponsesAPIStreamingResponse` schema. These adapters are passed to the iterator via `responses_api_provider_config`, keeping the core iterator provider-agnostic.

| Provider | Transformer Location |
|----------|---------------------|
| **OpenAI / Azure** | [`litellm/llms/openai/chat/streaming.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/openai/chat/streaming.py) (`OpenAIConfig.transform_streaming_response`) |
| **Anthropic** | [`litellm/llms/anthropic/experimental_pass_through/adapters/streaming_iterator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/anthropic/experimental_pass_through/adapters/streaming_iterator.py) |
| **Gemini** | [`litellm/llms/google_genai/streaming_iterator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/google_genai/streaming_iterator.py) |
| **Databricks** | [`litellm/llms/databricks/streaming_utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/databricks/streaming_utils.py) |

These transformers handle provider quirks—such as Anthropic’s `content_block_delta` events or Gemini’s `candidates` structure—while the core iterator remains unchanged.

## WebSocket Streaming Support

LiteLLM extends its streaming architecture to WebSocket transports for the Responses API (`v1/responses`).

### Direct WebSocket Forwarding

`ResponsesWebSocketStreaming` (lines 237–388 in [`streaming_iterator.py`](https://github.com/BerriAI/litellm/blob/main/streaming_iterator.py)) forwards events bidirectionally between a client WebSocket and the provider’s backend WebSocket, logging events through the standard pipeline.

### Managed WebSocket Handler

For providers that only expose HTTP streaming, `ManagedResponsesWebSocketHandler` (lines 409–754) bridges the gap:

1. Parses `response.create` events from the client WebSocket
2. Injects conversation history from an in-memory session cache
3. Calls `litellm.aresponses(stream=True)` under the hood
4. Serializes each chunk back to the client
5. Updates the session cache for multi-turn conversations

Both classes reuse `ResponsesAPIStreamingIterator`, ensuring identical event schemas whether the transport is HTTP or WebSocket.

## Practical Implementation Examples

### Async Streaming with Native Providers

```python
import litellm

async def stream_chat():
    async for chunk in litellm.aresponses(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
        stream=True,
    ):
        if chunk.type == "output_text.delta":
            print(chunk.delta, end="", flush=True)
        elif chunk.type == "response.completed":
            print("\n--- done ---")

```

### Sync Streaming with Automatic Fallback

```python
import litellm

def sync_stream():
    for chunk in litellm.responses(
        model="o1-mini",
        messages=[{"role": "user", "content": "Summarise the plot of Inception"}],
        stream=True,
    ):
        if chunk.type == "output_text.delta":
            print(chunk.delta, end="")
        elif chunk.type ==