LiteLLM Streaming Responses Across LLM Providers: Architecture and Implementation

LiteLLM provides a unified, provider-agnostic streaming interface that normalizes real-time responses from OpenAI, Anthropic, Gemini, Azure, and other LLM providers through a reusable three-layer architecture involving transport handling, chunk transformation, and integrated logging hooks.

LiteLLM enables developers to stream responses from any supported LLM provider using a single, consistent API. Whether you are consuming Server-Sent Events from OpenAI or handling WebSocket connections for real-time applications, the library abstracts transport differences into a standardized ResponsesAPIStreamingResponse model. This article examines the streaming architecture implemented in the BerriAI/litellm repository, detailing the iterator patterns, provider-specific transformers, and observability hooks that power the system.

The Three-Layer Streaming Architecture

The streaming stack in litellm/responses/streaming_iterator.py is built on three reusable layers that ensure consistent behavior regardless of the underlying provider.

Layer 1: Transport and Chunk Acquisition

The base transport layer reads line-by-line chunks from an httpx.Response and enforces maximum stream durations to prevent runaway connections. The BaseResponsesAPIStreamingIterator class manages both async (ResponsesAPIStreamingIterator) and sync (SyncResponsesAPIStreamingIterator) implementations.

Before yielding each chunk, the iterator calls _check_max_streaming_duration to verify the stream has not exceeded LITELLM_MAX_STREAMING_DURATION_SECONDS (defined in litellm/constants.py). This safety mechanism ensures that long-running streams against providers like Azure or Anthropic terminate predictably if they exceed configured timeouts.

Layer 2: Chunk Parsing and Transformation

The _process_chunk method (lines 104–176 in streaming_iterator.py) handles the normalization of provider-specific payloads. This layer:

  • Strips SSE wrappers using CustomStreamWrapper._strip_sse_data_from_chunk to remove the data: prefix
  • Detects the [DONE] marker via STREAM_SSE_DONE_STRING to set self.finished
  • Parses JSON and invokes the provider-specific transformer via responses_api_provider_config.transform_streaming_response
  • Converts native provider schemas into the canonical ResponsesAPIStreamingResponse model
  • Injects hidden metadata (model ID, API base) for downstream logging

For encrypted content, the iterator wraps ciphertext with the model-ID so downstream decryption routines know the source context.

Layer 3: Hook and Logging Pipeline

After transformation, the iterator executes the logging and hook pipeline through methods like _handle_logging_completed_response, _run_post_success_hooks, and _handle_failure (lines 170–258 and 295–336). This layer:

  • Creates a deep-copy of completed responses using Pydantic’s model_dump / model_validate to prevent mutation of user-facing objects
  • Fires async and sync success handlers (async_success_handler, success_handler)
  • Captures errors once via _failure_handled and routes them through async_failure_handler before re-raising

How the Streaming Iterator Works

Understanding the iteration flow reveals how LiteLLM maintains uniform behavior across HTTP and WebSocket transports.

Iterator Initialization

When you call litellm.aresponses() with stream=True, the library creates an httpx.AsyncClient, receives the raw response, and wraps it in a ResponsesAPIStreamingIterator:

response = httpx.get(url, stream=True)
iterator = ResponsesAPIStreamingIterator(
    response=response,
    model="gpt-4o-mini",
    responses_api_provider_config=provider_cfg,
    logging_obj=lite_logger,
    litellm_metadata=metadata,
)

The iterator stores the raw httpx.Response, model name, logging object, and any extra metadata needed for the provider-specific transformer.

Iteration and Duration Enforcement

Each call to __anext__ (or __next__ for sync) pulls the next line from response.aiter_lines() or response.iter_lines(). Before yielding, the iterator verifies the stream duration against LITELLM_MAX_STREAMING_DURATION_SECONDS, throwing a timeout error if the provider stalls.

Chunk Processing and Normalization

Raw lines pass through _process_chunk, which detects response completion events (RESPONSE_COMPLETED, RESPONSE_FAILED) and stores the final response in self.completed_response for cost calculation and logging.

Post-Streaming Hook Execution

The _call_post_streaming_deployment_hook method iterates over registered callbacks in litellm.callbacks, invoking async_post_call_streaming_deployment_hook on each. This allows users to modify chunks before they reach the caller—for example, injecting token-level telemetry or applying custom sanitization rules.

Mock Streaming for Non-Streaming Models

For models that do not natively support streaming (such as o1-pro), LiteLLM provides MockResponsesAPIStreamingIterator. This utility slices the full response text into 5-character deltas and yields synthetic ResponseCompletedEvent objects, guaranteeing a uniform streaming API even when the underlying provider only supports synchronous responses.

Provider-Specific Adapters

Each LLM provider implements a transformer that maps its native streaming payload to the canonical ResponsesAPIStreamingResponse schema. These adapters are passed to the iterator via responses_api_provider_config, keeping the core iterator provider-agnostic.

Provider Transformer Location
OpenAI / Azure litellm/llms/openai/chat/streaming.py (OpenAIConfig.transform_streaming_response)
Anthropic litellm/llms/anthropic/experimental_pass_through/adapters/streaming_iterator.py
Gemini litellm/llms/google_genai/streaming_iterator.py
Databricks litellm/llms/databricks/streaming_utils.py

These transformers handle provider quirks—such as Anthropic’s content_block_delta events or Gemini’s candidates structure—while the core iterator remains unchanged.

WebSocket Streaming Support

LiteLLM extends its streaming architecture to WebSocket transports for the Responses API (v1/responses).

Direct WebSocket Forwarding

ResponsesWebSocketStreaming (lines 237–388 in streaming_iterator.py) forwards events bidirectionally between a client WebSocket and the provider’s backend WebSocket, logging events through the standard pipeline.

Managed WebSocket Handler

For providers that only expose HTTP streaming, ManagedResponsesWebSocketHandler (lines 409–754) bridges the gap:

  1. Parses response.create events from the client WebSocket
  2. Injects conversation history from an in-memory session cache
  3. Calls litellm.aresponses(stream=True) under the hood
  4. Serializes each chunk back to the client
  5. Updates the session cache for multi-turn conversations

Both classes reuse ResponsesAPIStreamingIterator, ensuring identical event schemas whether the transport is HTTP or WebSocket.

Practical Implementation Examples

Async Streaming with Native Providers

import litellm

async def stream_chat():
    async for chunk in litellm.aresponses(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
        stream=True,
    ):
        if chunk.type == "output_text.delta":
            print(chunk.delta, end="", flush=True)
        elif chunk.type == "response.completed":
            print("\n--- done ---")

Sync Streaming with Automatic Fallback

import litellm

def sync_stream():
    for chunk in litellm.responses(
        model="o1-mini",
        messages=[{"role": "user", "content": "Summarise the plot of Inception"}],
        stream=True,
    ):
        if chunk.type == "output_text.delta":
            print(chunk.delta, end="")
        elif chunk.type ==

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →