LiteLLM Streaming Responses Across LLM Providers: Architecture and Implementation
LiteLLM provides a unified, provider-agnostic streaming interface that normalizes real-time responses from OpenAI, Anthropic, Gemini, Azure, and other LLM providers through a reusable three-layer architecture involving transport handling, chunk transformation, and integrated logging hooks.
LiteLLM enables developers to stream responses from any supported LLM provider using a single, consistent API. Whether you are consuming Server-Sent Events from OpenAI or handling WebSocket connections for real-time applications, the library abstracts transport differences into a standardized ResponsesAPIStreamingResponse model. This article examines the streaming architecture implemented in the BerriAI/litellm repository, detailing the iterator patterns, provider-specific transformers, and observability hooks that power the system.
The Three-Layer Streaming Architecture
The streaming stack in litellm/responses/streaming_iterator.py is built on three reusable layers that ensure consistent behavior regardless of the underlying provider.
Layer 1: Transport and Chunk Acquisition
The base transport layer reads line-by-line chunks from an httpx.Response and enforces maximum stream durations to prevent runaway connections. The BaseResponsesAPIStreamingIterator class manages both async (ResponsesAPIStreamingIterator) and sync (SyncResponsesAPIStreamingIterator) implementations.
Before yielding each chunk, the iterator calls _check_max_streaming_duration to verify the stream has not exceeded LITELLM_MAX_STREAMING_DURATION_SECONDS (defined in litellm/constants.py). This safety mechanism ensures that long-running streams against providers like Azure or Anthropic terminate predictably if they exceed configured timeouts.
Layer 2: Chunk Parsing and Transformation
The _process_chunk method (lines 104–176 in streaming_iterator.py) handles the normalization of provider-specific payloads. This layer:
- Strips SSE wrappers using
CustomStreamWrapper._strip_sse_data_from_chunkto remove thedata:prefix - Detects the
[DONE]marker viaSTREAM_SSE_DONE_STRINGto setself.finished - Parses JSON and invokes the provider-specific transformer via
responses_api_provider_config.transform_streaming_response - Converts native provider schemas into the canonical
ResponsesAPIStreamingResponsemodel - Injects hidden metadata (model ID, API base) for downstream logging
For encrypted content, the iterator wraps ciphertext with the model-ID so downstream decryption routines know the source context.
Layer 3: Hook and Logging Pipeline
After transformation, the iterator executes the logging and hook pipeline through methods like _handle_logging_completed_response, _run_post_success_hooks, and _handle_failure (lines 170–258 and 295–336). This layer:
- Creates a deep-copy of completed responses using Pydantic’s
model_dump/model_validateto prevent mutation of user-facing objects - Fires async and sync success handlers (
async_success_handler,success_handler) - Captures errors once via
_failure_handledand routes them throughasync_failure_handlerbefore re-raising
How the Streaming Iterator Works
Understanding the iteration flow reveals how LiteLLM maintains uniform behavior across HTTP and WebSocket transports.
Iterator Initialization
When you call litellm.aresponses() with stream=True, the library creates an httpx.AsyncClient, receives the raw response, and wraps it in a ResponsesAPIStreamingIterator:
response = httpx.get(url, stream=True)
iterator = ResponsesAPIStreamingIterator(
response=response,
model="gpt-4o-mini",
responses_api_provider_config=provider_cfg,
logging_obj=lite_logger,
litellm_metadata=metadata,
)
The iterator stores the raw httpx.Response, model name, logging object, and any extra metadata needed for the provider-specific transformer.
Iteration and Duration Enforcement
Each call to __anext__ (or __next__ for sync) pulls the next line from response.aiter_lines() or response.iter_lines(). Before yielding, the iterator verifies the stream duration against LITELLM_MAX_STREAMING_DURATION_SECONDS, throwing a timeout error if the provider stalls.
Chunk Processing and Normalization
Raw lines pass through _process_chunk, which detects response completion events (RESPONSE_COMPLETED, RESPONSE_FAILED) and stores the final response in self.completed_response for cost calculation and logging.
Post-Streaming Hook Execution
The _call_post_streaming_deployment_hook method iterates over registered callbacks in litellm.callbacks, invoking async_post_call_streaming_deployment_hook on each. This allows users to modify chunks before they reach the caller—for example, injecting token-level telemetry or applying custom sanitization rules.
Mock Streaming for Non-Streaming Models
For models that do not natively support streaming (such as o1-pro), LiteLLM provides MockResponsesAPIStreamingIterator. This utility slices the full response text into 5-character deltas and yields synthetic ResponseCompletedEvent objects, guaranteeing a uniform streaming API even when the underlying provider only supports synchronous responses.
Provider-Specific Adapters
Each LLM provider implements a transformer that maps its native streaming payload to the canonical ResponsesAPIStreamingResponse schema. These adapters are passed to the iterator via responses_api_provider_config, keeping the core iterator provider-agnostic.
| Provider | Transformer Location |
|---|---|
| OpenAI / Azure | litellm/llms/openai/chat/streaming.py (OpenAIConfig.transform_streaming_response) |
| Anthropic | litellm/llms/anthropic/experimental_pass_through/adapters/streaming_iterator.py |
| Gemini | litellm/llms/google_genai/streaming_iterator.py |
| Databricks | litellm/llms/databricks/streaming_utils.py |
These transformers handle provider quirks—such as Anthropic’s content_block_delta events or Gemini’s candidates structure—while the core iterator remains unchanged.
WebSocket Streaming Support
LiteLLM extends its streaming architecture to WebSocket transports for the Responses API (v1/responses).
Direct WebSocket Forwarding
ResponsesWebSocketStreaming (lines 237–388 in streaming_iterator.py) forwards events bidirectionally between a client WebSocket and the provider’s backend WebSocket, logging events through the standard pipeline.
Managed WebSocket Handler
For providers that only expose HTTP streaming, ManagedResponsesWebSocketHandler (lines 409–754) bridges the gap:
- Parses
response.createevents from the client WebSocket - Injects conversation history from an in-memory session cache
- Calls
litellm.aresponses(stream=True)under the hood - Serializes each chunk back to the client
- Updates the session cache for multi-turn conversations
Both classes reuse ResponsesAPIStreamingIterator, ensuring identical event schemas whether the transport is HTTP or WebSocket.
Practical Implementation Examples
Async Streaming with Native Providers
import litellm
async def stream_chat():
async for chunk in litellm.aresponses(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
stream=True,
):
if chunk.type == "output_text.delta":
print(chunk.delta, end="", flush=True)
elif chunk.type == "response.completed":
print("\n--- done ---")
Sync Streaming with Automatic Fallback
import litellm
def sync_stream():
for chunk in litellm.responses(
model="o1-mini",
messages=[{"role": "user", "content": "Summarise the plot of Inception"}],
stream=True,
):
if chunk.type == "output_text.delta":
print(chunk.delta, end="")
elif chunk.type ==Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →