How gpt4free Handles Streaming Responses: A Deep Dive into the Async Architecture

gpt4free handles streaming responses by converting asynchronous generators from providers into Python iterators through a four-stage pipeline involving request entry, provider generation, tool handling, and response normalization.

The xtekky/gpt4free library abstracts streaming capabilities across dozens of LLM providers into a unified interface. Understanding how gpt4free handles streaming responses reveals a sophisticated architecture that bridges network-level HTTP streaming with high-level Python iterators, allowing developers to consume partial model outputs in real time regardless of the underlying provider.

The Four-Stage Streaming Pipeline

The streaming flow follows a tightly-coupled chain of components that transform raw HTTP responses into structured chunks. Each stage serves a distinct purpose in the data transformation process.

Stage 1: Request Entry Point

Streaming begins when the public API receives stream=True. In g4f/client/__init__.py, the Completions.create method (lines 30-45) forwards the call to either iter_run_tools (synchronous) or async_iter_run_tools (asynchronous), passing the selected provider, model, and the critical stream flag.


# g4f/client/__init__.py

# Completions.create handles the initial stream=True parameter

def create(self, messages, model, stream=False, **kwargs):
    # Lines 30-45: forwards to iter_run_tools or async_iter_run_tools

    return iter_run_tools(stream=stream, ...)

Stage 2: Provider Generation

The selected provider implements create_async_generator to yield chunks as they arrive from the remote API. Most providers inherit from AsyncGeneratorProvider defined in g4f/providers/base_provider.py (lines 18-26). This base class guarantees that the method yields chunks containing str, FinishReason, ToolCalls, or other response types immediately upon receipt.


# g4f/providers/base_provider.py

class AsyncGeneratorProvider:
    @classmethod
    async def create_async_generator(cls, model, messages, stream=True, **kwargs):
        # Lines 18-26: yields chunks as they arrive from HTTP response

        yield chunk

Stage 3: Tool and Usage Handling

The async_iter_run_tools function in g4f/tools/run_tools.py (lines 24-48) wraps the provider generator. This stage adds optional tool-emulation, web-search capabilities, usage-tracking, and crucially preserves the stream flag when yielding each chunk downstream. The synchronous counterpart iter_run_tools (lines 71-94) performs the same function for blocking code.


# g4f/tools/run_tools.py

async def async_iter_run_tools(response, stream=True, **kwargs):
    # Lines 24-48: wraps provider generator, preserves stream flag

    async for chunk in response:
        if stream:
            yield chunk

Stage 4: Response Normalization

The low-level iterator passes to iter_response in g4f/client/__init__.py (lines 68-126). This function walks through every chunk, builds the final content string, collects Reasoning, ToolCalls, Usage, and stops when max_tokens or a FinishReason is encountered. When stream=True, it yields a ChatCompletionChunk after each processed chunk; otherwise, it returns a fully assembled ChatCompletion.


# g4f/client/__init__.py

def iter_response(response, stream=True, max_tokens=None, **kwargs):
    # Lines 68-126: normalizes chunks into ChatCompletionChunk or ChatCompletion

    for chunk in response:
        if stream:
            yield ChatCompletionChunk(delta=chunk)

How the Streaming Flag Propagates Through the System

The stream parameter travels through multiple layers to ensure consistent behavior:

  1. User call: client.chat.completions.create(messages, model, stream=True)
  2. Client layer: Inside Completions.create, the flag passes to iter_run_tools via stream=stream
  3. Tool layer: iter_run_tools forwards stream to Provider.async_create_function. Providers that support streaming have supports_stream = True in g4f/providers/types.py
  4. Response layer: iter_response receives the generator and, if stream is true, yields each intermediate ChatCompletionChunk immediately

This propagation ensures that network-level streaming translates into real-time Python iteration regardless of which provider handles the request.

Under-the-Hood HTTP Streaming Implementation

When providers use the curl_cffi request layer, they create a StreamResponse object with stream=True. In g4f/requests/curl_cffi.py (line 15), the underlying library reads the HTTP response line-by-line and forwards each line to the async generator exposed by the provider.


# g4f/requests/curl_cffi.py

def request(...):
    # Line 15: Creates StreamResponse with stream=True

    return StreamResponse(super().request(method, url, stream=True, verify=ssl, **kwargs))

Thus, network-level streaming transforms into Python-level async iteration, which then flows through the four-stage pipeline described above.

Practical Code Examples

Synchronous Streaming with Client

The standard Client class provides blocking streaming through a Python iterator:

import g4f

client = g4f.Client()
messages = [{"role": "user", "content": "Tell me a short story"}]

# stream=True makes the call return an iterator of chunks

for chunk in client.chat.completions.create(
    messages,
    model="gpt-4o-mini",
    stream=True
):
    # each chunk is a ChatCompletionChunk (or FinishReason)

    print(chunk.delta, end="", flush=True)

Asynchronous Streaming with AsyncClient

For non-blocking I/O, use AsyncClient with async for:

import asyncio
import g4f

async def main():
    client = g4f.AsyncClient()
    messages = [{"role": "user", "content": "Write a poem about moons"}]

    # stream=True returns an async iterator

    async for chunk in client.chat.completions.create(
        messages,
        model="gpt-4o-mini",
        stream=True
    ):
        print(chunk.delta, end="", flush=True)

asyncio.run(main())

Direct Provider Streaming

You can bypass the client abstraction and stream directly from a specific provider:

from g4f.Provider import OpenaiAccount

# Directly call the provider's async generator

async for part in OpenaiAccount.create_async_generator(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
    stream=True
):
    print(part)  # part can be a string, FinishReason, etc.

Summary

  • gpt4free handles streaming responses through a four-stage pipeline: request entry, provider generation, tool handling, and response normalization.
  • The architecture converts network-level HTTP streaming into Python async generators using curl_cffi with stream=True in g4f/requests/curl_cffi.py.
  • The stream flag propagates from Completions.create through iter_run_tools to provider implementations and finally to iter_response, which yields ChatCompletionChunk objects in real time.
  • Both synchronous (Client) and asynchronous (AsyncClient) interfaces support streaming through standard Python iteration protocols.

Frequently Asked Questions

What is the difference between synchronous and asynchronous streaming in gpt4free?

Synchronous streaming uses the standard g4f.Client class and returns a Python iterator that you consume with a for loop. Asynchronous streaming uses g4f.AsyncClient and returns an async generator consumed with async for. Both interfaces ultimately rely on the same underlying create_async_generator methods in the providers, but the async version avoids blocking the event loop during HTTP I/O operations.

How does gpt4free handle network-level streaming from different providers?

gpt4free uses the curl_cffi library to manage HTTP connections. In g4f/requests/curl_cffi.py, the request method creates a StreamResponse object with stream=True, which tells the underlying library to read the response line-by-line as data arrives. Each line is then yielded through the provider's create_async_generator method, creating a seamless bridge between raw HTTP streaming and Python async iteration.

Can I use tool calling with streaming responses in gpt4free?

Yes, tool calling works with streaming responses. The async_iter_run_tools function in g4f/tools/run_tools.py wraps the provider generator and handles tool emulation, web search, and usage tracking while preserving the stream=True flag. As chunks arrive, the tool handler processes them and yields intermediate results, allowing you to stream partial tool outputs or final completions depending on the implementation.

Where is the streaming flag checked in the gpt4free codebase?

The streaming flag is checked at multiple critical points throughout the codebase. Initially, it is received in Completions.create in g4f/client/__init__.py (lines 30-45). It then propagates to iter_run_tools and async_iter_run_tools in g4f/tools/run_tools.py (lines 24-48 and 71-94). Finally, the flag determines behavior in iter_response in g4f/client/__init__.py (lines 68-126), where it decides whether to yield ChatCompletionChunk objects or return a complete ChatCompletion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →