How to Use the Asynchronous Client in gpt4free: Complete Guide with Examples

The AsyncClient class in gpt4free provides a fully non-blocking interface to LLM providers using Python's asyncio, enabling high-performance chat completions and image generation through await client.chat.completions.create() and real-time streaming via async for loops.

The gpt4free library offers a robust asynchronous API that mirrors the OpenAI SDK while supporting multiple LLM providers without blocking your event loop. Whether you are building high-throughput applications or integrating AI into existing async frameworks like FastAPI or Sanic, understanding how to use the asynchronous client in gpt4free is essential for modern Python development.

Core Architecture of the Asynchronous Client

The async stack in gpt4free centers on three primary classes defined in g4f/client/__init__.py:

  • AsyncClient (lines 590-604): The top-level entry point that configures providers, authentication, and proxies.
  • AsyncChat (lines 606-610): Holds the chat-completion namespace accessible via client.chat.completions.
  • AsyncCompletions (lines 611-666): Implements the actual async create and stream methods that dispatch requests to the chosen provider.

Under the hood, the library uses async_iter_run_tools (from g4f/tools/run_tools.py) to wrap provider-specific capabilities like web search or code execution into an async iterator. The async_iter_response helper then normalizes raw provider streams into standard ChatCompletionChunk or ChatCompletion objects.

Creating an AsyncClient Instance

Basic Construction

The simplest way to start is by instantiating AsyncClient directly, which defaults to the AnyProvider that automatically selects the first working provider:

from g4f.client import AsyncClient

client = AsyncClient()

Using ClientFactory for Specific Providers

For production use cases requiring a specific provider or custom authentication, use the ClientFactory class. The create_async_client method (lines 987-1010 in g4f/client/__init__.py) handles provider resolution and setup:

from g4f.client import ClientFactory

client = ClientFactory.create_async_client(
    provider="PollinationsAI",              # Provider name or custom class

    base_url="https://api.example.com/v1",  # Optional custom endpoint

    api_key="YOUR_API_KEY",                 # Optional authentication

)

Performing Asynchronous Chat Completions

Non-Streaming Requests

To generate a complete response in a single call, use await with the create method. The AsyncCompletions.create implementation normalizes arguments, resolves the provider, and assembles the final ChatCompletion object:

response = await client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user",   "content": "Explain how async/await works in Python."}
    ],
    stream=False,
)

print(response.choices[0].message.content)

Streaming Responses

For real-time token generation, use the stream method (lines 682-688) which returns an async iterator yielding ChatCompletionChunk objects:

async for chunk in client.chat.completions.stream(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Write a haiku about coding."}]
):
    # Access incremental content via the delta field

    print(chunk.choices[0].delta.content, end="", flush=True)

Generating Images Asynchronously

The async image API mirrors the synchronous interface but leverages await for non-blocking I/O. The AsyncImages.generate method (lines 698-706) supports URL, base64, or local file responses:

images = await client.images.generate(
    prompt="A photorealistic cat wearing a space helmet",
    model="dalle-3",
    response_format="url"  # Options: "url", "b64_json", or None (default)

)

for img in images.data:
    print(img.url)  # Direct URL to the generated image

Key Implementation Files

Understanding the source structure helps when debugging or extending the async functionality:

File Primary Responsibility
g4f/client/__init__.py Core async classes (AsyncClient, AsyncCompletions, AsyncImages) and ClientFactory
g4f/client/service.py Provider mapping via convert_to_provider
g4f/tools/run_tools.py Async tooling pipeline (async_iter_run_tools)
etc/examples/text_completions_demo_async.py Minimal runnable demonstration

Tips and Best Practices

  • Avoid blocking the event loop: Never call asyncio.run() inside an already-running async context (e.g., within a FastAPI endpoint). Use await directly instead.
  • Provider fallback: When using ClientFactory.create_async_client with provider=None, the library automatically selects the first working provider from the available pool.
  • Custom endpoints: Pass base_url and api_key to the factory to create a CustomProvider for self-hosted OpenAI-compatible APIs.
  • Tool execution: The async client automatically routes tool-enabled requests through async_iter_run_tools. No additional configuration is required beyond setting appropriate parameters.
  • Memory efficiency: For long-running streams, process chunks immediately rather than accumulating them in memory to avoid excessive RAM usage.

Summary

  • The AsyncClient class in g4f/client/__init__.py provides the main entry point for non-blocking LLM interactions.
  • Use await client.chat.completions.create() for standard requests and async for chunk in client.chat.completions.stream() for real-time output.
  • ClientFactory.create_async_client() enables provider-specific configuration, custom base URLs, and authentication.
  • Image generation follows the same pattern via await client.images.generate().
  • The architecture relies on async_iter_run_tools and async_iter_response to normalize provider-specific behavior into standard OpenAI-compatible objects.

Frequently Asked Questions

How do I handle streaming responses with the asynchronous client in gpt4free?

Use the stream method available on client.chat.completions, which returns an async iterator yielding ChatCompletionChunk objects. Access incremental content through chunk.choices[0].delta.content inside an async for loop. This implementation resides in AsyncCompletions.stream (lines 682-688 of g4f/client/__init__.py).

Can I use a specific provider instead of the default AnyProvider with AsyncClient?

Yes. Instead of instantiating AsyncClient directly, use ClientFactory.create_async_client() and pass the provider name as a string (e.g., "PollinationsAI" or "DeepInfra") or a custom provider class. This factory method, located at lines 987-1010 in g4f/client/__init__.py, handles provider resolution, authentication, and custom base URL configuration.

What is the difference between the response_format options for async image generation?

The response_format parameter in client.images.generate() accepts three values: "url" returns direct URLs to hosted images without local download, "b64_json" returns base64-encoded image data suitable for embedding directly in HTML or JSON responses, and None (default) typically triggers local file download. This logic is implemented in AsyncImages.generate (lines 698-706).

How does the async client handle tool calling and function execution automatically?

The AsyncCompletions class automatically routes requests through async_iter_run_tools (defined in g4f/tools/run_tools.py) when it detects tool configurations. This helper wraps provider-specific capabilities like web search or code execution into an async iterator that chains with async_iter_response to normalize output into standard ChatCompletion objects. No manual intervention is required beyond setting appropriate parameters in the create or stream calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →