# How to Use the Asynchronous Client in gpt4free: Complete Guide with Examples

> Master the gpt4free asynchronous client for high-performance LLM interactions. Learn to use asyncio for non-blocking chat completions and real-time streaming with await and async for.

- Repository: [Tekky/gpt4free](https://github.com/xtekky/gpt4free)
- Tags: how-to-guide
- Published: 2026-03-04

---

**The `AsyncClient` class in gpt4free provides a fully non-blocking interface to LLM providers using Python's `asyncio`, enabling high-performance chat completions and image generation through `await client.chat.completions.create()` and real-time streaming via `async for` loops.**

The gpt4free library offers a robust asynchronous API that mirrors the OpenAI SDK while supporting multiple LLM providers without blocking your event loop. Whether you are building high-throughput applications or integrating AI into existing async frameworks like FastAPI or Sanic, understanding how to use the asynchronous client in gpt4free is essential for modern Python development.

## Core Architecture of the Asynchronous Client

The async stack in gpt4free centers on three primary classes defined in [`g4f/client/__init__.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/client/__init__.py):

- **`AsyncClient`** (lines 590-604): The top-level entry point that configures providers, authentication, and proxies.
- **`AsyncChat`** (lines 606-610): Holds the chat-completion namespace accessible via `client.chat.completions`.
- **`AsyncCompletions`** (lines 611-666): Implements the actual async `create` and `stream` methods that dispatch requests to the chosen provider.

Under the hood, the library uses `async_iter_run_tools` (from [`g4f/tools/run_tools.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/tools/run_tools.py)) to wrap provider-specific capabilities like web search or code execution into an async iterator. The `async_iter_response` helper then normalizes raw provider streams into standard `ChatCompletionChunk` or `ChatCompletion` objects.

## Creating an AsyncClient Instance

### Basic Construction

The simplest way to start is by instantiating `AsyncClient` directly, which defaults to the `AnyProvider` that automatically selects the first working provider:

```python
from g4f.client import AsyncClient

client = AsyncClient()

```

### Using ClientFactory for Specific Providers

For production use cases requiring a specific provider or custom authentication, use the `ClientFactory` class. The `create_async_client` method (lines 987-1010 in [`g4f/client/__init__.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/client/__init__.py)) handles provider resolution and setup:

```python
from g4f.client import ClientFactory

client = ClientFactory.create_async_client(
    provider="PollinationsAI",              # Provider name or custom class

    base_url="https://api.example.com/v1",  # Optional custom endpoint

    api_key="YOUR_API_KEY",                 # Optional authentication

)

```

## Performing Asynchronous Chat Completions

### Non-Streaming Requests

To generate a complete response in a single call, use `await` with the `create` method. The `AsyncCompletions.create` implementation normalizes arguments, resolves the provider, and assembles the final `ChatCompletion` object:

```python
response = await client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user",   "content": "Explain how async/await works in Python."}
    ],
    stream=False,
)

print(response.choices[0].message.content)

```

### Streaming Responses

For real-time token generation, use the `stream` method (lines 682-688) which returns an async iterator yielding `ChatCompletionChunk` objects:

```python
async for chunk in client.chat.completions.stream(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Write a haiku about coding."}]
):
    # Access incremental content via the delta field

    print(chunk.choices[0].delta.content, end="", flush=True)

```

## Generating Images Asynchronously

The async image API mirrors the synchronous interface but leverages `await` for non-blocking I/O. The `AsyncImages.generate` method (lines 698-706) supports URL, base64, or local file responses:

```python
images = await client.images.generate(
    prompt="A photorealistic cat wearing a space helmet",
    model="dalle-3",
    response_format="url"  # Options: "url", "b64_json", or None (default)

)

for img in images.data:
    print(img.url)  # Direct URL to the generated image

```

## Key Implementation Files

Understanding the source structure helps when debugging or extending the async functionality:

| File | Primary Responsibility |
|------|------------------------|
| [`g4f/client/__init__.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/client/__init__.py) | Core async classes (`AsyncClient`, `AsyncCompletions`, `AsyncImages`) and `ClientFactory` |
| [`g4f/client/service.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/client/service.py) | Provider mapping via `convert_to_provider` |
| [`g4f/tools/run_tools.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/tools/run_tools.py) | Async tooling pipeline (`async_iter_run_tools`) |
| [`etc/examples/text_completions_demo_async.py`](https://github.com/xtekky/gpt4free/blob/main/etc/examples/text_completions_demo_async.py) | Minimal runnable demonstration |

## Tips and Best Practices

- **Avoid blocking the event loop**: Never call `asyncio.run()` inside an already-running async context (e.g., within a FastAPI endpoint). Use `await` directly instead.
- **Provider fallback**: When using `ClientFactory.create_async_client` with `provider=None`, the library automatically selects the first working provider from the available pool.
- **Custom endpoints**: Pass `base_url` and `api_key` to the factory to create a `CustomProvider` for self-hosted OpenAI-compatible APIs.
- **Tool execution**: The async client automatically routes tool-enabled requests through `async_iter_run_tools`. No additional configuration is required beyond setting appropriate parameters.
- **Memory efficiency**: For long-running streams, process chunks immediately rather than accumulating them in memory to avoid excessive RAM usage.

## Summary

- The `AsyncClient` class in [`g4f/client/__init__.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/client/__init__.py) provides the main entry point for non-blocking LLM interactions.
- Use `await client.chat.completions.create()` for standard requests and `async for chunk in client.chat.completions.stream()` for real-time output.
- `ClientFactory.create_async_client()` enables provider-specific configuration, custom base URLs, and authentication.
- Image generation follows the same pattern via `await client.images.generate()`.
- The architecture relies on `async_iter_run_tools` and `async_iter_response` to normalize provider-specific behavior into standard OpenAI-compatible objects.

## Frequently Asked Questions

### How do I handle streaming responses with the asynchronous client in gpt4free?

Use the `stream` method available on `client.chat.completions`, which returns an async iterator yielding `ChatCompletionChunk` objects. Access incremental content through `chunk.choices[0].delta.content` inside an `async for` loop. This implementation resides in `AsyncCompletions.stream` (lines 682-688 of [`g4f/client/__init__.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/client/__init__.py)).

### Can I use a specific provider instead of the default AnyProvider with AsyncClient?

Yes. Instead of instantiating `AsyncClient` directly, use `ClientFactory.create_async_client()` and pass the provider name as a string (e.g., `"PollinationsAI"` or `"DeepInfra"`) or a custom provider class. This factory method, located at lines 987-1010 in [`g4f/client/__init__.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/client/__init__.py), handles provider resolution, authentication, and custom base URL configuration.

### What is the difference between the response_format options for async image generation?

The `response_format` parameter in `client.images.generate()` accepts three values: `"url"` returns direct URLs to hosted images without local download, `"b64_json"` returns base64-encoded image data suitable for embedding directly in HTML or JSON responses, and `None` (default) typically triggers local file download. This logic is implemented in `AsyncImages.generate` (lines 698-706).

### How does the async client handle tool calling and function execution automatically?

The `AsyncCompletions` class automatically routes requests through `async_iter_run_tools` (defined in [`g4f/tools/run_tools.py`](https://github.com/xtekky/gpt4free/blob/main/g4f/tools/run_tools.py)) when it detects tool configurations. This helper wraps provider-specific capabilities like web search or code execution into an async iterator that chains with `async_iter_response` to normalize output into standard `ChatCompletion` objects. No manual intervention is required beyond setting appropriate parameters in the `create` or `stream` calls.