What Is the Interference API in gpt4free? Complete Technical Guide

The Interference API in gpt4free is an OpenAI-compatible REST server built on FastAPI that exposes the library's multi-provider backend through standard HTTP endpoints like /v1/chat/completions, enabling drop-in replacement for official OpenAI services.

The xtekky/gpt4free repository aggregates access to multiple AI providers through a unified asynchronous client. The Interference API serves as the HTTP façade that transforms this internal AsyncClient into a production-ready web service, allowing any application using standard OpenAI client libraries to route requests through gpt4free's provider-agnostic pipeline without modifying existing code.

Core Architecture and Implementation Files

The Interference API implementation resides primarily in g4f/api/__init__.py, which constructs the FastAPI application, registers route handlers, and performs request validation. This module creates a lightweight HTTP server that forwards every validated request to the internal AsyncClient (defined in g4f/client_async.py) for provider resolution and response generation.

The bootstrap entry point is located in g4f/api/run.py, a minimal wrapper that invokes g4f.api.run_api() to initialize the server. Provider implementations available for routing are located in g4f/Provider/, and the selection logic automatically applies based on incoming request parameters or global configuration.

OpenAI-Compatible Endpoint Specifications

The API mirrors OpenAI's REST contract exactly, implementing the same path structures and JSON schemas. The Api.chat_completions method (defined around lines 45-48 in g4f/api/__init__.py) handles POST requests to /v1/chat/completions, while Api.models (lines 126-152) serves GET requests to /v1/models returning consolidated provider model lists.

Provider-agnostic routing occurs automatically for every call. The system selects a concrete provider (e.g., Perplexity, Gemini, or local models) based on either the request's provider field or the global AppConfig.provider setting. This routing logic lives in g4f/Provider/ and requires no code changes when new providers are added.

Authentication and Configuration Modes

The Interference API supports two distinct authentication patterns controlled via AppConfig.

Demo mode: When AppConfig.demo is set to true, the server runs without requiring an API key, returning a demo-user context for all incoming requests. This is useful for local development and testing.

Key-based authentication: For production deployments, set g4f.config.g4f_api_key. When configured, the API requires the g4f-api-key header on all requests, rejecting unauthorized calls with standard HTTP 401 responses.

Streaming Responses and Media Support

The API supports both synchronous JSON responses and Server-Sent Events (SSE) streaming. When stream: true is passed in the request body, the Api.chat_completions method (around lines 223-237) returns a StreamingResponse object that yields content chunks as SSE data: events.

Beyond text generation, the API handles multimedia through dedicated routes. The Api.generate_image method (lines 557-587) processes /v1/media/generate and /v1/images/generations endpoints, forwarding requests to image-specific providers and returning generated URLs. Additional routes support audio transcription, speech synthesis, and arbitrary file upload/download operations.

Starting the Interference API Server

You can start the server via command line or programmatically.

CLI startup (most common):

python -m g4f --port 8080 --debug

The --debug flag enables request logging through the FastAPI stack.

Programmatic startup:

import g4f.api
g4f.api.run_api(debug=True)

This pattern is used internally by g4f/api/run.py when executing the module entry point.

Integration Examples

Standard Chat Completions with the OpenAI SDK

Point the official OpenAI client at your local Interference API:

import openai

openai.api_base = "http://localhost:8080/v1"
openai.api_key = "any-value-if-key-required"

resp = openai.ChatCompletion.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain the Interference API"}],
    stream=False
)

print(resp["choices"][0]["message"]["content"])

This request hits the /v1/chat/completions route in g4f/api/__init__.py and processes through Api.chat_completions.

Streaming Responses via Server-Sent Events

import requests

resp = requests.post(
    "http://localhost:8080/v1/chat/completions",
    json={
        "model": "gpt-4o-mini",
        "messages": [{"role": "user", "content": "Stream this response"}],
        "stream": True
    },
    stream=True
)

for line in resp.iter_lines():
    if line:
        print(line.decode())

The endpoint returns a StreamingResponse where each chunk is yielded as an SSE event.

Image Generation Endpoints

import openai

openai.api_base = "http://localhost:8080/v1"
openai.api_key = "dummy"

result = openai.Image.create(
    model="flux",
    prompt="a futuristic cityscape at sunset",
    n=1,
    size="1024x1024"
)

print(result["data"][0]["url"])

This maps to /v1/media/generate and is handled by Api.generate_image in the main API module.

Raw HTTP Requests

List available models via direct HTTP:

curl -X POST http://localhost:8080/v1/models

This calls Api.models() and returns the aggregated model list from all configured providers.

Summary

  • The Interference API is implemented in g4f/api/__init__.py as a FastAPI application that wraps the internal AsyncClient.
  • It provides OpenAI-compatible endpoints including /v1/chat/completions, /v1/models, and media routes, enabling drop-in SDK replacement.
  • Provider routing is automatic based on the provider request field or global AppConfig.provider settings.
  • Authentication is optional in demo mode (AppConfig.demo = true) or enforced via the g4f-api-key header when g4f.config.g4f_api_key is configured.
  • Streaming is supported through Server-Sent Events (StreamingResponse), and media generation (images, audio) is available via dedicated endpoints.

Frequently Asked Questions

What is the Interference API in gpt4free used for?

The Interference API provides a standards-compliant HTTP interface that exposes gpt4free's multi-provider backend. It allows developers to use existing OpenAI client libraries and tools while routing traffic through gpt4free's free provider ecosystem, effectively creating a drop-in alternative to OpenAI's commercial API.

How do I start the Interference API server locally?

Execute python -m g4f --port 8080 --debug from your terminal. This command invokes g4f.api.run_api() via the bootstrap file g4f/api/run.py, starting the FastAPI server on the specified port with optional debug logging enabled.

Does the Interference API support real-time streaming responses?

Yes. When you set stream: true in your request payload to /v1/chat/completions, the API returns a StreamingResponse (implemented around lines 223-237 in g4f/api/__init__.py) that yields content via Server-Sent Events (SSE), compatible with OpenAI's streaming protocol.

Which source files contain the core API implementation?

The main FastAPI application logic, route definitions, and request handling reside in g4f/api/__init__.py. The bootstrap entry point is g4f/api/run.py. Underlying client functionality used by the API is located in g4f/client_async.py and g4f/client.py, while available providers are defined in the g4f/Provider/ directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →