What Is the gpt4free FastAPI Integration? A Complete Technical Guide
The gpt4free FastAPI integration transforms the Python library into a production-ready HTTP API server that exposes OpenAI-compatible endpoints for chat completions, model listing, and media generation through a configurable FastAPI application.
The gpt4free repository by xtekky provides a unified interface to multiple AI providers, and its FastAPI integration allows you to run this functionality as a standalone web service. This integration converts the asynchronous client into RESTful endpoints that mirror the OpenAI API specification, enabling seamless integration with existing tools and frontend applications. Whether deploying locally, containerizing, or running behind a reverse proxy, the FastAPI layer provides the necessary HTTP interface to access the library's provider-agnostic capabilities.
Core Architecture of the gpt4free FastAPI Integration
The integration centers on a factory pattern that constructs the FastAPI application with proper lifecycle management, security middleware, and comprehensive route registration.
Application Factory and Lifespan Management
The entry point create_app() in g4f/api/__init__.py initializes the FastAPI instance with a custom lifespan context manager that handles startup and shutdown procedures. According to the source code at lines 16-25, this lifespan function loads cookie files required for provider authentication and ensures browser locks are cleaned up on shutdown, preventing resource leaks when using browser-based providers.
The module exposes multiple factory variants for different deployment scenarios:
create_app()– Standard production configurationcreate_app_debug()– Development mode with enhanced loggingcreate_app_with_gui_and_debug()– Includes the built-in HTML interfacecreate_app_with_demo_and_debug()– Enables demo mode bypass for testing
These factories read runtime parameters from the AppConfig class defined in g4f/config.py, which sources configuration from environment variables or programmatic settings at lines 37-65.
CORS and Middleware Configuration
Before route registration, the integration applies CORS middleware to allow cross-origin requests from any domain. As implemented in g4f/api/__init__.py at lines 20-27, this middleware configuration enables frontend applications hosted on different origins to communicate with the API without proxy complications, making it suitable for browser-based GUI consumption.
Route Registration and Endpoint Structure
The Api.register_routes() method at lines 31-52 mounts the core HTTP endpoints that provide the public interface:
/v1/models– Returns a JSON list of all supported models and their associated providers/v1/chat/completions– The primary chat endpoint that mirrors OpenAI's completion specification, accepting model names, message arrays, and streaming parameters/v1/media/generate– Handles image, video, and audio generation requests through compatible providers/api/{provider}/models– Provider-specific model enumeration for direct provider access/api/{provider}/quota– Exposes quota and rate limit information for specific providers/docs– Interactive Swagger UI for API exploration and testing
These routes forward incoming HTTP requests to the underlying async client defined in g4f/client/*.py, which then dispatches to the appropriate provider adapter in g4f/providers/.
Authentication and Authorization Middleware
Security is implemented through Api.register_authorization() at lines 22-32, which supports multiple authentication strategies:
- G4F API Key – Optional global API key validation via environment variables
- Encrypted Tokens – Support for token-based authentication with encryption
- Demo Mode – Configurable bypass for development and testing environments
The middleware intercepts requests before they reach the endpoint handlers, validating credentials against the configuration before allowing access to provider resources.
Error Handling and Validation
The integration overrides FastAPI's default validation errors with a custom exception handler defined at lines 95-108 in g4f/api/__init__.py. This handler normalizes Pydantic validation messages into a consistent JSON format, ensuring that clients receive predictable error structures when submitting malformed requests or invalid model parameters.
Configuration and Runtime Settings
Server behavior is controlled through the AppConfig class in g4f/config.py. Key configuration parameters include:
- Port – Default 1337, configurable via
AppConfig.DEFAULT_PORT - Timeout – Request timeout thresholds for provider responses
- Default Model – Fallback model when none is specified in requests
- GUI Flag – Boolean enabling the HTML interface when true
Configuration can be set programmatically or through environment variables, allowing containerized deployments to inject settings without code modification.
Running the FastAPI Server
Deploying the API requires installing optional FastAPI dependencies and invoking the application factory through an ASGI server.
Installation and Startup
Install the required dependencies and start the server using uvicorn:
# Install FastAPI extras
pip install "gpt4free[fastapi]"
# Start the production server
uvicorn g4f.api:create_app --host 0.0.0.0 --port 1337
For development with auto-reload and debug output:
uvicorn g4f.api:create_app_debug --host 0.0.0.0 --port 1337 --reload
Enabling the Web GUI
To serve the built-in HTML interface alongside the API endpoints, install the GUI extras and use the combined factory:
pip install "gpt4free[gui]"
uvicorn g4f.api:create_app_with_gui_and_debug --host 0.0.0.0 --port 1337
When AppConfig.gui evaluates to true, the integration mounts the GUI application using a2wsgi at the root path /, as referenced in the conditional block at lines 35-40 of g4f/api/__init__.py.
API Endpoints and Usage Examples
The exposed endpoints follow RESTful conventions and OpenAI API compatibility standards.
OpenAI-Compatible Chat Completions
The /v1/chat/completions endpoint accepts POST requests with JSON payloads matching the OpenAI schema:
curl -X POST http://localhost:1337/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-api-key" \
-d '{
"model": "openai/gpt-4",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain the gpt4free FastAPI integration"}
],
"stream": false,
"temperature": 0.7
}'
Responses conform to the standard OpenAI completion format, including id, object, created timestamp, model identifier, and choices array containing message content.
Model Listing and Provider-Specific Routes
Retrieve available models through the standard or provider-specific endpoints:
# List all supported models across all providers
curl http://localhost:1337/v1/models
# List models for a specific provider
curl http://localhost:1337/api/openai/models
Media Generation Endpoints
For image or video generation, use the media endpoint with appropriate model specifications:
curl -X POST http://localhost:1337/v1/media/generate \
-H "Content-Type: application/json" \
-d '{
"model": "openai/dall-e-3",
"prompt": "A diagram showing FastAPI architecture",
"size": "1024x1024"
}'
Python Client Integration
The g4f Python client can target the local FastAPI instance by specifying the base_url parameter:
import asyncio
import g4f
async def query_local_api():
client = g4f.AsyncClient(base_url="http://localhost:1337")
response = await client.ChatCompletion.create(
model="openai/gpt-4",
messages=[{"role": "user", "content": "What is FastAPI?"}],
temperature=0.5
)
print(response.choices[0].message.content)
asyncio.run(query_local_api())
This approach allows existing codebases using the g4f client library to seamlessly switch between direct provider calls and the HTTP API without changing request structures.
Summary
The gpt4free FastAPI integration provides a comprehensive HTTP layer for the library's AI provider aggregation capabilities:
- Factory functions like
create_app()ing4f/api/__init__.pygenerate configurable FastAPI instances with proper lifespan management for resource cleanup - OpenAI-compatible endpoints at
/v1/chat/completionsand/v1/modelsenable drop-in replacement for existing OpenAI API consumers - Flexible authentication through
Api.register_authorization()supports API keys, encrypted tokens, and demo modes - Runtime configuration via
AppConfiging4f/config.pyallows environment-based deployment customization - Optional GUI mounting through
a2wsgiprovides a visual interface when runningcreate_app_with_gui_and_debug() - Comprehensive error handling normalizes validation errors for consistent client-side error management
Frequently Asked Questions
How do I secure the gpt4free FastAPI server in production?
Set the G4F_API_KEY environment variable before starting the server to enable API key validation through the authorization middleware defined in g4f/api/__init__.py. The Api.register_authorization() function checks this key against the Authorization header on every request. For additional security, run the server behind a reverse proxy with HTTPS termination and restrict direct port access using firewall rules.
Can I use the gpt4free FastAPI integration without installing the GUI dependencies?
Yes. The GUI components are entirely optional. Install only the FastAPI extras using pip install "gpt4free[fastapi]" and use create_app() or create_app_debug() instead of the GUI-enabled factory functions. The application will expose all API endpoints without mounting the HTML interface, reducing the dependency footprint.
What is the difference between the /v1/chat/completions and provider-specific endpoints?
The /v1/chat/completions endpoint follows the OpenAI API specification and automatically routes requests to appropriate providers based on the model parameter, while provider-specific paths like /api/{provider}/quota expose direct provider metadata and quota information. Use the standard endpoint for OpenAI compatibility and provider-specific routes when you need to query individual provider capabilities or bypass the automatic provider selection logic.
How does the gpt4free FastAPI server handle streaming responses?
The /v1/chat/completions endpoint supports streaming by setting "stream": true in the request payload. When enabled, the server returns Server-Sent Events (SSE) formatted according to the OpenAI streaming specification, transmitting response tokens as they are generated by the underlying provider rather than buffering the complete response. This behavior is implemented in the route handlers registered by Api.register_routes() and maintains compatibility with standard OpenAI streaming clients.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →