Understanding Model Definitions in gpt4free: Architecture and Usage Guide

Model definitions in gpt4free are centralized dataclass instances stored in g4f/models.py that map model names to their base providers and preferred implementation providers through an auto-registering global registry system.

According to the xtekky/gpt4free source code, the library maintains a comprehensive catalog of large language, image, audio, and vision models through a structured definition system. These model definitions serve as the single source of truth for provider resolution, enabling the library to route requests to the appropriate backend automatically. All model metadata lives in a single file and leverages Python dataclasses for type-safe configuration management.

Model Definition Dataclass Architecture

At the heart of the system is the Model dataclass defined in g4f/models.py. This class encapsulates the metadata required to identify and route requests for every supported AI model.

Core Model Class

The base Model class uses the @dataclass(unsafe_hash=True) decorator and defines four critical fields:

@dataclass(unsafe_hash=True)
class Model:
    name: str
    base_provider: str
    best_provider: ProviderType = None
    long_name: Optional[str] = None
  • name: The canonical model identifier (e.g., "gpt-4", "gemini-2.0")
  • base_provider: The originating service (e.g., "OpenAI", "Google", "Meta")
  • best_provider: The preferred implementation, often wrapped in an IterListProvider for automatic fallback cycling
  • long_name: Optional extended identifier for publishing contexts

Specialized Model Subclasses

The architecture supports modality-specific variants that inherit from Model but signal different capabilities:

  • ImageModel: For text-to-image generators like DALL-E 3 and Stable Diffusion
  • VisionModel: For multimodal models that process images and text (e.g., GPT-4o vision)
  • AudioModel: For speech synthesis and recognition models
  • VideoModel: For video generation capabilities

These subclasses contain identical fields but enable the registry to filter models by capability type.

Global Model Registry System

The ModelRegistry class implements a global singleton pattern that maintains two internal dictionaries: _models (mapping canonical names to instances) and _aliases (mapping shortcuts to canonical names).

Automatic Registration Mechanism

Every Model instance auto-registers upon instantiation through the __post_init__ hook:

def __post_init__(self):
    """Auto-register model after initialization"""
    if self.name:
        ModelRegistry.register(self)

This design ensures that simply importing g4f/models.py populates the entire catalog without manual registration calls. When the module loads, each Model(...) declaration executes, triggering registration in the global namespace.

Registry API Methods

The registry exposes several authority methods for model retrieval:

  • register(model, aliases=None): Adds a model with optional name aliases
  • get(name): Resolves canonical names or aliases to the underlying Model instance
  • all_models(): Returns a shallow copy of the complete _models dictionary
  • list_models_by_provider(provider_name): Filters the catalog by base provider (e.g., "Cloudflare" or "OpenAI")
  • validate_all_models(): Performs sanity checks for missing required fields, returning a dictionary of validation errors

ModelUtils Convenience API

While ModelRegistry handles backend storage, the ModelUtils class provides the public API surface that most client code interacts with. This utility class maintains a static convert dictionary that mirrors the registry for fast lookups.

Key utility methods include:

  • refresh(): Synchronizes ModelUtils.convert with the current state of ModelRegistry
  • get_model(name): Wrapper around ModelRegistry.get() that returns Optional[Model]
  • register_alias(alias, model_name): Runtime alias creation for custom shortcuts

The convert dictionary serves as the primary interface for the rest of the codebase, including g4f/client/service.py which routes API calls based on these definitions.

Complete Model Catalog Overview

The g4f/models.py file defines approximately 300 model instances across multiple families. The catalog organizes models by their base_provider field and typically assigns an IterListProvider to the best_provider field for resilience through provider cycling.

Major Model Families

OpenAI Models

  • Chat: gpt_4, gpt_4o, o1, o1_mini (base provider: "OpenAI")
  • Vision: gpt_4o as VisionModel (preferred providers: OpenaiChat via IterListProvider)
  • Image: dall_e_3, gpt_image (preferred providers: CopilotAccount and others)

Meta Llama Family

  • llama_2_7b, llama_3_70b, llama_4_maverick (base provider: "Meta")
  • Best providers vary by model: Together, Cloudflare, or HuggingChat

Google DeepMind (Gemini)

  • gemini, gemini_2_5_flash (base provider: "Google")
  • Aliases map "gemini" to "gemini-2.0" for convenience
  • Preferred providers: Gemini, GeminiPro, GeminiCLI

Specialized Providers

  • Mistral AI: mistral_7b, mixtral_8x7b (often routed through Together)
  • DeepSeek: deepseek_v3, deepseek_r1 with fallback providers
  • Stability AI: sdxl_turbo, sd_3_5_large (ImageModel instances)
  • Black Forest Labs: flux, flux_pro, flux_dev for image generation
  • x.ai: grok_2, grok_3 (base provider: "x.ai")
  • Perplexity AI: sonar, sonar_pro (routed through PuterJS)

Working with Model Definitions

The registry system enables runtime model discovery and manipulation without hardcoding provider logic.

Retrieve a Specific Model

Access any model definition through the utility API:

from g4f.models import ModelUtils

# Resolve by canonical name

gpt4 = ModelUtils.get_model("gpt-4")
print(gpt4.base_provider)  # Output: OpenAI

# Resolve by alias

gemini = ModelUtils.get_model("gemini")
print(gemini.name)  # Output: gemini-2.0

List and Filter Available Models

Enumerate the complete catalog or filter by infrastructure provider:

from g4f.models import ModelUtils, ModelRegistry

# List all model identifiers

all_ids = list(ModelUtils.convert.keys())
print(f"Total available models: {len(all_ids)}")

# Filter by specific provider implementation

cf_models = ModelRegistry.list_models_by_provider("Cloudflare")
print(f"Cloudflare-backed models: {cf_models}")

Register Custom Aliases

Extend the catalog at runtime for application-specific naming conventions:

from g4f.models import ModelUtils, ModelRegistry

ModelUtils.register_alias("production-llm", "gpt-4")
assert ModelRegistry.get("production-llm") is ModelRegistry.get("gpt-4")

Validate Registry Integrity

Debug model configurations using the built-in validator:

from g4f.models import ModelRegistry

issues = ModelRegistry.validate_all_models()
if issues:
    for model_name, problems in issues.items():
        print(f"{model_name}: {', '.join(problems)}")
else:
    print("All model definitions are valid.")

Summary

  • Model definitions in xtekky/gpt4free are stored as dataclass instances in g4f/models.py, encapsulating model names, base providers, and best provider implementations.
  • The ModelRegistry provides global auto-registration through __post_init__, maintaining mappings for both canonical names and aliases.
  • ModelUtils offers a convenience API with static dictionaries and helper methods for runtime model retrieval.
  • The catalog supports 300+ models across text, image, audio, and vision modalities, with specialized subclasses like ImageModel and VisionModel.
  • Provider resolution uses IterListProvider for automatic fallback cycling when multiple backends support the same model.

Frequently Asked Questions

How are model definitions registered automatically in gpt4free?

The Model dataclass implements a __post_init__ method that calls ModelRegistry.register(self) immediately after instantiation. When g4f/models.py imports, Python executes the module-level Model(...) calls, triggering automatic registration without explicit initialization code.

What is the difference between base_provider and best_provider in model definitions?

The base_provider field identifies the original model host (e.g., "OpenAI" or "Meta"), while the best_provider field specifies the concrete implementation class that gpt4free will use to execute requests. The best provider often wraps multiple options in an IterListProvider to enable automatic fallback if one provider fails.

Can I add custom models to the gpt4free registry at runtime?

Yes. Instantiate a new Model or subclass (e.g., ImageModel) with your configuration, which auto-registers via __post_init__. Alternatively, use ModelUtils.register_alias() to create shortcuts to existing models without defining new instances.

How do I find which provider will handle a specific model request?

Query the model definition through ModelUtils.get_model("model-name") and inspect the best_provider attribute. If it contains an IterListProvider, the library cycles through its internal list; otherwise, it uses the single provider class directly. For filtering by infrastructure type, use ModelRegistry.list_models_by_provider("ProviderName").

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →