How FreeLLMAPI Aggregates Free LLM Providers: A Technical Deep Dive

FreeLLMAPI unifies 34 independent free-tier LLM services into a single OpenAI-compatible endpoint through a five-layer architecture involving declarative config sync, provider abstraction, and intelligent quota-aware routing.

The tashfeenahmed/freellmapi repository implements a self-updating aggregation layer that eliminates the complexity of managing multiple API keys and rate limits across disparate providers. By treating free LLM tiers as a unified commodity resource, the system automatically balances load across approximately 635 model endpoints while respecting each provider's specific constraints.

The Five-Layer Aggregation Architecture

FreeLLMAPI's aggregation mechanism relies on tightly coupled components that synchronize state, abstract provider differences, and enforce usage constraints without manual intervention.

1. Declarative Model Catalog Synchronization

The aggregation pipeline begins with server/src/services/declarative-config.ts, which periodically pulls a signed JSON catalog from freellmapi.co twice daily. This declarative approach allows the router to "know" about new models without requiring code changes or redeployment.

The synchronization process follows three steps:

  1. Download: Fetches the latest catalog containing model metadata, provider mappings, quota limits, and known quirks.
  2. Verify: Validates the Ed25519 cryptographic signature to ensure catalog integrity.
  3. Persist: Writes verified entries into a local SQLite database, updating the router's runtime view of available capabilities.

This design decouples the physical deployment from the logical service catalog, enabling the addition of new free providers through configuration alone.

2. Provider Abstraction Layer

Each supported service implements the BaseProvider interface defined in server/src/providers/base.ts. Provider-specific adapters reside in server/src/providers/ (e.g., openai-compat.ts for OpenAI-compatible endpoints), normalizing heterogeneous APIs into a common request/response format.

The server/src/services/router.ts file dynamically loads these adapters using getProvider() and hasProvider() functions. When a request arrives specifying a model ID, the router resolves the appropriate adapter instance. If the primary provider is unavailable, the system consults a fallback chain to route the request to the next viable candidate.

3. Intelligent Routing and Quota Enforcement

Before dispatching any request, the router consults two critical services to prevent rate limit violations and optimize performance:

  • server/src/services/provider-quota.ts: Maintains per-key, per-model, per-provider counters for daily request caps (RPD), per-minute caps (RPM), and token-based limits (TPD/TPM). When a provider's quota exhausts, the router immediately skips that provider and attempts the next in the chain.
  • server/src/services/ratelimit.ts: Enforces cooldown periods and sliding window rate limits across the entire provider pool.

The system also incorporates a speed/intelligence scoring mechanism (server/src/services/scoring.ts) that ranks providers based on latency and output quality. When a user specifies model: "auto", the router selects the highest-scoring provider with available quota.

4. Custom Endpoint Discovery

FreeLLMAPI extends aggregation to user-supplied infrastructure through server/src/services/model-discovery.ts. The discoverEndpointModels() function enables integration with arbitrary OpenAI-compatible endpoints such as Ollama or LM Studio.

When registering a custom endpoint, the system issues a GET /v1/models request to the target URL, parses the returned model catalog, and registers discovered models in the SQLite database. This treats local or private endpoints as first-class citizens within the same routing and quota framework as commercial providers.

5. Request Handling and Transparency

The main HTTP handler in server/src/services/router.ts executes the final aggregation step:

  1. Decrypt the stored provider API key in-memory (keys are never logged)
  2. Build the provider-specific request payload
  3. Forward the request to the selected provider
  4. Inject the X-Routed-Via response header to identify which provider actually served the request

This transparency allows clients to audit routing decisions and debug provider-specific behavior while maintaining the unified interface.

Implementation Walkthrough

The following examples demonstrate how to interact with the aggregated endpoint and extend it with custom providers.

Querying the Aggregated API

Use any OpenAI-compatible client to access the unified endpoint. The auto model selection delegates routing to the scoring and quota systems:

import { OpenAI } from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",
  apiKey: "freellmapi-your-unified-key",
});

const resp = await client.chat.completions.create({
  model: "auto",
  messages: [{ 
    role: "user", 
    content: "Explain quantum tunnelling in one sentence." 
  }],
});

console.log(resp.choices[0].message.content);
console.log("Provider:", resp.headers.get("x-routed-via"));

Verifying Provider Selection

To inspect which provider handled a specific request without writing code, use the HTTP interface and check the response headers:

curl -X POST http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","messages":[{"role":"user","content":"Summarise the plot of Romeo & Juliet."}]}' \
  -v

The X-Routed-Via header reveals the actual provider (e.g., groq, anthropic, ollama-local) that served the request.

Adding Custom OpenAI-Compatible Endpoints

Integrate local or private endpoints by triggering the discovery mechanism:

await fetch("http://localhost:3001/api/keys/custom/discover-models", {
  method: "POST",
  headers: { 
    "Authorization": "Bearer freellmapi-your-unified-key", 
    "Content-Type": "application/json" 
  },
  body: JSON.stringify({ 
    baseUrl: "http://localhost:11434/v1", 
    apiKey: "" 
  })
});

After discovery, models from this endpoint participate in the same fallback chains and quota tracking as managed providers.

Summary

FreeLLMAPI aggregates free LLM providers through a robust pipeline that ensures reliability and transparency:

Frequently Asked Questions

How does FreeLLMAPI handle provider failures?

When a provider returns an error or times out, the server/src/services/router.ts implementation immediately attempts the next provider in the fallback chain. This retry logic executes within the same request lifecycle, ensuring clients receive a valid response from an alternative provider without manual intervention. The system marks exhausted providers as unavailable for subsequent requests until their quotas reset or health checks pass.

What quota limits does FreeLLMAPI enforce?

The aggregation layer enforces four quota dimensions tracked in server/src/services/provider-quota.ts: requests per day (RPD), requests per minute (RPM), tokens per day (TPD), and tokens per minute (TPM). Each counter operates at the granularity of API key, model, and provider. When any limit exhausts, the router automatically excludes that provider from the candidate pool for the current request.

Can I use my own local LLM with FreeLLMAPI?

Yes. The discoverEndpointModels() function in server/src/services/model-discovery.ts supports arbitrary OpenAI-compatible endpoints. Register your local Ollama, LM Studio, or vLLM instance by providing its base URL; the system queries /v1/models, registers discovered models in the SQLite database, and applies the same routing, scoring, and fallback logic used for commercial providers.

How often is the model catalog updated?

The server/src/services/declarative-config.ts service synchronizes with freellmapi.co every 12 hours. After each sync, it verifies the Ed25519 signature of the downloaded catalog before updating the local SQLite database. This schedule balances freshness with system stability, ensuring new free tiers appear automatically without requiring service restarts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →