# How FreeLLMAPI Aggregates Free LLM Providers: A Technical Deep Dive

> Discover how FreeLLMAPI aggregates 34 free LLM providers into one OpenAI-compatible endpoint using a five-layer architecture for unified access and intelligent routing.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-09-04

---

**FreeLLMAPI unifies 34 independent free-tier LLM services into a single OpenAI-compatible endpoint through a five-layer architecture involving declarative config sync, provider abstraction, and intelligent quota-aware routing.**

The `tashfeenahmed/freellmapi` repository implements a self-updating aggregation layer that eliminates the complexity of managing multiple API keys and rate limits across disparate providers. By treating free LLM tiers as a unified commodity resource, the system automatically balances load across approximately 635 model endpoints while respecting each provider's specific constraints.

## The Five-Layer Aggregation Architecture

FreeLLMAPI's aggregation mechanism relies on tightly coupled components that synchronize state, abstract provider differences, and enforce usage constraints without manual intervention.

### 1. Declarative Model Catalog Synchronization

The aggregation pipeline begins with [`server/src/services/declarative-config.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/declarative-config.ts), which periodically pulls a signed JSON catalog from **freellmapi.co** twice daily. This declarative approach allows the router to "know" about new models without requiring code changes or redeployment.

The synchronization process follows three steps:

1. **Download**: Fetches the latest catalog containing model metadata, provider mappings, quota limits, and known quirks.
2. **Verify**: Validates the Ed25519 cryptographic signature to ensure catalog integrity.
3. **Persist**: Writes verified entries into a local SQLite database, updating the router's runtime view of available capabilities.

This design decouples the physical deployment from the logical service catalog, enabling the addition of new free providers through configuration alone.

### 2. Provider Abstraction Layer

Each supported service implements the **BaseProvider** interface defined in [`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts). Provider-specific adapters reside in `server/src/providers/` (e.g., [`openai-compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/openai-compat.ts) for OpenAI-compatible endpoints), normalizing heterogeneous APIs into a common request/response format.

The [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) file dynamically loads these adapters using `getProvider()` and `hasProvider()` functions. When a request arrives specifying a model ID, the router resolves the appropriate adapter instance. If the primary provider is unavailable, the system consults a **fallback chain** to route the request to the next viable candidate.

### 3. Intelligent Routing and Quota Enforcement

Before dispatching any request, the router consults two critical services to prevent rate limit violations and optimize performance:

- **[`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts)**: Maintains per-key, per-model, per-provider counters for daily request caps (RPD), per-minute caps (RPM), and token-based limits (TPD/TPM). When a provider's quota exhausts, the router immediately skips that provider and attempts the next in the chain.
- **[`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts)**: Enforces cooldown periods and sliding window rate limits across the entire provider pool.

The system also incorporates a **speed/intelligence scoring** mechanism ([`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts)) that ranks providers based on latency and output quality. When a user specifies `model: "auto"`, the router selects the highest-scoring provider with available quota.

### 4. Custom Endpoint Discovery

FreeLLMAPI extends aggregation to user-supplied infrastructure through [`server/src/services/model-discovery.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-discovery.ts). The `discoverEndpointModels()` function enables integration with arbitrary OpenAI-compatible endpoints such as Ollama or LM Studio.

When registering a custom endpoint, the system issues a **GET /v1/models** request to the target URL, parses the returned model catalog, and registers discovered models in the SQLite database. This treats local or private endpoints as first-class citizens within the same routing and quota framework as commercial providers.

### 5. Request Handling and Transparency

The main HTTP handler in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) executes the final aggregation step:

1. **Decrypt** the stored provider API key in-memory (keys are never logged)
2. **Build** the provider-specific request payload
3. **Forward** the request to the selected provider
4. **Inject** the `X-Routed-Via` response header to identify which provider actually served the request

This transparency allows clients to audit routing decisions and debug provider-specific behavior while maintaining the unified interface.

## Implementation Walkthrough

The following examples demonstrate how to interact with the aggregated endpoint and extend it with custom providers.

### Querying the Aggregated API

Use any OpenAI-compatible client to access the unified endpoint. The `auto` model selection delegates routing to the scoring and quota systems:

```typescript
import { OpenAI } from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",
  apiKey: "freellmapi-your-unified-key",
});

const resp = await client.chat.completions.create({
  model: "auto",
  messages: [{ 
    role: "user", 
    content: "Explain quantum tunnelling in one sentence." 
  }],
});

console.log(resp.choices[0].message.content);
console.log("Provider:", resp.headers.get("x-routed-via"));

```

### Verifying Provider Selection

To inspect which provider handled a specific request without writing code, use the HTTP interface and check the response headers:

```bash
curl -X POST http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","messages":[{"role":"user","content":"Summarise the plot of Romeo & Juliet."}]}' \
  -v

```

The `X-Routed-Via` header reveals the actual provider (e.g., `groq`, `anthropic`, `ollama-local`) that served the request.

### Adding Custom OpenAI-Compatible Endpoints

Integrate local or private endpoints by triggering the discovery mechanism:

```typescript
await fetch("http://localhost:3001/api/keys/custom/discover-models", {
  method: "POST",
  headers: { 
    "Authorization": "Bearer freellmapi-your-unified-key", 
    "Content-Type": "application/json" 
  },
  body: JSON.stringify({ 
    baseUrl: "http://localhost:11434/v1", 
    apiKey: "" 
  })
});

```

After discovery, models from this endpoint participate in the same fallback chains and quota tracking as managed providers.

## Summary

FreeLLMAPI aggregates free LLM providers through a robust pipeline that ensures reliability and transparency:

- **Self-updating catalogs** synchronize model availability via cryptographically signed JSON feeds in [`server/src/services/declarative-config.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/declarative-config.ts).
- **Provider abstraction** normalizes disparate APIs through the `BaseProvider` interface and dynamic adapter loading in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts).
- **Quota enforcement** prevents 429/402 errors through [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts) and [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts), tracking RPD, RPM, TPD, and TPM limits per key.
- **Intelligent selection** uses speed/intelligence scoring in [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) to optimize the `auto` model routing.
- **Extensible discovery** incorporates custom endpoints via [`server/src/services/model-discovery.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-discovery.ts), querying their `/v1/models` endpoints automatically.
- **Request transparency** via the `X-Routed-Via` header allows clients to trace which provider served each request.

## Frequently Asked Questions

### How does FreeLLMAPI handle provider failures?

When a provider returns an error or times out, the [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) implementation immediately attempts the next provider in the fallback chain. This retry logic executes within the same request lifecycle, ensuring clients receive a valid response from an alternative provider without manual intervention. The system marks exhausted providers as unavailable for subsequent requests until their quotas reset or health checks pass.

### What quota limits does FreeLLMAPI enforce?

The aggregation layer enforces four quota dimensions tracked in [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts): **requests per day (RPD)**, **requests per minute (RPM)**, **tokens per day (TPD)**, and **tokens per minute (TPM)**. Each counter operates at the granularity of API key, model, and provider. When any limit exhausts, the router automatically excludes that provider from the candidate pool for the current request.

### Can I use my own local LLM with FreeLLMAPI?

Yes. The `discoverEndpointModels()` function in [`server/src/services/model-discovery.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-discovery.ts) supports arbitrary OpenAI-compatible endpoints. Register your local Ollama, LM Studio, or vLLM instance by providing its base URL; the system queries `/v1/models`, registers discovered models in the SQLite database, and applies the same routing, scoring, and fallback logic used for commercial providers.

### How often is the model catalog updated?

The [`server/src/services/declarative-config.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/declarative-config.ts) service synchronizes with freellmapi.co every 12 hours. After each sync, it verifies the Ed25519 signature of the downloaded catalog before updating the local SQLite database. This schedule balances freshness with system stability, ensuring new free tiers appear automatically without requiring service restarts.