# Core Components of the FreeLLMAPI Architecture: Technical Deep Dive

> Explore the core components of the FreeLLMAPI architecture including its Express proxy router rate limiting and secure key management for self-hosted LLMs.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-09-02

---

**FreeLLMAPI is a self-hosted, OpenAI-compatible gateway that aggregates free LLM tiers through a modular architecture comprising an Express proxy layer, intelligent router with Thompson-sampling bandit scoring, provider-specific adapters, SQLite-backed rate limiting, and AES-256-GCM encrypted key management.**

FreeLLMAPI serves as a unified API layer that transforms dozens of disparate free-tier LLM providers into a single OpenAI-compatible endpoint. Understanding the FreeLLMAPI architecture reveals how it handles request routing, provider failover, and rate-limit compliance while maintaining security and performance. The system is implemented in TypeScript/Node.js and exposes standard endpoints like `/v1/chat/completions` while managing complex internal state across approximately 635 model endpoints.

## Express Proxy and API Layer

The **Express Proxy** acts as the entry point for all client traffic, exposing familiar OpenAI-style endpoints including `/v1/chat/completions`, `/v1/embeddings`, and `/v1/models`. According to the `tashfeenahmed/freellmapi` source code, this layer is implemented in [`server/src/app.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/app.ts) and handles HTTP request forwarding to the internal router.

The proxy maintains strict compatibility with the OpenAI API specification, allowing existing clients to migrate by simply changing the `basePath` configuration while using a unified API key format (`freellmapi-…`).

## Intelligent Request Routing

The **Router** component in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) constitutes the decision-making core of the FreeLLMAPI architecture. It implements the `routeRequest` function, which performs several critical operations:

- Resolves user-requested model IDs to concrete catalog entries via `resolveRequestedIdForDispatch`
- Loads fallback chains using `getActiveProfileId` and `loadChainRows`
- Selects the highest-scoring viable provider through iterative evaluation
- Decrypts provider keys in-memory using `decrypt` from [`server/src/lib/crypto.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/crypto.ts)

The router implements sophisticated **fallback-chain logic** with cooldown handling and per-key diagnostics. If all models in a chain are exhausted, it throws a `RouteError` with status 429 and detailed diagnostics.

### Model Scoring and Selection

Model selection relies on **Thompson-sampling bandit algorithms** implemented in [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts). This system scores models based on reliability, speed, intelligence metrics, and current quota headroom to optimize routing decisions dynamically.

## Provider Adapter System

FreeLLMAPI abstracts vendor-specific SDKs through a standardized **Provider Adapter** pattern. The `BaseProvider` interface in [`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts) defines common methods including `chatCompletion()` and `streamChatCompletion()`.

Individual adapters (such as [`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts) for Google, and similar files for Groq and OpenRouter) encapsulate vendor-specific API calls and error mapping. This architecture allows the router to treat diverse providers uniformly while handling provider-specific quirks internally.

## Rate Limiting and Quota Management

The **Rate-Limit Ledger** in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) implements a sliding-window algorithm backed by SQLite to track RPM (requests per minute), RPD (requests per day), TPM (tokens per minute), and TPD (tokens per day) counters.

Key features include:

- Granular tracking for each `(platform, model, key)` tuple
- Concurrency "leases" to prevent check-then-act race conditions
- Atomic operations ensuring accurate quota enforcement across concurrent requests

The **Health and Quota Service** in [`server/src/services/health.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/health.ts) complements this by periodically probing key health and deriving per-provider quota headroom, updating cooldown status accordingly.

## Model Discovery and Catalog Management

The system maintains a signed, self-updating **Model Catalog** covering approximately 635 free model endpoints. The discovery logic in [`server/src/services/model-discovery.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-discovery.ts) synchronizes available models with the scoring system to ensure the router only considers viable candidates.

## Security and Encryption Architecture

FreeLLMAPI implements defense-in-depth for credential protection:

- **AES-256-GCM encryption** for provider API keys at rest, implemented in [`server/src/lib/crypto.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/crypto.ts)
- **Key Storage** in SQLite with encrypted fields
- **In-memory decryption** per request via [`server/src/lib/key-proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/key-proxy.ts), ensuring keys exist in plaintext only during active request processing

Backup operations in [`server/src/lib/db-backup.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/db-backup.ts) maintain encrypted database snapshots for disaster recovery.

## Advanced Processing Pipeline

### Prompt Compression

The optional **Prompt Compression Pipeline** in [`server/src/services/compression/pipeline.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/pipeline.ts) performs deduplication, tool output filtering, JSON compression, and stale context trimming before cache lookup and routing. This reduces token costs and improves latency.

### Tool-Call Rescue

Some free-tier models return tool calls as plain text rather than structured objects. The **Tool-Call Rescue** module in [`server/src/lib/tool-call-rescue.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/tool-call-rescue.ts) normalizes these responses into proper `tool_calls` structures, ensuring compatibility with downstream agents expecting OpenAI-compliant formats.

## Management and Control Interfaces

### Admin Dashboard

The **Admin Dashboard** provides a React + Vite interface located in the `client/` directory. It enables key management, fallback chain editing, analytics visualization, and one-click agent configuration.

### Model Control Protocol (MCP) Server

FreeLLMAPI generates an **MCP Server** at runtime (implemented in [`server/src/lib/mcp.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/mcp.ts)) that exposes introspection endpoints under `/mcp`. Agents can query available models, health status, and routing strategies through this protocol.

## Practical Usage Examples

Connecting to FreeLLMAPI requires minimal configuration changes to existing OpenAI clients:

```javascript
// Client-side usage with standard OpenAI SDK
import { Configuration, OpenAIApi } from "openai";

const cfg = new Configuration({
  apiKey: "freellmapi-xxxxxxxxxxxxxxxxxxxx",
  basePath: "http://localhost:3001/v1",
});
const client = new OpenAIApi(cfg);

const resp = await client.createChatCompletion({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Explain the core of FreeLLMAPI." }],
});

```

For agent setup, the CLI helper configures environment variables automatically:

```bash
npx freellmapi setup-claude

# Generates .env with unified token and http://localhost:3001/v1 endpoint

```

## Summary

- **FreeLLMAPI** aggregates free LLM tiers behind a single OpenAI-compatible API using a modular Node.js/TypeScript architecture.
- The **Express Proxy** ([`server/src/app.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/app.ts)) handles incoming requests while the **Router** ([`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)) implements Thompson-sampling bandit algorithms for intelligent model selection.
- **Provider Adapters** ([`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts)) abstract vendor-specific implementations behind a unified interface.
- **Rate limiting** ([`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts)) uses SQLite-backed sliding windows with concurrency leases to respect provider quotas.
- **Security** relies on AES-256-GCM encryption ([`server/src/lib/crypto.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/crypto.ts)) with in-memory-only key decryption during request processing.
- **Advanced features** include prompt compression, tool-call rescue normalization, and an MCP introspection server for agent integration.

## Frequently Asked Questions

### How does FreeLLMAPI handle provider failover?

FreeLLMAPI implements cascading fallback chains where the router iterates through prioritized provider lists stored in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts). If a provider returns errors or exceeds rate limits, the system automatically attempts the next candidate in the chain, using health status from [`server/src/services/health.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/health.ts) to skip known-unhealthy endpoints.

### What encryption standard protects API keys in FreeLLMAPI?

The system uses **AES-256-GCM** encryption implemented in [`server/src/lib/crypto.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/crypto.ts) to protect provider keys at rest in SQLite. Keys are decrypted in-memory only during active request processing via [`server/src/lib/key-proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/key-proxy.ts), minimizing exposure windows.

### How does the router decide which model to use for each request?

The router employs **Thompson-sampling bandit algorithms** defined in [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) to score available models based on real-time reliability, latency, intelligence metrics, and remaining quota headroom. It selects the highest-scoring viable model that can satisfy the current request without violating rate limits.

### What rate-limiting algorithm does FreeLLMAPI employ?

FreeLLMAPI uses a **sliding-window algorithm** with atomic lease management implemented in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts). This tracks RPM, RPD, TPM, and TPD counters per `(platform, model, key)` tuple in SQLite, preventing race conditions through concurrency leases rather than simple check-then-act logic.