Core Components of the FreeLLMAPI Architecture: Technical Deep Dive
FreeLLMAPI is a self-hosted, OpenAI-compatible gateway that aggregates free LLM tiers through a modular architecture comprising an Express proxy layer, intelligent router with Thompson-sampling bandit scoring, provider-specific adapters, SQLite-backed rate limiting, and AES-256-GCM encrypted key management.
FreeLLMAPI serves as a unified API layer that transforms dozens of disparate free-tier LLM providers into a single OpenAI-compatible endpoint. Understanding the FreeLLMAPI architecture reveals how it handles request routing, provider failover, and rate-limit compliance while maintaining security and performance. The system is implemented in TypeScript/Node.js and exposes standard endpoints like /v1/chat/completions while managing complex internal state across approximately 635 model endpoints.
Express Proxy and API Layer
The Express Proxy acts as the entry point for all client traffic, exposing familiar OpenAI-style endpoints including /v1/chat/completions, /v1/embeddings, and /v1/models. According to the tashfeenahmed/freellmapi source code, this layer is implemented in server/src/app.ts and handles HTTP request forwarding to the internal router.
The proxy maintains strict compatibility with the OpenAI API specification, allowing existing clients to migrate by simply changing the basePath configuration while using a unified API key format (freellmapi-…).
Intelligent Request Routing
The Router component in server/src/services/router.ts constitutes the decision-making core of the FreeLLMAPI architecture. It implements the routeRequest function, which performs several critical operations:
- Resolves user-requested model IDs to concrete catalog entries via
resolveRequestedIdForDispatch - Loads fallback chains using
getActiveProfileIdandloadChainRows - Selects the highest-scoring viable provider through iterative evaluation
- Decrypts provider keys in-memory using
decryptfromserver/src/lib/crypto.ts
The router implements sophisticated fallback-chain logic with cooldown handling and per-key diagnostics. If all models in a chain are exhausted, it throws a RouteError with status 429 and detailed diagnostics.
Model Scoring and Selection
Model selection relies on Thompson-sampling bandit algorithms implemented in server/src/services/scoring.ts. This system scores models based on reliability, speed, intelligence metrics, and current quota headroom to optimize routing decisions dynamically.
Provider Adapter System
FreeLLMAPI abstracts vendor-specific SDKs through a standardized Provider Adapter pattern. The BaseProvider interface in server/src/providers/base.ts defines common methods including chatCompletion() and streamChatCompletion().
Individual adapters (such as server/src/providers/google.ts for Google, and similar files for Groq and OpenRouter) encapsulate vendor-specific API calls and error mapping. This architecture allows the router to treat diverse providers uniformly while handling provider-specific quirks internally.
Rate Limiting and Quota Management
The Rate-Limit Ledger in server/src/services/ratelimit.ts implements a sliding-window algorithm backed by SQLite to track RPM (requests per minute), RPD (requests per day), TPM (tokens per minute), and TPD (tokens per day) counters.
Key features include:
- Granular tracking for each
(platform, model, key)tuple - Concurrency "leases" to prevent check-then-act race conditions
- Atomic operations ensuring accurate quota enforcement across concurrent requests
The Health and Quota Service in server/src/services/health.ts complements this by periodically probing key health and deriving per-provider quota headroom, updating cooldown status accordingly.
Model Discovery and Catalog Management
The system maintains a signed, self-updating Model Catalog covering approximately 635 free model endpoints. The discovery logic in server/src/services/model-discovery.ts synchronizes available models with the scoring system to ensure the router only considers viable candidates.
Security and Encryption Architecture
FreeLLMAPI implements defense-in-depth for credential protection:
- AES-256-GCM encryption for provider API keys at rest, implemented in
server/src/lib/crypto.ts - Key Storage in SQLite with encrypted fields
- In-memory decryption per request via
server/src/lib/key-proxy.ts, ensuring keys exist in plaintext only during active request processing
Backup operations in server/src/lib/db-backup.ts maintain encrypted database snapshots for disaster recovery.
Advanced Processing Pipeline
Prompt Compression
The optional Prompt Compression Pipeline in server/src/services/compression/pipeline.ts performs deduplication, tool output filtering, JSON compression, and stale context trimming before cache lookup and routing. This reduces token costs and improves latency.
Tool-Call Rescue
Some free-tier models return tool calls as plain text rather than structured objects. The Tool-Call Rescue module in server/src/lib/tool-call-rescue.ts normalizes these responses into proper tool_calls structures, ensuring compatibility with downstream agents expecting OpenAI-compliant formats.
Management and Control Interfaces
Admin Dashboard
The Admin Dashboard provides a React + Vite interface located in the client/ directory. It enables key management, fallback chain editing, analytics visualization, and one-click agent configuration.
Model Control Protocol (MCP) Server
FreeLLMAPI generates an MCP Server at runtime (implemented in server/src/lib/mcp.ts) that exposes introspection endpoints under /mcp. Agents can query available models, health status, and routing strategies through this protocol.
Practical Usage Examples
Connecting to FreeLLMAPI requires minimal configuration changes to existing OpenAI clients:
// Client-side usage with standard OpenAI SDK
import { Configuration, OpenAIApi } from "openai";
const cfg = new Configuration({
apiKey: "freellmapi-xxxxxxxxxxxxxxxxxxxx",
basePath: "http://localhost:3001/v1",
});
const client = new OpenAIApi(cfg);
const resp = await client.createChatCompletion({
model: "gpt-4o-mini",
messages: [{ role: "user", content: "Explain the core of FreeLLMAPI." }],
});
For agent setup, the CLI helper configures environment variables automatically:
npx freellmapi setup-claude
# Generates .env with unified token and http://localhost:3001/v1 endpoint
Summary
- FreeLLMAPI aggregates free LLM tiers behind a single OpenAI-compatible API using a modular Node.js/TypeScript architecture.
- The Express Proxy (
server/src/app.ts) handles incoming requests while the Router (server/src/services/router.ts) implements Thompson-sampling bandit algorithms for intelligent model selection. - Provider Adapters (
server/src/providers/base.ts) abstract vendor-specific implementations behind a unified interface. - Rate limiting (
server/src/services/ratelimit.ts) uses SQLite-backed sliding windows with concurrency leases to respect provider quotas. - Security relies on AES-256-GCM encryption (
server/src/lib/crypto.ts) with in-memory-only key decryption during request processing. - Advanced features include prompt compression, tool-call rescue normalization, and an MCP introspection server for agent integration.
Frequently Asked Questions
How does FreeLLMAPI handle provider failover?
FreeLLMAPI implements cascading fallback chains where the router iterates through prioritized provider lists stored in server/src/services/router.ts. If a provider returns errors or exceeds rate limits, the system automatically attempts the next candidate in the chain, using health status from server/src/services/health.ts to skip known-unhealthy endpoints.
What encryption standard protects API keys in FreeLLMAPI?
The system uses AES-256-GCM encryption implemented in server/src/lib/crypto.ts to protect provider keys at rest in SQLite. Keys are decrypted in-memory only during active request processing via server/src/lib/key-proxy.ts, minimizing exposure windows.
How does the router decide which model to use for each request?
The router employs Thompson-sampling bandit algorithms defined in server/src/services/scoring.ts to score available models based on real-time reliability, latency, intelligence metrics, and remaining quota headroom. It selects the highest-scoring viable model that can satisfy the current request without violating rate limits.
What rate-limiting algorithm does FreeLLMAPI employ?
FreeLLMAPI uses a sliding-window algorithm with atomic lease management implemented in server/src/services/ratelimit.ts. This tracks RPM, RPD, TPM, and TPD counters per (platform, model, key) tuple in SQLite, preventing race conditions through concurrency leases rather than simple check-then-act logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →