How to Self-Host FreeLLMAPI: The Complete Deployment Guide
FreeLLMAPI is a self-hosted Node.js proxy that aggregates free LLM tiers into a single OpenAI-compatible endpoint, deployable via Docker, plain Node.js, or desktop app with built-in AES-256-GCM encryption, intelligent routing, and automatic failover.
FreeLLMAPI (available at tashfeenahmed/freellmapi) unifies access to dozens of free provider tiers behind one local API. Its modular architecture—comprising an Express proxy, SQLite-backed storage, and a React dashboard—lets you self-host with full control over encryption, rate limits, and model fallback chains.
Deployment Methods
FreeLLMAPI supports five distinct deployment patterns, ranging from one-liner Docker scripts to native desktop applications.
Docker One-Liner (Quickest Start)
For Linux, macOS, or WSL, run the official installer which generates an encryption key, pulls the image, and starts the container:
curl -fsSL https://freellmapi.co/install.sh | bash
This script automatically creates a 32-byte hex ENCRYPTION_KEY and exposes the service on port 3001.
Docker Compose (Recommended Production)
For persistent deployments with volume-mounted SQLite storage:
git clone https://github.com/tashfeenahmed/freellmapi.git
cd freellmapi
ENCRYPTION_KEY="$(openssl rand -hex 32)" printf "ENCRYPTION_KEY=%s\nPORT=3001\n" > .env
docker compose up -d
The docker-compose.yml mounts freellmapi-data for database persistence, ensuring your provider keys and request logs survive container restarts.
Local Development Setup
To modify the source or run the Vite-powered dashboard on port 5173:
npm install
ENCRYPTION_KEY="$(node -e 'console.log(require("crypto").randomBytes(32).toString("hex"))')" printf "ENCRYPTION_KEY=%s\nPORT=3001\n" > .env
npm run dev
This launches the Express proxy in server/src/index.ts alongside the React client in client/src/App.tsx with hot-reload enabled.
Production Build (Node.js Binary)
When Docker is unavailable, build and run the compiled JavaScript directly:
npm run build
node server/dist/index.js
Set the NODE_ENV=production variable to disable development logging and enable request compression.
Desktop Application
Download the OS-specific installer from the GitHub Releases page. The desktop bundle includes both the router (from server/src/services/router.ts) and the dashboard UI, requiring no command-line configuration for single-user workstations.
Configuration and Environment Variables
FreeLLMAPI reads runtime options from a .env file or environment variables. The following are required for any self-hosted instance:
| Variable | Purpose | Required |
|---|---|---|
ENCRYPTION_KEY |
32-byte hex string for AES-256-GCM encryption of provider keys | Yes |
PORT |
HTTP port for the API and dashboard (default: 3001) |
No |
HOST_BIND |
Interface to bind (default: 127.0.0.1; set to 0.0.0.0 for LAN exposure) |
No |
FREEAPI_DB_PATH |
Custom SQLite database location (default: ./data/freellmapi.db) |
No |
FREEAPI_CONFIG_PATH |
Path to JSON/YAML config for idempotent model provisioning | No |
In server/src/services/config.ts, the application validates that ENCRYPTION_KEY is exactly 64 hexadecimal characters (32 bytes) before initializing the crypto module in server/lib/crypto.ts.
Architecture and Request Flow
Understanding how requests traverse the system helps optimize your deployment.
The Routing Pipeline
When a client POSTs to /:3001/v1/chat/completions, the following occurs:
-
Request Reception: The Express server in
server/src/index.tsvalidates the unified bearer token (freellmapi-...) and forwards the payload to the router. -
Model Selection:
server/src/services/router.tsimplements a Thompson-sampling bandit algorithm that scores models by reliability, latency, and capability. It queriesserver/src/services/ratelimit.tsto check per-key RPM/RPD/TPM/TPD counters and applies cooldown penalties viarecordRateLimitHit()orrecordModelFailure()when upstreams return 429/5xx errors. -
Key Decryption: The selected provider key is decrypted using
decrypt()fromserver/lib/crypto.tswith theENCRYPTION_KEY. -
Provider Adaptation: The router instantiates the appropriate adapter from
server/src/providers/*.ts(e.g.,google.ts,groq.ts) viagetProvider()and streams the request. -
Automatic Fallback: If the upstream fails, the router marks the key as
isOnCooldownand retries up to 20 times against the next model in the fallback chain. -
Analytics Logging: Successfully completed requests update the SQLite
requeststable (managed inserver/src/db/index.ts) for dashboard visualization.
Health Monitoring
The server/src/services/health.ts module periodically probes each provider key to maintain status states (healthy, rate_limited, invalid, error), feeding real-time availability data to the router without blocking user requests.
Securing Your Installation
FreeLLMAPI encrypts all provider keys at rest using AES-256-GCM. The ENCRYPTION_KEY is never persisted to disk; it must reside in the environment or a .encryption-key file that the application reads on startup.
To add a key securely via the HTTP API:
curl -X POST http://localhost:3001/api/keys \
-H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxxxxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"platform": "groq",
"key": "gsk_XXXXXXXXXXXXXXXX",
"label": "my-groq-key",
"enabled": true
}'
The plaintext key is encrypted immediately before storage in the SQLite database, and only decrypted transiently in memory during active requests.
Practical Usage Examples
OpenAI SDK Integration
Point any OpenAI-compatible client to your local instance:
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "freellmapi-xxxxxxxxxxxxxxxxxxxxxx",
baseURL: "http://localhost:3001/v1",
});
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: "Explain self-hosting FreeLLMAPI" }],
temperature: 0.7
});
console.log(response.choices[0].message.content);
Curated Fallback Configuration
Mount a JSON configuration file to define your model hierarchy idempotently:
{
"fallbackEnabled": true,
"routing": { "strategy": "balanced" },
"models": [
{
"platform": "groq",
"modelId": "llama-3.1-70b",
"displayName": "Llama 3.1 70B",
"supportsTools": true,
"fallbackEnabled": true
},
{
"platform": "google",
"modelId": "gemini-1.5-flash",
"displayName": "Gemini 1.5 Flash",
"supportsTools": true,
"fallbackEnabled": true
}
]
}
Start the server with FREEAPI_CONFIG_PATH=/path/to/freellmapi.config.json to apply these settings on each boot.
Enabling Exploration Mode
Allow the router to test unmeasured models 10% of the time to gather reliability data:
curl -X PUT http://localhost:3001/api/routing/explore \
-H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxxxxxxxxxx" \
-d '{"enabled": true}'
This Thompson-sampling exploration improves the quality of future routing decisions without impacting overall reliability.
Backup and Restore
Automate encrypted SQLite snapshots using the backup service defined in server/src/services/backup.ts:
curl -X POST http://localhost:3001/api/backup \
-H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxxxxxxxxxx" \
-d '{"target": "/mnt/backups/freellmapi-$(date +%F).db.enc"}'
If the database file is missing on startup, FreeLLMAPI automatically restores from the most recent backup target.
Summary
- FreeLLMAPI self-hosts via Docker, Node.js, or desktop app, exposing an OpenAI-compatible endpoint on port
3001. - All provider keys are encrypted with AES-256-GCM using a 32-byte
ENCRYPTION_KEYthat must be set via environment variables. - The router (
server/src/services/router.ts) implements intelligent fallback with Thompson-sampling bandit scoring and cooldown management viaserver/src/services/ratelimit.ts. - Health monitoring (
server/src/services/health.ts) tracks key status asynchronously to prevent routing to failed providers. - Idempotent configuration via JSON files allows reproducible deployments across multiple instances.
Frequently Asked Questions
What are the minimum hardware requirements for self-hosting FreeLLMAPI?
FreeLLMAPI runs on any machine capable of executing Node.js 18+ or Docker. For personal use, 512MB RAM and a single CPU core suffice; production workloads handling 100+ concurrent requests benefit from 2GB RAM and SSD storage for the SQLite database in server/src/db/index.ts.
How does FreeLLMAPI secure my third-party API keys?
According to the server/lib/crypto.ts implementation, each key is encrypted with AES-256-GCM before storage in SQLite. The encryption key must be provided via the ENCRYPTION_KEY environment variable at runtime; without it, the database contents remain unreadable even if the filesystem is compromised.
Can I deploy FreeLLMAPI without using Docker?
Yes. Run npm run build and execute node server/dist/index.js after setting the required ENCRYPTION_KEY in your environment. The server/src/services/config.ts module loads variables from .env files automatically, making bare-metal deployment straightforward.
How does automatic fallback work when a provider rate-limits me?
When the router receives a 429 or 5xx error, server/src/services/ratelimit.ts invokes recordRateLimitHit() to place the key on a temporary cooldown. The router then retries the request against the next available model in the fallback chain, repeating this process up to 20 times until successful or exhausting all options.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →