How to Self-Host FreeLLMAPI: The Complete Deployment Guide

FreeLLMAPI is a self-hosted Node.js proxy that aggregates free LLM tiers into a single OpenAI-compatible endpoint, deployable via Docker, plain Node.js, or desktop app with built-in AES-256-GCM encryption, intelligent routing, and automatic failover.

FreeLLMAPI (available at tashfeenahmed/freellmapi) unifies access to dozens of free provider tiers behind one local API. Its modular architecture—comprising an Express proxy, SQLite-backed storage, and a React dashboard—lets you self-host with full control over encryption, rate limits, and model fallback chains.

Deployment Methods

FreeLLMAPI supports five distinct deployment patterns, ranging from one-liner Docker scripts to native desktop applications.

Docker One-Liner (Quickest Start)

For Linux, macOS, or WSL, run the official installer which generates an encryption key, pulls the image, and starts the container:

curl -fsSL https://freellmapi.co/install.sh | bash

This script automatically creates a 32-byte hex ENCRYPTION_KEY and exposes the service on port 3001.

For persistent deployments with volume-mounted SQLite storage:

git clone https://github.com/tashfeenahmed/freellmapi.git
cd freellmapi
ENCRYPTION_KEY="$(openssl rand -hex 32)" printf "ENCRYPTION_KEY=%s\nPORT=3001\n" > .env
docker compose up -d

The docker-compose.yml mounts freellmapi-data for database persistence, ensuring your provider keys and request logs survive container restarts.

Local Development Setup

To modify the source or run the Vite-powered dashboard on port 5173:

npm install
ENCRYPTION_KEY="$(node -e 'console.log(require("crypto").randomBytes(32).toString("hex"))')" printf "ENCRYPTION_KEY=%s\nPORT=3001\n" > .env
npm run dev

This launches the Express proxy in server/src/index.ts alongside the React client in client/src/App.tsx with hot-reload enabled.

Production Build (Node.js Binary)

When Docker is unavailable, build and run the compiled JavaScript directly:

npm run build
node server/dist/index.js

Set the NODE_ENV=production variable to disable development logging and enable request compression.

Desktop Application

Download the OS-specific installer from the GitHub Releases page. The desktop bundle includes both the router (from server/src/services/router.ts) and the dashboard UI, requiring no command-line configuration for single-user workstations.

Configuration and Environment Variables

FreeLLMAPI reads runtime options from a .env file or environment variables. The following are required for any self-hosted instance:

Variable Purpose Required
ENCRYPTION_KEY 32-byte hex string for AES-256-GCM encryption of provider keys Yes
PORT HTTP port for the API and dashboard (default: 3001) No
HOST_BIND Interface to bind (default: 127.0.0.1; set to 0.0.0.0 for LAN exposure) No
FREEAPI_DB_PATH Custom SQLite database location (default: ./data/freellmapi.db) No
FREEAPI_CONFIG_PATH Path to JSON/YAML config for idempotent model provisioning No

In server/src/services/config.ts, the application validates that ENCRYPTION_KEY is exactly 64 hexadecimal characters (32 bytes) before initializing the crypto module in server/lib/crypto.ts.

Architecture and Request Flow

Understanding how requests traverse the system helps optimize your deployment.

The Routing Pipeline

When a client POSTs to /:3001/v1/chat/completions, the following occurs:

  1. Request Reception: The Express server in server/src/index.ts validates the unified bearer token (freellmapi-...) and forwards the payload to the router.

  2. Model Selection: server/src/services/router.ts implements a Thompson-sampling bandit algorithm that scores models by reliability, latency, and capability. It queries server/src/services/ratelimit.ts to check per-key RPM/RPD/TPM/TPD counters and applies cooldown penalties via recordRateLimitHit() or recordModelFailure() when upstreams return 429/5xx errors.

  3. Key Decryption: The selected provider key is decrypted using decrypt() from server/lib/crypto.ts with the ENCRYPTION_KEY.

  4. Provider Adaptation: The router instantiates the appropriate adapter from server/src/providers/*.ts (e.g., google.ts, groq.ts) via getProvider() and streams the request.

  5. Automatic Fallback: If the upstream fails, the router marks the key as isOnCooldown and retries up to 20 times against the next model in the fallback chain.

  6. Analytics Logging: Successfully completed requests update the SQLite requests table (managed in server/src/db/index.ts) for dashboard visualization.

Health Monitoring

The server/src/services/health.ts module periodically probes each provider key to maintain status states (healthy, rate_limited, invalid, error), feeding real-time availability data to the router without blocking user requests.

Securing Your Installation

FreeLLMAPI encrypts all provider keys at rest using AES-256-GCM. The ENCRYPTION_KEY is never persisted to disk; it must reside in the environment or a .encryption-key file that the application reads on startup.

To add a key securely via the HTTP API:

curl -X POST http://localhost:3001/api/keys \
  -H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
        "platform": "groq",
        "key": "gsk_XXXXXXXXXXXXXXXX",
        "label": "my-groq-key",
        "enabled": true
      }'

The plaintext key is encrypted immediately before storage in the SQLite database, and only decrypted transiently in memory during active requests.

Practical Usage Examples

OpenAI SDK Integration

Point any OpenAI-compatible client to your local instance:

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "freellmapi-xxxxxxxxxxxxxxxxxxxxxx",
  baseURL: "http://localhost:3001/v1",
});

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Explain self-hosting FreeLLMAPI" }],
  temperature: 0.7
});
console.log(response.choices[0].message.content);

Curated Fallback Configuration

Mount a JSON configuration file to define your model hierarchy idempotently:

{
  "fallbackEnabled": true,
  "routing": { "strategy": "balanced" },
  "models": [
    {
      "platform": "groq",
      "modelId": "llama-3.1-70b",
      "displayName": "Llama 3.1 70B",
      "supportsTools": true,
      "fallbackEnabled": true
    },
    {
      "platform": "google",
      "modelId": "gemini-1.5-flash",
      "displayName": "Gemini 1.5 Flash",
      "supportsTools": true,
      "fallbackEnabled": true
    }
  ]
}

Start the server with FREEAPI_CONFIG_PATH=/path/to/freellmapi.config.json to apply these settings on each boot.

Enabling Exploration Mode

Allow the router to test unmeasured models 10% of the time to gather reliability data:

curl -X PUT http://localhost:3001/api/routing/explore \
  -H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxxxxxxxxxx" \
  -d '{"enabled": true}'

This Thompson-sampling exploration improves the quality of future routing decisions without impacting overall reliability.

Backup and Restore

Automate encrypted SQLite snapshots using the backup service defined in server/src/services/backup.ts:

curl -X POST http://localhost:3001/api/backup \
  -H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxxxxxxxxxx" \
  -d '{"target": "/mnt/backups/freellmapi-$(date +%F).db.enc"}'

If the database file is missing on startup, FreeLLMAPI automatically restores from the most recent backup target.

Summary

  • FreeLLMAPI self-hosts via Docker, Node.js, or desktop app, exposing an OpenAI-compatible endpoint on port 3001.
  • All provider keys are encrypted with AES-256-GCM using a 32-byte ENCRYPTION_KEY that must be set via environment variables.
  • The router (server/src/services/router.ts) implements intelligent fallback with Thompson-sampling bandit scoring and cooldown management via server/src/services/ratelimit.ts.
  • Health monitoring (server/src/services/health.ts) tracks key status asynchronously to prevent routing to failed providers.
  • Idempotent configuration via JSON files allows reproducible deployments across multiple instances.

Frequently Asked Questions

What are the minimum hardware requirements for self-hosting FreeLLMAPI?

FreeLLMAPI runs on any machine capable of executing Node.js 18+ or Docker. For personal use, 512MB RAM and a single CPU core suffice; production workloads handling 100+ concurrent requests benefit from 2GB RAM and SSD storage for the SQLite database in server/src/db/index.ts.

How does FreeLLMAPI secure my third-party API keys?

According to the server/lib/crypto.ts implementation, each key is encrypted with AES-256-GCM before storage in SQLite. The encryption key must be provided via the ENCRYPTION_KEY environment variable at runtime; without it, the database contents remain unreadable even if the filesystem is compromised.

Can I deploy FreeLLMAPI without using Docker?

Yes. Run npm run build and execute node server/dist/index.js after setting the required ENCRYPTION_KEY in your environment. The server/src/services/config.ts module loads variables from .env files automatically, making bare-metal deployment straightforward.

How does automatic fallback work when a provider rate-limits me?

When the router receives a 429 or 5xx error, server/src/services/ratelimit.ts invokes recordRateLimitHit() to place the key on a temporary cooldown. The router then retries the request against the next available model in the fallback chain, repeating this process up to 20 times until successful or exhausting all options.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →