How to Set Up FreeLLMAPI Self-Hosted: Complete Docker Deployment Guide

FreeLLMAPI is a self-hosted gateway that aggregates multiple large-language-model providers (OpenAI, Anthropic, Azure) behind a unified OpenAI-compatible REST API, deployable via Docker Compose using a single environment file.

Setting up FreeLLMAPI self-hosted gives you a private, unified interface to route requests across multiple LLM providers with intelligent failover and quota management. The tashfeenahmed/freellmapi repository provides a containerized Fastify server that exposes standard OpenAI endpoints while the routing engine distributes traffic based on health checks and scoring rules. All configuration lives in a single .env file, making deployment reproducible and version-controlled.

Prerequisites

Before deploying FreeLLMAPI self-hosted, ensure your environment meets the following requirements:

  • Docker Engine (or Docker Desktop) and Docker Compose installed on your host machine. See the official installation guide at docs/install/OVERVIEW.md for platform-specific instructions.
  • API keys for at least one supported provider (OpenAI, Anthropic, Azure OpenAI, or others).
  • Git to clone the repository.

Installation Steps

Deploying FreeLLMAPI involves cloning the repository, configuring environment variables, and starting the Docker containers.

1. Clone the Repository

git clone https://github.com/tashfeenahmed/freellmapi.git
cd freellmapi

2. Configure Environment Variables

Copy the example environment file and edit it to include your provider API keys:

cp .env.example .env

Edit .env and populate the required fields. At minimum, configure one provider:

PORT=8080
OPENAI_API_KEY=sk-your-openai-key-here
ANTHROPIC_API_KEY=sk-your-anthropic-key-here
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com/
AZURE_OPENAI_API_KEY=your-azure-key

The .env file controls all configurable values including rate limits (RATE_LIMIT_GLOBAL, RATE_LIMIT_USER), model weights, and retention settings.

3. Start the Service

Launch the container stack in detached mode:

docker compose up -d

This command pulls the official image ghcr.io/tashfeenahmed/freellmapi:latest, mounts your .env file, and exposes port 8080 (configurable via the PORT variable). The service defined in docker-compose.yml includes a restart policy to ensure high availability.

4. Verify the Deployment

Test that the gateway is responding by requesting the models list:

curl http://localhost:8080/v1/models

You should receive a JSON array of available models that the router currently exposes from your configured providers.

Testing API Calls

Once running, FreeLLMAPI accepts standard OpenAI-compatible requests. Send a chat completion through your local gateway:

curl http://localhost:8080/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role":"user","content":"Tell me a joke"}]
  }'

The router logic in server/src/services/router.ts processes this request, applies bandit scoring to select the optimal provider, and enforces rate limits defined in server/src/services/ratelimit.ts before forwarding the call.

Using the CLI Client

The repository includes a TypeScript CLI for management and diagnostics. Install it globally from the repo root:

npm install -g .

Run health checks and test requests:

llmapi doctor
llmapi chat "Hello, world!"

The CLI entry point resides in cli/src/index.ts and provides commands for key management and service diagnostics.

Core Architecture Components

Understanding the three main components helps with troubleshooting and customization:

Production Configuration

Data Persistence

By default, the server maintains no state on disk. To enable request retention, persistent rate-limit counters, or model catalog synchronization, mount a volume to /app/data and enable persistence flags in .env:

RETENTION_ENABLED=true

Details on degraded mode and failover configuration are available in docs/architecture/04-degraded-mode-and-failover.md.

Updating the Service

To upgrade to the latest version:

git pull
docker compose pull
docker compose up -d --force-recreate

This pulls the latest image and recreates containers without downtime.

Summary

  • FreeLLMAPI self-hosted deploys via Docker Compose using a single .env file for all configuration.
  • The gateway aggregates OpenAI, Anthropic, Azure, and other providers behind a unified OpenAI-compatible API endpoint.
  • Core routing logic in server/src/services/router.ts provides intelligent load balancing and failover across backends.
  • Rate limiting is enforced through server/src/services/ratelimit.ts while quota tracking occurs in server/src/services/provider-quota.ts.
  • Optional data persistence and the llmapi CLI client support production operational requirements.

Frequently Asked Questions

What providers does FreeLLMAPI support?

FreeLLMAPI supports any provider compatible with the OpenAI API format, including OpenAI, Anthropic, Azure OpenAI, and custom endpoints. Configuration occurs entirely through environment variables in .env, allowing you to mix multiple providers simultaneously. The router automatically load-balances requests across healthy providers based on latency and quota availability.

How does the routing logic select between providers?

The Bandit Scoring algorithm implemented in server/src/services/router.ts evaluates provider health, remaining quota, and historical performance to select the optimal backend for each request. If a provider fails health checks or exceeds its configured quota (tracked in server/src/services/provider-quota.ts), the router automatically fails over to the next available option. See docs/architecture/01-routing-and-bandit-scoring.md for the scoring mathematics.

Can I deploy FreeLLMAPI without Docker?

While Docker is the officially supported deployment method documented in docs/install/OVERVIEW.md, the service is a standard Node.js/Fastify application. You can run it directly using npm install and npm start from the server directory, though you must manually manage environment variables and process management. Docker remains the recommended approach for production stability.

How do I monitor quota usage and rate limits?

The gateway tracks consumption in real-time through the Provider Quota Engine (server/src/services/provider-quota.ts). You can inspect current limits via the CLI using llmapi doctor, or query the internal health endpoints. Per-user and global rate limits are enforced by server/src/services/ratelimit.ts, with breach events logged to stdout for integration with external monitoring systems.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →