Routing Pipeline for Requests in FreeLLMAPI: The Complete 8-Stage Workflow

FreeLLMAPI processes every incoming completion request through an eight-stage pipeline that authenticates the caller, scores available models against rate-limit budgets, decrypts provider keys in-memory, forwards the call with automatic failover, and returns streaming responses while updating usage counters.

FreeLLMAPI is an open-source, OpenAI-compatible proxy that aggregates free-tier LLM providers into a unified API endpoint. Understanding its request routing pipeline is essential for optimizing latency, managing rate limits across multiple providers like Groq, Google, and OpenRouter, and ensuring high availability through intelligent fallback chains.

The Eight-Stage Routing Pipeline

1. Request Entry and Authentication

All traffic enters through the Express route definition in server/src/app.ts. The server exposes a fully OpenAI-compatible endpoint at POST /v1/chat/completions that accepts standard chat completion payloads.

The authentication middleware extracts the bearer token—formatted as freellmapi-xxxxxxxxxxxx—from the Authorization header. This unified API key identifies the user and initiates the routing context without exposing any upstream provider credentials.

2. Request Validation and Context Building

After authentication, the middleware constructs a RequestContext object that encapsulates the raw payload, user authentication metadata, and a unique RequestId for distributed tracing. This context object travels through the entire pipeline, enabling consistent logging and error tracking across asynchronous boundaries.

3. Router Selection and Model Scoring

The core routing logic resides in server/src/services/router.ts. The Router service implements a scoring algorithm that evaluates every configured model against the active routing strategy—which can be set to priority, balanced, smartest, or custom weightings.

For each candidate model, the router checks the rate-limit ledger (server/src/services/ratelimit.ts) to verify that the request stays within the key’s allocated RPM (requests per minute), RPD (requests per day), TPM (tokens per minute), and TPD (tokens per day) caps. The highest-priority model with a healthy key and available quota wins the selection.

4. API Key Decryption

Once the router selects a provider, it retrieves the encrypted API key from an SQLite-backed key store. The decryption occurs in-memory within server/src/services/key-store.ts using AES-256-GCM. No plaintext key ever touches the filesystem; the decryption key is injected via environment variables at runtime.

5. Provider Adapter Invocation

The router delegates the actual HTTP call to a provider-specific adapter located in server/src/providers/*.ts. Each adapter implements the Provider base class interface, exposing two primary methods:

  • chatCompletion() for synchronous, non-streaming requests
  • streamChatCompletion() for token-by-token streaming responses

The adapter receives the original request payload, injects the decrypted provider key into the headers, and forwards the call to the upstream SDK or raw HTTP endpoint.

6. Automatic Failover and Retry Logic

If the provider returns HTTP 429 (rate limited), 5xx errors, or triggers a timeout, the router initiates an automatic failover sequence. According to the source code in server/src/services/router.ts, the system:

  1. Places the failing key on a short cooldown period recorded in the rate-limit ledger
  2. Retries the request with the next model in the fallback chain
  3. Continues this cycle for up to approximately 20 attempts, bounded by a wall-clock budget to prevent infinite loops

If all candidates exhaust, the proxy returns a detailed error trail to the client indicating which providers were attempted and why they failed.

7. Response Streaming and Cache Headers

For streaming requests, the provider’s byte stream pipes directly back to the client through the Express response object in server/src/app.ts, preserving token-by-token delivery without buffering. Non-streaming calls return the complete JSON payload after the upstream connection closes.

The proxy injects diagnostic headers into every response:

  • X-FreeLLM-Provider: Identifies the actual upstream service (e.g., groq, google, openrouter) that fulfilled the request
  • X-FreeLLM-Cache: Indicates hit or miss for the in-memory LRU cache layer

8. Post-Request Bookkeeping

After a successful response, the pipeline updates the rate-limit counters in server/src/services/ratelimit.ts to deduct tokens and increment request counts. Simultaneously, the server/src/services/health.ts service periodically probes provider keys in the background to refresh their status markers (healthy, rate-limited, or invalid), ensuring future routing decisions use current availability data.

Code Examples

Sending a Request with curl

curl -X POST https://my-freellmapi.example.com/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-xxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{"role":"user","content":"Explain the routing pipeline"}],
        "stream": false
      }'

Using the OpenAI SDK (Node.js)

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://my-freellmapi.example.com/v1",
  apiKey: "freellmapi-xxxxxxxxxxxx",
});

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Explain the routing pipeline" }],
  stream: true,
});

for await (const chunk of response) {
  process.stdout.write(chunk.choices[0].delta?.content || "");
}

Inspecting Provider Routing via Headers

curl -i -X POST https://my-freellmapi.example.com/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-xxxx" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Hi"}]}'

Typical response headers:


X-FreeLLM-Provider: groq
X-FreeLLM-Cache: miss

Core Architecture and Source Files

The routing pipeline spans multiple specialized services:

  • server/src/app.ts: Express server initialization, route mounting, authentication middleware, and response streaming pipes
  • server/src/services/router.ts: Central orchestrator implementing model scoring, fallback chain traversal, and retry logic with exponential backoff
  • server/src/services/ratelimit.ts: Hybrid in-memory and SQLite ledger tracking RPM, RPD, TPM, and TPD consumption per API key
  • server/src/providers/*.ts: Adapter implementations for each supported LLM provider, all extending the Provider abstract class
  • server/src/services/key-store.ts: Cryptographic layer handling AES-256-GCM encryption and decryption of provider credentials
  • server/src/services/health.ts: Background worker that probes keys to maintain accurate health statuses for the router’s scoring algorithm

Summary

  • FreeLLMAPI exposes an OpenAI-compatible endpoint at POST /v1/chat/completions that accepts standard chat completion requests with a unified freellmapi-* bearer token.
  • The Router service in server/src/services/router.ts scores models using configurable strategies and enforces rate limits via server/src/services/ratelimit.ts.
  • Provider API keys remain encrypted at rest in SQLite and decrypt only in-memory using AES-256-GCM within server/src/services/key-store.ts.
  • Automatic failover retries failed requests across up to 20 alternative providers, applying cooldowns to throttled keys before reattempting.
  • Streaming responses pipe directly from the upstream provider to the client, while the system updates usage counters and health status after every call.

Frequently Asked Questions

How does FreeLLMAPI handle rate limiting across different providers?

The system maintains a rate-limit ledger in server/src/services/ratelimit.ts that tracks RPM, RPD, TPM, and TPD counters for every configured API key. Before routing, the Router service checks these counters against the request’s estimated token cost; if a key exceeds its quota, the router skips that provider and attempts the next candidate in the fallback chain.

What happens when every provider in the fallback chain fails?

If all approximately 20 retry attempts exhaust across the fallback chain, the router terminates the pipeline and returns a structured error response to the client. This response includes a detailed trail indicating which providers were attempted, the specific HTTP status codes or timeout errors encountered, and which rate-limit caps were hit, enabling debugging without exposing decrypted keys.

How does the routing strategy setting affect which model gets selected?

The routing strategy—configured as priority, balanced, smartest, or custom—alters the scoring weights in server/src/services/router.ts. Priority selects the first healthy model in the user-defined list regardless of load, balanced distributes traffic across healthy keys to maximize throughput, and smartest uses latency metrics and token-cost heuristics to minimize response time and maximize free-tier utilization.

Is the storage of upstream API keys secure?

Yes. Upstream provider keys are encrypted using AES-256-GCM before being written to the SQLite database in server/src/services/key-store.ts. The decryption key is supplied via environment variables and never persisted to disk; decryption occurs only in-memory at request time, ensuring that even a full database compromise would not reveal plaintext provider credentials.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →