How FreeLLMAPI Request Flow Works: A Complete Technical Guide

FreeLLMAPI operates as an OpenAI-compatible proxy that routes incoming requests through a seven-stage pipeline—authentication, parameter sanitization, intelligent model selection via Thompson-sampling bandits, AES-256-GCM key decryption, provider SDK invocation, automatic failover with rate-limit tracking, and response finalization—to deliver free-tier LLM access with enterprise-grade reliability.

FreeLLMAPI is an open-source, unified gateway that aggregates multiple free-tier large language model providers behind a single OpenAI-compatible API endpoint. Understanding the FreeLLMAPI request flow is essential for developers who want to leverage its intelligent routing, automatic failover mechanisms, and cryptographic key management. This article examines the complete request lifecycle as implemented in the tashfeenahmed/freellmapi repository, from the initial HTTP ingress in Express to the final provider response.

Step 1: Ingress and Authentication

Every FreeLLMAPI request begins at the Express server entry point in server/src/app.ts. The application exposes standard OpenAI-style routes such as POST /v1/chat/completions and validates incoming requests using a unified bearer token format (freellmapi-…).

The Express application handles request parsing and wires the authentication middleware before forwarding traffic to the core router service. This thin ingress layer ensures that only properly formatted requests with valid FreeLLMAPI credentials proceed to the processing pipeline.

Step 2: Parameter Sanitization and Normalization

Once authenticated, the request enters the router service (server/src/services/router.ts), where the system performs parameter filtering and normalization around lines 420-440. This stage ensures provider compatibility by:

  • Filtering unsupported parameters per provider (for example, dropping logprobs for Groq)
  • Normalizing deprecated keys such as converting max_completion_tokens to max_tokens
  • Validating request structure against the selected provider's capabilities

This sanitization prevents provider-specific errors and ensures consistent behavior across the diverse set of supported LLM backends.

Step 3: Intelligent Model Selection

The core of the FreeLLMAPI request flow resides in the routing engine (server/src/services/router.ts), which implements a sophisticated selection algorithm using Thompson-sampling bandits. The orderChain and scoreChainEntry functions (lines 1025-1064) evaluate candidate models based on:

  • Health status: API key validity and current availability
  • Rate limit headroom: Remaining capacity under RPM (requests per minute), RPD (requests per day), TPM (tokens per minute), and TPD (tokens per day) caps
  • Capability matching: Support for vision, tool-calling, or structured output requirements
  • Performance metrics: Historical reliability, speed, and intelligence scores

The router builds a prioritized fallback chain and selects the highest-scoring healthy model, ensuring optimal resource utilization across the free-tier provider pool.

Step 4: Cryptographic Key Decryption

After selecting a model/key pair, the router decrypts the stored credentials using AES-256-GCM encryption. The decrypt() function in lib/crypto.ts handles the primary API key decryption, while decryptProxyUrl() manages per-key proxy overrides.

This decryption occurs per-request, ensuring that credentials remain encrypted at rest and are only exposed transiently during the provider communication window. The decrypted key is then prepared for injection into the provider's authentication header.

Step 5: Provider SDK Invocation

With decrypted credentials in hand, the router invokes the appropriate provider adapter from server/src/providers/*.ts (such as google.ts or groq.ts). Each adapter implements a minimal subset of the OpenAI SDK interface, specifically:

  • chatCompletion() for synchronous responses
  • streamChatCompletion() for Server-Sent Events (SSE) streaming

The base provider class in server/src/providers/base.ts defines the standard interface, ensuring consistent behavior whether the backend is Groq, Google, or another supported service. The decrypted API key is passed as the authorization header to the underlying provider.

Step 6: Error Handling and Automatic Retries

If the provider returns a 429 (rate limit), 5xx (server error), or timeout, the FreeLLMAPI request flow enters its resilience phase. The router:

  1. Registers a cooldown in the rate-limit ledger (server/src/services/ratelimit.ts)
  2. Records a penalty via recordRateLimitHit() to update the model's reliability score
  3. Retries the request with the next model in the fallback chain

The system attempts up to approximately 20 retries across the entire provider chain before exhausting options. This automatic failover occurs transparently to the client, requiring no manual retry logic in consuming applications.

Step 7: Response Finalization and Lease Management

Upon successful provider response, the router streams the result back to the client and executes cleanup operations. The releaseLease() function returns the in-flight request slot to the pool, while analytics caches update with latency and token consumption metrics.

If all models in the chain fail, the router throws a RouteError with the diagnostic summary "All models exhausted…", providing detailed failure information for debugging. This finalization stage ensures resource cleanup and accurate telemetry regardless of success or failure.

Client Integration Examples

You can interact with FreeLLMAPI using standard HTTP clients or the official OpenAI SDK, as the proxy maintains full API compatibility.

Using cURL

curl https://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{"role":"user","content":"Explain the request flow"}],
        "max_tokens": 256
      }'

Using the OpenAI SDK

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",            // FreeLLMAPI proxy
  apiKey: "freellmapi-xxxxxxxxxxxxxx",            // unified token
});

const resp = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Explain the request flow" }],
  max_tokens: 256,
});

console.log(resp.choices[0].message.content);

Both examples trigger the complete FreeLLMAPI request flow, including automatic failover if the initially selected provider hits quota limits.

Summary

  • FreeLLMAPI acts as a transparent OpenAI-compatible proxy aggregating multiple free-tier providers.
  • The routing engine in server/src/services/router.ts implements Thompson-sampling bandits for intelligent model selection based on health, rate limits, and performance.
  • AES-256-GCM encryption protects API keys at rest, with per-request decryption handled in lib/crypto.ts.
  • The system provides automatic failover with up to approximately 20 retry attempts across the provider chain when encountering 429 or 5xx errors.
  • Rate limiting is tracked via an SQLite-backed ledger in server/src/services/ratelimit.ts, monitoring RPM/RPD/TPM/TPD metrics per key.

Frequently Asked Questions

How does FreeLLMAPI handle provider rate limits?

FreeLLMAPI maintains an in-memory rate-limit ledger in server/src/services/ratelimit.ts that tracks RPM, RPD, TPM, and TPD consumption per API key. When a provider returns a 429 error, the router invokes recordRateLimitHit(), registers a cooldown period, and automatically retries with the next available model in the fallback chain.

What encryption standard protects API keys in FreeLLMAPI?

The system uses AES-256-GCM encryption to secure provider API keys at rest. The decrypt() function in lib/crypto.ts handles decryption during the request flow, while decryptProxyUrl() manages encrypted proxy overrides, ensuring keys are only exposed transiently during active provider communication.

How many retry attempts does FreeLLMAPI attempt before failing?

The router implements a resilient retry mechanism that attempts approximately 20 fallback models before exhausting the chain. Each failure (429, 5xx, or timeout) triggers a retry with the next highest-priority healthy model, making the process transparent to API consumers.

Which file contains the core routing logic for model selection?

The primary routing intelligence resides in server/src/services/router.ts, specifically within the orderChain and scoreChainEntry functions (lines 1025-1064). This file handles model scoring using Thompson-sampling bandits, parameter sanitization, key decryption, and the retry orchestration logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →