# How FreeLLMAPI Request Flow Works: A Complete Technical Guide

> Discover FreeLLMAPI's request flow: authentication, sanitization, model selection, decryption, SDK invocation, failover, and finalization. Get free LLM access with enterprise reliability.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: technical-guide
- Published: 2026-08-29

---

**FreeLLMAPI operates as an OpenAI-compatible proxy that routes incoming requests through a seven-stage pipeline—authentication, parameter sanitization, intelligent model selection via Thompson-sampling bandits, AES-256-GCM key decryption, provider SDK invocation, automatic failover with rate-limit tracking, and response finalization—to deliver free-tier LLM access with enterprise-grade reliability.**

FreeLLMAPI is an open-source, unified gateway that aggregates multiple free-tier large language model providers behind a single OpenAI-compatible API endpoint. Understanding the FreeLLMAPI request flow is essential for developers who want to leverage its intelligent routing, automatic failover mechanisms, and cryptographic key management. This article examines the complete request lifecycle as implemented in the `tashfeenahmed/freellmapi` repository, from the initial HTTP ingress in Express to the final provider response.

## Step 1: Ingress and Authentication

Every FreeLLMAPI request begins at the **Express server** entry point in [`server/src/app.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/app.ts). The application exposes standard OpenAI-style routes such as `POST /v1/chat/completions` and validates incoming requests using a unified bearer token format (`freellmapi-…`).

The Express application handles request parsing and wires the authentication middleware before forwarding traffic to the core router service. This thin ingress layer ensures that only properly formatted requests with valid FreeLLMAPI credentials proceed to the processing pipeline.

## Step 2: Parameter Sanitization and Normalization

Once authenticated, the request enters the **router service** ([`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)), where the system performs parameter filtering and normalization around lines 420-440. This stage ensures provider compatibility by:

- Filtering unsupported parameters per provider (for example, dropping `logprobs` for Groq)
- Normalizing deprecated keys such as converting `max_completion_tokens` to `max_tokens`
- Validating request structure against the selected provider's capabilities

This sanitization prevents provider-specific errors and ensures consistent behavior across the diverse set of supported LLM backends.

## Step 3: Intelligent Model Selection

The core of the FreeLLMAPI request flow resides in the **routing engine** ([`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)), which implements a sophisticated selection algorithm using **Thompson-sampling bandits**. The `orderChain` and `scoreChainEntry` functions (lines 1025-1064) evaluate candidate models based on:

- **Health status**: API key validity and current availability
- **Rate limit headroom**: Remaining capacity under RPM (requests per minute), RPD (requests per day), TPM (tokens per minute), and TPD (tokens per day) caps
- **Capability matching**: Support for vision, tool-calling, or structured output requirements
- **Performance metrics**: Historical reliability, speed, and intelligence scores

The router builds a prioritized fallback chain and selects the highest-scoring healthy model, ensuring optimal resource utilization across the free-tier provider pool.

## Step 4: Cryptographic Key Decryption

After selecting a model/key pair, the router decrypts the stored credentials using **AES-256-GCM** encryption. The `decrypt()` function in [`lib/crypto.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/lib/crypto.ts) handles the primary API key decryption, while `decryptProxyUrl()` manages per-key proxy overrides.

This decryption occurs per-request, ensuring that credentials remain encrypted at rest and are only exposed transiently during the provider communication window. The decrypted key is then prepared for injection into the provider's authentication header.

## Step 5: Provider SDK Invocation

With decrypted credentials in hand, the router invokes the appropriate **provider adapter** from `server/src/providers/*.ts` (such as [`google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/google.ts) or [`groq.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/groq.ts)). Each adapter implements a minimal subset of the OpenAI SDK interface, specifically:

- `chatCompletion()` for synchronous responses
- `streamChatCompletion()` for Server-Sent Events (SSE) streaming

The base provider class in [`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts) defines the standard interface, ensuring consistent behavior whether the backend is Groq, Google, or another supported service. The decrypted API key is passed as the authorization header to the underlying provider.

## Step 6: Error Handling and Automatic Retries

If the provider returns a **429** (rate limit), **5xx** (server error), or timeout, the FreeLLMAPI request flow enters its resilience phase. The router:

1. Registers a cooldown in the **rate-limit ledger** ([`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts))
2. Records a penalty via `recordRateLimitHit()` to update the model's reliability score
3. Retries the request with the **next model** in the fallback chain

The system attempts up to approximately 20 retries across the entire provider chain before exhausting options. This automatic failover occurs transparently to the client, requiring no manual retry logic in consuming applications.

## Step 7: Response Finalization and Lease Management

Upon successful provider response, the router streams the result back to the client and executes cleanup operations. The `releaseLease()` function returns the in-flight request slot to the pool, while analytics caches update with latency and token consumption metrics.

If all models in the chain fail, the router throws a `RouteError` with the diagnostic summary "All models exhausted…", providing detailed failure information for debugging. This finalization stage ensures resource cleanup and accurate telemetry regardless of success or failure.

## Client Integration Examples

You can interact with FreeLLMAPI using standard HTTP clients or the official OpenAI SDK, as the proxy maintains full API compatibility.

### Using cURL

```bash
curl https://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-xxxxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{"role":"user","content":"Explain the request flow"}],
        "max_tokens": 256
      }'

```

### Using the OpenAI SDK

```javascript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",            // FreeLLMAPI proxy
  apiKey: "freellmapi-xxxxxxxxxxxxxx",            // unified token
});

const resp = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Explain the request flow" }],
  max_tokens: 256,
});

console.log(resp.choices[0].message.content);

```

Both examples trigger the complete FreeLLMAPI request flow, including automatic failover if the initially selected provider hits quota limits.

## Summary

- **FreeLLMAPI** acts as a transparent OpenAI-compatible proxy aggregating multiple free-tier providers.
- The routing engine in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) implements **Thompson-sampling bandits** for intelligent model selection based on health, rate limits, and performance.
- **AES-256-GCM encryption** protects API keys at rest, with per-request decryption handled in [`lib/crypto.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/lib/crypto.ts).
- The system provides **automatic failover** with up to approximately 20 retry attempts across the provider chain when encountering 429 or 5xx errors.
- **Rate limiting** is tracked via an SQLite-backed ledger in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts), monitoring RPM/RPD/TPM/TPD metrics per key.

## Frequently Asked Questions

### How does FreeLLMAPI handle provider rate limits?

FreeLLMAPI maintains an in-memory rate-limit ledger in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) that tracks RPM, RPD, TPM, and TPD consumption per API key. When a provider returns a 429 error, the router invokes `recordRateLimitHit()`, registers a cooldown period, and automatically retries with the next available model in the fallback chain.

### What encryption standard protects API keys in FreeLLMAPI?

The system uses **AES-256-GCM** encryption to secure provider API keys at rest. The `decrypt()` function in [`lib/crypto.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/lib/crypto.ts) handles decryption during the request flow, while `decryptProxyUrl()` manages encrypted proxy overrides, ensuring keys are only exposed transiently during active provider communication.

### How many retry attempts does FreeLLMAPI attempt before failing?

The router implements a resilient retry mechanism that attempts approximately **20 fallback models** before exhausting the chain. Each failure (429, 5xx, or timeout) triggers a retry with the next highest-priority healthy model, making the process transparent to API consumers.

### Which file contains the core routing logic for model selection?

The primary routing intelligence resides in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), specifically within the `orderChain` and `scoreChainEntry` functions (lines 1025-1064). This file handles model scoring using Thompson-sampling bandits, parameter sanitization, key decryption, and the retry orchestration logic.