# How the FreeLLMAPI Request Flow Works: A Step-by-Step Technical Breakdown

> Understand the FreeLLMAPI request flow. Explore its eight-stage pipeline from authentication and model scoring to optimal provider dispatch and normalized response delivery.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-09-02

---

**When you send a request to FreeLLMAPI, it traverses an eight-stage pipeline that authenticates the client, scores candidate models using Thompson-sampling bandits, acquires per-key leases, and dispatches to the optimal provider before returning a normalized response.**

The FreeLLMAPI request flow orchestrates multi-provider LLM routing through a sophisticated architecture designed to minimize latency and maximize reliability. Implemented in the `tashfeenahmed/freellmapi` repository, this system transforms a simple `POST /v1/chat/completions` call into an intelligent, quota-aware dispatch operation. Understanding how the server processes incoming requests reveals the mechanics behind its dynamic scoring, fallback chains, and real-time penalty systems.

## The FreeLLMAPI Request Pipeline

### Entry and Authentication Layer

Every request first hits the route handlers in `server/src/routes/…`, where middleware performs essential gatekeeping. The [`requireAuth.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/requireAuth.ts) middleware validates the incoming API key, while [`rateLimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/rateLimit.ts) enforces global request-level throttling before any routing logic executes. This ensures only authenticated, non-overflow traffic enters the pipeline.

### Request Normalization and Token Guardrails

Once past middleware, the route extracts and validates the payload. The `routingReserveTokens()` function (lines 41-50 in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)) caps the output token reservation to prevent starving the routing pool. This normalization step ensures that `max_tokens` and other optional fields are sanitized before model selection begins.

### Intelligent Model Selection

The router builds a **chain of candidate models** (`ChainRow`) and computes a priority order. For each candidate, `scoreChainEntry()` (lines 443-511 in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) ) calculates a composite score combining reliability, speed, intelligence, head-room, and any active 429-penalties. These scores pull from decay-weighted statistics stored in the in-memory `statsCache`.

The active **routing strategy**—configured via `getRoutingStrategy` as `priority`, `balanced`, or `fastest`—determines how the chain is ordered. When `sampled = true`, the system runs a **Thompson-sampling bandit** algorithm through `orderChain()` (lines 525-565 in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)) to balance exploration and exploitation across providers.

### API Key Leasing and Concurrency Control

After selecting a model, the router determines which API key to use based on the `key_selection_strategy` setting (`auto` for round-robin or `least-remaining`). The `acquireLease()` and `releaseLease()` functions in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) enforce strict per-key concurrency limits, ensuring no single credential exceeds its allowed parallel requests.

### Provider Dispatch and Response Handling

With a concrete `(provider, modelId, apiKey)` tuple, the request enters the provider layer. Each vendor implements the `BaseProvider` interface defined in [`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts), translating the generic LLMAPI payload into vendor-specific HTTP calls (e.g., OpenAI formats in [`server/src/providers/openai-compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/openai-compat.ts)). The raw response returns to the route handler for thin post-processing, such as injecting the `model` field or handling streaming chunks. If the provider returns a 429 error, `recordRateLimitHit()` (lines 91-93 in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)) applies a penalty that demotes the model in future routing decisions until it recovers via `DECAY_INTERVAL_MS`.

## Dynamic Routing Algorithms

### The scoreChainEntry Reliability Engine

The **dynamic scoring** system treats reliability as a beta-posterior (`reliabilityPosterior`), calculates speed as a composite of token-throughput and latency (`speedScore`), and factors in model tier through `intelligenceComposite`. The **head-room factor** (`headroomFactor`) monitors monthly budgets and per-window usage, demoting models before they exhaust quotas.

### Thompson-Sampling Bandit Optimization

When the routing strategy enables sampling (`sampled = true`), the system applies Thompson sampling to the chain. This statistical method balances trying high-performing models against exploring potentially better alternatives, optimizing long-term success rates across the provider pool.

### Rate Limit Penalties and Recovery

The **penalty system** distinguishes between generic failures (`recordModelFailure`) and rate-limit hits (`recordRateLimitHit`), with 429 responses carrying heavier penalties. Penalized models sink in the ordering chain until statistical decay restores their standing, creating automatic circuit-breaker behavior.

## Practical Integration Examples

### Direct cURL Request

Send a chat completion through the public endpoint:

```bash
curl -X POST https://api.free-llmapi.com/v1/chat/completions \
     -H "Authorization: Bearer $FREELLMAPI_KEY" \
     -H "Content-Type: application/json" \
     -d '{
           "model": "gpt-4o-mini",
           "messages": [{"role":"user","content":"What is the capital of France?"}],
           "max_tokens": 512
         }'

```

### Node.js SDK Integration

Use the bundled SDK from [`cli/src/tools.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/cli/src/tools.ts) for programmatic access:

```javascript
import { fetchChatCompletion } from './cli/src/tools.js';

const response = await fetchChatCompletion({
  apiKey: process.env.FREELLMAPI_KEY,
  model: 'gpt-4o-mini',
  messages: [{ role: 'user', content: 'Explain quantum tunnelling in one sentence.' }],
  maxTokens: 256,
});

console.log(response.choices[0].message.content);

```

### Manual Routing Inspection

Inspect the routing decision and fallback chain manually:

```javascript
import { routeRequest, RouteError } from '../server/src/services/router.js';

try {
  const result = await routeRequest({
    model: 'gpt-4o-mini',
    messages: [{ role: 'user', content: 'Ping?' }],
    maxTokens: 128,
  });
  console.log('Routed to', result.provider.name, 'via key', result.keyLabel);
} catch (e) {
  if (e instanceof RouteError) {
    console.error('All models exhausted:', e.diagnostics);
  }
}

```

## Summary

- **Authentication and rate-limiting** occur first in `server/src/middleware/` before any routing logic executes.
- **Token reservation** via `routingReserveTokens()` prevents pool starvation by capping output allocations.
- **Model scoring** uses `scoreChainEntry()` to combine reliability, speed, and intelligence metrics with real-time quota head-room.
- **Thompson-sampling bandits** optimize the routing strategy when `sampled = true`, balancing exploration against exploitation.
- **API key selection** follows `auto` (round-robin) or `least-remaining` strategies with lease-based concurrency enforcement in [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts).
- **Provider abstraction** through `BaseProvider` enables uniform handling of OpenAI, Anthropic, Google, and other backends.
- **Penalty decay** automatically recovers models from 429 errors via `recordRateLimitHit()` and `DECAY_INTERVAL_MS`.

## Frequently Asked Questions

### How does FreeLLMAPI handle rate limiting across multiple providers?

FreeLLMAPI tracks rate-limit hits through `recordRateLimitHit()` in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), which applies decay-weighted penalties to models returning 429 errors. The system maintains an in-memory `statsCache` that demotes penalized providers in the routing chain until their statistics recover, effectively creating a self-healing circuit breaker that avoids repeatedly hitting exhausted quotas.

### What routing strategies are available in FreeLLMAPI?

The platform supports multiple strategies configurable via `getRoutingStrategy`: **priority** respects manual ordering, **balanced** optimizes for cost-performance ratios, and **fastest** prioritizes low-latency providers. When Thompson sampling is enabled (`sampled = true`), the `orderChain()` function applies probabilistic bandit algorithms to dynamically optimize selection based on historical performance metrics.

### How does FreeLLMAPI choose between different API keys for the same provider?

Key selection follows the `key_selection_strategy` setting, supporting `auto` (round-robin distribution) or `least-remaining` (selecting the key with the most quota left). The `acquireLease()` and `releaseLease()` functions in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) enforce per-key concurrency limits, preventing any single credential from exceeding its parallel request capacity.

### What happens when all models in the chain fail?

If the primary model fails, the router iterates through the ordered `ChainRow` list, attempting each candidate until one succeeds or the chain exhausts. Upon total failure, the system throws a `RouteError` containing diagnostic metadata via `summarizeExhaustion`, allowing client code to inspect which providers were attempted and why each failed.