How the FreeLLMAPI Request Flow Works: A Step-by-Step Technical Breakdown

When you send a request to FreeLLMAPI, it traverses an eight-stage pipeline that authenticates the client, scores candidate models using Thompson-sampling bandits, acquires per-key leases, and dispatches to the optimal provider before returning a normalized response.

The FreeLLMAPI request flow orchestrates multi-provider LLM routing through a sophisticated architecture designed to minimize latency and maximize reliability. Implemented in the tashfeenahmed/freellmapi repository, this system transforms a simple POST /v1/chat/completions call into an intelligent, quota-aware dispatch operation. Understanding how the server processes incoming requests reveals the mechanics behind its dynamic scoring, fallback chains, and real-time penalty systems.

The FreeLLMAPI Request Pipeline

Entry and Authentication Layer

Every request first hits the route handlers in server/src/routes/…, where middleware performs essential gatekeeping. The requireAuth.ts middleware validates the incoming API key, while rateLimit.ts enforces global request-level throttling before any routing logic executes. This ensures only authenticated, non-overflow traffic enters the pipeline.

Request Normalization and Token Guardrails

Once past middleware, the route extracts and validates the payload. The routingReserveTokens() function (lines 41-50 in server/src/services/router.ts) caps the output token reservation to prevent starving the routing pool. This normalization step ensures that max_tokens and other optional fields are sanitized before model selection begins.

Intelligent Model Selection

The router builds a chain of candidate models (ChainRow) and computes a priority order. For each candidate, scoreChainEntry() (lines 443-511 in server/src/services/router.ts ) calculates a composite score combining reliability, speed, intelligence, head-room, and any active 429-penalties. These scores pull from decay-weighted statistics stored in the in-memory statsCache.

The active routing strategy—configured via getRoutingStrategy as priority, balanced, or fastest—determines how the chain is ordered. When sampled = true, the system runs a Thompson-sampling bandit algorithm through orderChain() (lines 525-565 in server/src/services/router.ts) to balance exploration and exploitation across providers.

API Key Leasing and Concurrency Control

After selecting a model, the router determines which API key to use based on the key_selection_strategy setting (auto for round-robin or least-remaining). The acquireLease() and releaseLease() functions in server/src/services/ratelimit.ts enforce strict per-key concurrency limits, ensuring no single credential exceeds its allowed parallel requests.

Provider Dispatch and Response Handling

With a concrete (provider, modelId, apiKey) tuple, the request enters the provider layer. Each vendor implements the BaseProvider interface defined in server/src/providers/base.ts, translating the generic LLMAPI payload into vendor-specific HTTP calls (e.g., OpenAI formats in server/src/providers/openai-compat.ts). The raw response returns to the route handler for thin post-processing, such as injecting the model field or handling streaming chunks. If the provider returns a 429 error, recordRateLimitHit() (lines 91-93 in server/src/services/router.ts) applies a penalty that demotes the model in future routing decisions until it recovers via DECAY_INTERVAL_MS.

Dynamic Routing Algorithms

The scoreChainEntry Reliability Engine

The dynamic scoring system treats reliability as a beta-posterior (reliabilityPosterior), calculates speed as a composite of token-throughput and latency (speedScore), and factors in model tier through intelligenceComposite. The head-room factor (headroomFactor) monitors monthly budgets and per-window usage, demoting models before they exhaust quotas.

Thompson-Sampling Bandit Optimization

When the routing strategy enables sampling (sampled = true), the system applies Thompson sampling to the chain. This statistical method balances trying high-performing models against exploring potentially better alternatives, optimizing long-term success rates across the provider pool.

Rate Limit Penalties and Recovery

The penalty system distinguishes between generic failures (recordModelFailure) and rate-limit hits (recordRateLimitHit), with 429 responses carrying heavier penalties. Penalized models sink in the ordering chain until statistical decay restores their standing, creating automatic circuit-breaker behavior.

Practical Integration Examples

Direct cURL Request

Send a chat completion through the public endpoint:

curl -X POST https://api.free-llmapi.com/v1/chat/completions \
     -H "Authorization: Bearer $FREELLMAPI_KEY" \
     -H "Content-Type: application/json" \
     -d '{
           "model": "gpt-4o-mini",
           "messages": [{"role":"user","content":"What is the capital of France?"}],
           "max_tokens": 512
         }'

Node.js SDK Integration

Use the bundled SDK from cli/src/tools.ts for programmatic access:

import { fetchChatCompletion } from './cli/src/tools.js';

const response = await fetchChatCompletion({
  apiKey: process.env.FREELLMAPI_KEY,
  model: 'gpt-4o-mini',
  messages: [{ role: 'user', content: 'Explain quantum tunnelling in one sentence.' }],
  maxTokens: 256,
});

console.log(response.choices[0].message.content);

Manual Routing Inspection

Inspect the routing decision and fallback chain manually:

import { routeRequest, RouteError } from '../server/src/services/router.js';

try {
  const result = await routeRequest({
    model: 'gpt-4o-mini',
    messages: [{ role: 'user', content: 'Ping?' }],
    maxTokens: 128,
  });
  console.log('Routed to', result.provider.name, 'via key', result.keyLabel);
} catch (e) {
  if (e instanceof RouteError) {
    console.error('All models exhausted:', e.diagnostics);
  }
}

Summary

  • Authentication and rate-limiting occur first in server/src/middleware/ before any routing logic executes.
  • Token reservation via routingReserveTokens() prevents pool starvation by capping output allocations.
  • Model scoring uses scoreChainEntry() to combine reliability, speed, and intelligence metrics with real-time quota head-room.
  • Thompson-sampling bandits optimize the routing strategy when sampled = true, balancing exploration against exploitation.
  • API key selection follows auto (round-robin) or least-remaining strategies with lease-based concurrency enforcement in ratelimit.ts.
  • Provider abstraction through BaseProvider enables uniform handling of OpenAI, Anthropic, Google, and other backends.
  • Penalty decay automatically recovers models from 429 errors via recordRateLimitHit() and DECAY_INTERVAL_MS.

Frequently Asked Questions

How does FreeLLMAPI handle rate limiting across multiple providers?

FreeLLMAPI tracks rate-limit hits through recordRateLimitHit() in server/src/services/router.ts, which applies decay-weighted penalties to models returning 429 errors. The system maintains an in-memory statsCache that demotes penalized providers in the routing chain until their statistics recover, effectively creating a self-healing circuit breaker that avoids repeatedly hitting exhausted quotas.

What routing strategies are available in FreeLLMAPI?

The platform supports multiple strategies configurable via getRoutingStrategy: priority respects manual ordering, balanced optimizes for cost-performance ratios, and fastest prioritizes low-latency providers. When Thompson sampling is enabled (sampled = true), the orderChain() function applies probabilistic bandit algorithms to dynamically optimize selection based on historical performance metrics.

How does FreeLLMAPI choose between different API keys for the same provider?

Key selection follows the key_selection_strategy setting, supporting auto (round-robin distribution) or least-remaining (selecting the key with the most quota left). The acquireLease() and releaseLease() functions in server/src/services/ratelimit.ts enforce per-key concurrency limits, preventing any single credential from exceeding its parallel request capacity.

What happens when all models in the chain fail?

If the primary model fails, the router iterates through the ordered ChainRow list, attempting each candidate until one succeeds or the chain exhausts. Upon total failure, the system throws a RouteError containing diagnostic metadata via summarizeExhaustion, allowing client code to inspect which providers were attempted and why each failed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →