How FreeLLMAPI Manages Rate Limiting: A Deep Dive into Multi-Layer Quota Controls

FreeLLMAPI enforces rate limits through a hierarchical three-layer architecture that combines per-key sliding windows, provider-wide caps, and intelligent cooldown mechanisms, all centralized in server/src/services/ratelimit.ts.

The tashfeenahmed/freellmapi repository implements a sophisticated traffic management system designed to maximize free-tier LLM provider utilization without hitting hard quotas. Unlike simple token bucket algorithms, this implementation tracks requests per minute (rpm), tokens per minute (tpm), requests per day (rpd), and tokens per day (tpd) across multiple dimensions while maintaining ACID-like persistence through SQLite.

Three-Layer Rate Limiting Architecture

Per-Key, Per-Model Windows

At the granular level, FreeLLMAPI tracks consumption through sliding-window counters defined in server/src/services/ratelimit.ts. The functions requestCount() and tokenCount() query the SQLite rate_limit_usage table when available, falling back to in-memory windows (memoryRequestCount, memoryTokenCount) during database unavailability.

The gate functions canMakeRequest() and canUseTokens() execute the core admission logic. These functions combine persisted counts, active in-flight leases, and configured limits to determine whether a new call is permitted. When checking a request, the system validates against both minute-level (rpm/tpm) and day-level (rpd/tpd) constraints simultaneously.

Provider-Wide Daily Caps

FreeLLMAPI aggregates consumption across all models for a given provider key using buildWindowUsageSnapshot(). The canUseProvider() function enforces hard daily limits by comparing the aggregated providerDailyRequestCount against environment-specific caps.

Operators configure these limits through environment variables: PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> overrides the default values defined in DEFAULT_PROVIDER_DAILY_REQUEST_CAPS. This allows the system to adapt to fluctuating free-tier quotas from providers like OpenRouter (approximately 1000 requests/day) without code deployments.

Provider-Wide Minute Caps

Similar to daily enforcement, canUseProviderMinute() applies per-minute request limits using providerMinuteRequestCount. This prevents burst traffic from exhausting minute-level quotas, particularly relevant for providers like NVIDIA NIM that enforce approximately 40 requests per minute.

Cooldown Handling for Rate-Limit Errors

When upstream providers return 429 errors or quota-exceeded signals, the getCooldownDecisionForLimit() function calculates temporary backoff periods. The decision engine considers whether the quota is known, historical hit counts, and any Retry-After header values.

Local endpoints (loopback URLs) receive a fixed short cooldown (LOCAL_ENDPOINT_COOLDOWN_MS), while external providers follow escalating cooldown durations (COOLDOWN_DURATIONS). The system applies extended cooldowns when daily limits exhaust or when heuristic thresholds (two hits within one hour) trigger. These cooldowns persist in the rate_limit_cooldowns table and expire automatically via clearCooldown().

Core Rate Limiting Mechanisms

Lease-Based Concurrency Control

Before dispatching any request, acquireLease() creates an in-memory lease representing the anticipated resource consumption. This lease gets counted against both minute and day windows immediately, preventing race conditions between the check and act phases. The lease ensures that concurrent requests respect quota limits even before the database records the actual usage.

Persistent SQLite Storage

All successful request and token recordings flow through recordUsage(), which writes to the rate_limit_usage table. If the database write fails, the system gracefully degrades to in-memory counters, ensuring availability during SQLite outages while potentially sacrificing accuracy until recovery.

Snapshot Caching for Performance

To minimize database query overhead, FreeLLMAPI maintains a five-second TTL snapshot (windowUsageSnapshot) that aggregates per-key usage statistics. This caching layer allows canMakeRequest() and canUseTokens() to perform fast guard-rail calculations without executing multiple tiny queries for every incoming request.

Configuration and Environment Variables

FreeLLMAPI exposes provider-specific caps through environment variables, enabling runtime configuration of rate limiting behavior:

  • PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>: Sets the maximum daily requests for a provider account
  • PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>: Configures per-minute request thresholds

These variables interface directly with the limit-checking logic in server/src/services/ratelimit.ts (lines 61-68 and 95-103), allowing operators to adjust to evolving free-tier quotas without restarting the service or modifying source code.

Implementation Examples

Checking Request Permissions

import { canMakeRequest, canUseTokens } from './services/ratelimit.js';

// Verify if a key can issue another request for a specific model
const canCall = canMakeRequest('openrouter', 'gpt-4o-mini', 42, {
  rpm: 60,
  rpd: 1000,
  tpm: null,
  tpd: null,
});

if (canCall) {
  // Safe to dispatch the request
}

Validating Token Budgets

// Check token budget before sending large prompts
const canSpend = canUseTokens('openrouter', 'gpt-4o-mini', 42, 5000, {
  tpm: 200_000,
  tpd: null,
});

if (canSpend) {
  // Proceed with the API call
}

Provider-Wide Quota Verification

import { canUseProvider } from './services/ratelimit.js';

// Check if the entire provider account has remaining budget
const accountOk = canUseProvider('openrouter', 42);

if (!accountOk) {
  // Skip all OpenRouter models for this key today
}

Applying Post-Error Cooldowns

import { setCooldown } from './services/ratelimit.js';

// Apply a 2-minute cooldown after receiving a 429 response
setCooldown('openrouter', 'gpt-4o-mini', 42, 2 * 60_000, 'heuristic');

Summary

  • Multi-layer enforcement: FreeLLMAPI applies per-key/model windows, provider-wide daily caps, and provider-wide minute caps through server/src/services/ratelimit.ts.
  • Race-condition prevention: The acquireLease() function creates in-memory leases before request dispatch to prevent quota overruns during concurrent access.
  • Persistent tracking: SQLite tables rate_limit_usage and rate_limit_cooldowns maintain state across restarts, with automatic fallback to memory during database failures.
  • Intelligent backoff: The getCooldownDecisionForLimit() function implements escalating delays based on error types, headers, and historical patterns.
  • Environment configuration: Provider caps are tunable via PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> and PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM> variables.

Frequently Asked Questions

How does FreeLLMAPI prevent concurrent requests from exceeding rate limits?

The system implements lease-based concurrency control through the acquireLease() function in server/src/services/ratelimit.ts. Before dispatching any request, FreeLLMAPI reserves the expected capacity in sliding-window counters, ensuring that simultaneous checks respect the same quota limits without race conditions.

What happens to rate limit data if the SQLite database becomes unavailable?

FreeLLMAPI degrades gracefully to in-memory windows (memoryRequestCount, memoryTokenCount) when the rate_limit_usage table is inaccessible. The recordUsage() function handles write failures by maintaining counters in memory, preserving basic protection at the cost of persistence until database connectivity restores.

How does the cooldown system handle different types of rate limit errors?

The getCooldownDecisionForLimit() function analyzes error context to determine backoff durations. Local endpoints receive fixed short timeouts (LOCAL_ENDPOINT_COOLDOWN_MS), while external providers follow escalating cooldown patterns (COOLDOWN_DURATIONS). When providers include Retry-After headers, the system respects these values; otherwise, it applies heuristics based on daily limit exhaustion or hourly hit thresholds.

Can operators adjust rate limits without modifying code?

Yes. FreeLLMAPI reads provider-wide caps from environment variables (PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> and PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>) at runtime. These variables override default constants in server/src/services/ratelimit.ts, allowing immediate adaptation to changing free-tier quotas from providers like OpenRouter or NVIDIA NIM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →