How to View FreeLLMAPI Rate Limit Counters: HTTP Headers and Provider Status Guide
FreeLLMAPI exposes rate limit counters through OpenAI-style HTTP headers on every request and aggregates provider-specific throttling status via the /v1/providers endpoint.
FreeLLMAPI is an open-source unified API gateway that aggregates multiple LLM providers into a single interface. Monitoring your usage against rate limits is essential for production workloads to prevent service interruptions. According to the tashfeenahmed/freellmapi source code, the system tracks both per-IP request quotas and upstream provider status through specific middleware and route handlers.
Checking Per-IP Rate Limits via HTTP Headers
The server/src/middleware/rateLimit.ts file implements middleware that attaches three standard headers to every response from the public /v1/* proxy. These headers follow the OpenAI format and reflect the global request quota for the caller's IP address.
Understanding the X-RateLimit Headers
At lines 60-62 of the middleware file, the system sets the following headers:
X-RateLimit-Limit: The maximum allowed calls per minute (defaults to 120 RPM)X-RateLimit-Remaining: How many calls are left in the current windowX-RateLimit-Reset: Epoch timestamp in seconds when the window expires
Detecting 429 Errors and Retry-After
When the per-IP limit is exceeded, the middleware at lines 64-73 returns HTTP 429 with a Retry-After header indicating seconds until the reset. This allows clients to implement automatic backoff logic.
Use the following curl command to inspect these headers on any API call:
curl -s -D - http://localhost:3001/v1/models \
-H "Authorization: Bearer freellmapi-your-unified-key" \
-H "Content-Type: application/json" \
-o /dev/null
Typical successful response headers:
HTTP/1.1 200 OK
X-RateLimit-Limit: 120
X-RateLimit-Remaining: 87
X-RateLimit-Reset: 1725629385
When rate limited:
HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 120
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1725629410
Monitoring Global Rate Limits via the Providers Endpoint
Beyond per-IP tracking, FreeLLMAPI exposes aggregated rate limit counters through the /v1/providers endpoint defined in server/src/routes/status.ts. This route queries the SQLite database and returns a JSON summary of all upstream providers.
The counts Object and Provider Status
The providersRouter.get('/providers') handler starting at line 29 returns a response containing a counts object with four keys: healthy, rate_limited, invalid, and unknown. The logic that identifies rate-limited providers resides at lines 74-78 of the file.
Tracking Cooldown Periods with resume_at
Each provider marked as rate-limited includes a resume_at ISO 8601 timestamp indicating when the cooldown period ends and requests can resume.
Query the endpoint using:
curl -s http://localhost:3001/v1/providers \
-H "Authorization: Bearer freellmapi-your-unified-key"
Example JSON response:
{
"counts": {
"healthy": 3,
"rate_limited": 1,
"invalid": 0,
"unknown": 0
},
"providers": [
{
"platform": "groq",
"name": "Groq",
"status": "healthy",
"keys": 2
},
{
"platform": "openai",
"name": "OpenAI",
"status": "rate_limited",
"keys": 1,
"resume_at": "2024-09-06T12:34:56.000Z"
}
]
}
The counts.rate_limited field shows how many upstream providers are currently throttled, while individual resume_at fields provide exact unlock times.
Summary
- Per-IP counters appear in
X-RateLimit-*headers set byserver/src/middleware/rateLimit.ts(lines 60-62), showing remaining calls and reset times for the default 120 RPM limit. - 429 responses include
Retry-Afterheaders when limits are exceeded (lines 64-73), enabling client-side retry logic. - Aggregated status is available via the
/v1/providersendpoint implemented inserver/src/routes/status.ts(lines 29-33), which returns acountsobject with current provider states. - Rate-limited providers expose
resume_attimestamps (lines 74-78) indicating when service resumes.
Frequently Asked Questions
What is the default rate limit for FreeLLMAPI requests?
The default rate limit is 120 requests per minute (RPM) per IP address. This value appears in the X-RateLimit-Limit header as implemented in server/src/middleware/rateLimit.ts.
How do I determine when my rate limit window resets?
Check the X-RateLimit-Reset header on any API response. This value represents Unix epoch seconds indicating when the current window expires and your quota refreshes to the full 120 requests.
What does the counts.rate_limited field indicate in the providers response?
The counts.rate_limited field in the /v1/providers JSON response indicates how many upstream LLM providers are currently throttled due to their own rate limits. This number is calculated in server/src/routes/status.ts at lines 74-78 by querying the SQLite database for providers with active cooldown periods.
Where does FreeLLMAPI store rate limit tracking data?
FreeLLMAPI stores provider-specific rate limit status and cooldown timestamps in a SQLite database. The providersRouter in server/src/routes/status.ts queries this database to build the aggregated counts and individual provider status objects returned by the API.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →