How the FreeLLMAPI Health Service Works: Architecture, Probes, and Degradation Logic
FreeLLMAPI continuously monitors every enabled API key through a coordinated health subsystem that combines scheduled probes, cooldown recovery, and operational endpoints to route requests only to healthy providers.
The FreeLLMAPI health service is the gatekeeper that determines which LLM providers can receive traffic. Built into the tashfeenahmed/freellmapi repository, this system automatically validates credentials, detects failures, and degrades gracefully when upstream services falter. This article explains the three core components—scheduled health checks, cooldown probes, and status endpoints—that keep the gateway reliable.
Core Health Service Components
The health subsystem spans three coordinated parts across server/src/services/health.ts, server/src/services/cooldown-probe.ts, and server/src/routes/status.ts.
| Component | Purpose | Source File |
|---|---|---|
| Scheduled health checker | Probes each enabled key every ~5 minutes, updates status, auto-disables after failures | server/src/services/health.ts |
| Cooldown-probe | Lightweight per-minute check to clear rate-limit cooldowns without DB writes | server/src/services/cooldown-probe.ts |
| Operational endpoints | Expose aggregated health data (/v1/providers, /livez, /readyz) |
server/src/routes/status.ts |
Scheduled Health Pass
The scheduled health pass in server/src/services/health.ts runs the validation cycle that keeps provider status current.
Timing and Concurrency
The pass executes every 5 minutes with controlled randomness:
CHECK_INTERVAL_MS = 5 * 60 * 1000sets the base cadence- ±20% jitter via
nextHealthCheckDelayMsprevents lock-step probing across multiple gateway instances HEALTH_CHECK_CONCURRENCYlimits parallel probes to 8 keys (configurable)HEALTH_CHECK_MIN_SPACING_MSenforces minimum gaps between probes to the same provider
Keys checked less than 3.5 minutes ago are skipped via RECENT_CHECK_SKIP_MS. Keys in error state are never skipped—they need fresh validation to recover.
Probe Execution Flow
The checkKeyHealth function implements the validation sequence at lines 97-103:
// Simplified probe logic from server/src/services/health.ts
async function checkKeyHealth(key: ApiKey): Promise<HealthResult> {
// 1. Decrypt the stored key
const decryptedKey = await decryptKey(key.encryptedKey);
// 2. Resolve provider implementation
const provider = resolveProvider(key.platform);
// 3. Validate through the key's own proxy (if configured)
const result = await provider.validateKey(decryptedKey, {
proxy: key.proxyConfig
});
// 4. Interpret and store status
const status = result.valid ? 'healthy' : 'invalid';
await updateKeyStatus(key.id, status, result.error);
return { keyId: key.id, status };
}
Failure Handling and Auto-Disable
The health pass distinguishes between credential failures and transport errors:
invalidresponses: Increment failure counter viarecordInvalidFailure; auto-disable after 3 consecutive failures (lines 68-80)- Transport errors (DNS timeout, TLS failures): Recorded as
errorstatus, do not increment failure counter (lines 31-55)
This separation prevents transient network issues from permanently disabling valid keys.
Triggering Health Passes Manually
For debugging or CI pipelines, force an immediate scan:
import { checkAllKeys } from './services/health.js';
// Bypass recent-check skipping
await checkAllKeys({ force: true });
Cooldown-Probe Service
The cooldown-probe in server/src/services/cooldown-probe.ts operates independently from the main health pass. It runs approximately once per minute through the provider-quota scheduler.
Key characteristics of probeKeyValidity:
- Uses the same cheap
validateKeycall as the main health pass - Never writes to the database—no status updates, no timestamp changes, no failure counter increments
- Never affects the failure counter
- Sole purpose: detect when a rate-limited provider has recovered and clear its cooldown state
This lightweight design allows rapid recovery detection without the overhead of full health passes.
Operational Health Endpoints
The status endpoints in server/src/routes/status.ts expose health data to external orchestrators.
/v1/providers Aggregated Status
The providersRouter.get('/providers', …) endpoint (lines 6-8) returns provider roll-ups:
curl -H "Authorization: Bearer $FREE_LLMAPI_KEY" \
https://api.freellmapi.com/v1/providers
Response structure:
{
"providers": [
{
"platform": "openrouter",
"name": "OpenRouter",
"status": "healthy",
"keys": 3,
"requests_remaining_pct": 92
},
{
"platform": "groq",
"name": "Groq",
"status": "rate_limited",
"keys": 2,
"resume_at": "2024-07-01T12:34:56.000Z"
}
],
"counts": { "healthy": 1, "rate_limited": 1, "invalid": 0, "unknown": 0 }
}
Status values derive from KeyStatus in shared/types.ts: 'healthy' | 'rate_limited' | 'invalid' | 'error' | 'unknown'.
Liveness and Readiness Checks
/livez(lines 32-52): Confirms database connectivity and encryption key availability/readyz(lines 54-90): Confirms at least one upstream provider is serviceable
These enable Kubernetes-style health probes and load-balancer integration.
Degradation Mode
When healthy providers fall below a configured threshold, the router enters degradation mode via server/src/services/degradation.ts. The health pass updates this state after each run through updateDegradationState.
In degradation mode:
- Router prefers the healthiest remaining providers
- May reject traffic earlier to protect callers from prolonged outages
- Threshold and behavior are configurable per deployment
Request-Time Health Updates
Successful requests can optimistically mark keys healthy:
import { markKeyHealthyFromRequest } from './services/health.js';
// Called from request handler after successful response
markKeyHealthyFromRequest(keyId);
This bypasses waiting for the next scheduled probe when a key proves functional through actual traffic.
Health Service File Reference
| File | Responsibility |
|---|---|
server/src/services/health.ts |
Core health checker, status persistence, auto-disable logic |
server/src/services/cooldown-probe.ts |
Rate-limit recovery detection without DB overhead |
server/src/routes/status.ts |
Public health endpoints (/v1/providers, /livez, /readyz) |
server/src/services/degradation.ts |
Degraded mode switching based on provider health ratio |
shared/types.ts |
Type definitions for KeyStatus enum |
Summary
- Scheduled health passes run every ~5 minutes with jitter, validating up to 8 keys concurrently while respecting minimum spacing between probes to the same provider
checkKeyHealthdecrypts keys, resolves providers, validates through configured proxies, and writes status to theapi_keystable; three consecutiveinvalidresults trigger auto-disable- Cooldown-probes run ~1 minute intervals for rapid rate-limit recovery without database writes or failure counting
- Status endpoints expose aggregated provider health, liveness, and readiness for external orchestration
- Degradation mode activates when healthy provider ratios drop, shifting routing strategy to protect callers
Frequently Asked Questions
How often does FreeLLMAPI check provider health?
The scheduled health pass runs every 5 minutes (CHECK_INTERVAL_MS = 5 * 60 * 1000) with ±20% random jitter to prevent synchronized probing. The separate cooldown-probe runs approximately once per minute for rate-limit recovery detection.
What happens when an API key fails validation three times?
After three consecutive invalid validation responses, the key is automatically disabled. This counter increments only on explicit invalid results from provider.validateKey—transport errors like DNS timeouts or TLS failures are recorded as error status without affecting the failure counter.
Can I force a health check outside the normal schedule?
Yes. Import checkAllKeys from server/src/services/health.js and call await checkAllKeys({ force: true }) to bypass the RECENT_CHECK_SKIP_MS window and trigger immediate validation of all enabled keys.
How does FreeLLMAPI expose health information to load balancers?
The /livez endpoint confirms basic connectivity and encryption key availability; /readyz confirms at least one upstream provider is serviceable. The /v1/providers endpoint returns detailed provider status, key counts, and quota information for routing decisions by meta-gateways like LibreChat or LiteLLM Proxy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →