How Rate Limiting Works in the G0DM0D3 API: Implementation and Configuration

The G0DM0D3 API enforces tier-based usage limits through a middleware that maintains an in-memory Map<string, number[]> to track request timestamps per API key, applying sliding-window checks for per-minute, per-day, and total limits before rejecting excess requests with HTTP 429 responses.

The G0DM0D3 project implements request throttling through a dedicated middleware layer defined in src/api/middleware/rateLimit.ts. This system ensures fair resource allocation across different subscription tiers by tracking usage patterns in real-time without requiring external database calls.

Architecture Overview

The rate limiting implementation relies on a three-layer approach: tier-based configuration, in-memory state management, and sliding-window validation. Each API key is associated with a subscription tier defined in src/api/lib/tiers.ts, which determines specific constraints on request volume. The middleware intercepts every authenticated request to validate against these constraints before allowing access to route handlers.

Tier-Based Configuration System

The foundation of the rate limiting strategy resides in src/api/lib/tiers.ts, where each subscription tier (Free, Pro, Enterprise) declares specific constraints:

  • rateLimit.total: Maximum cumulative requests allowed per API key
  • rateLimit.perMinute: Burst capacity for short time windows
  • rateLimit.perDay: Daily allocation limits

When a tier omits specific values, the middleware falls back to environment variables: RATE_LIMIT_TOTAL, RATE_LIMIT_PER_MINUTE, and RATE_LIMIT_PER_DAY. These defaults are parsed as integers at runtime to ensure consistent numeric comparisons.

In-Memory Request Tracking

Rather than querying a database on every request, the middleware utilizes a high-performance in-memory structure defined in src/api/middleware/rateLimit.ts:

rateLimitMap: Map<string, number[]>

The Map key represents the API key identifier, while the value stores an array of UNIX timestamps marking each request's arrival time. This design enables O(1) lookups with minimal latency overhead, eliminating I/O bottlenecks that would occur with persistent storage solutions.

Sliding Window Enforcement Logic

For every incoming request, the middleware performs three concurrent checks against the timestamp array:

  1. Per-minute validation: Filters the array to retain only timestamps from the last 60 seconds, comparing the count against MINUTE_LIMIT
  2. Per-day validation: Retains only timestamps from the last 24 hours (86,400 seconds), checked against DAY_LIMIT
  3. Total usage validation: Compares the complete array length against TOTAL_LIMIT

If any threshold exceeds its configured limit, the middleware immediately terminates the request chain before route handlers execute.

HTTP 429 Error Handling

When limits are breached, the API returns a structured JSON response with standard HTTP semantics:

{
  "error": "Rate limit exceeded",
  "detail": "Per-minute limit of 60 requests reached"
}

This response includes the HTTP 429 status code, preventing unnecessary processing of rejected requests and signaling to clients that they should implement exponential backoff strategies.

Memory Management and Housekeeping

To prevent unbounded memory growth from abandoned or infrequently used API keys, the implementation includes automatic housekeeping logic. When the rateLimitMap exceeds 10,000 distinct keys, the system purges the oldest entries based on earliest activity timestamps. This mechanism maintains predictable memory consumption in long-running production environments.

Server Integration and Route Protection

The middleware applies globally to protected routes in src/api/server.ts using Express.js composition:

app.use('/v1/chat', apiKeyAuth, rateLimit, chatRoutes)

This declaration ensures every chat completion request passes through authentication followed by rate limiting before reaching business logic. The Hugging Face variant at src/HF/api/middleware/rateLimit.ts implements identical logic for the HF-hosted deployment, ensuring consistent enforcement across both hosting environments.

Practical Usage Examples

Testing within limits:

curl -H "X-API-Key: abc123" https://api.g0dm0d3.com/v1/chat \
  -d '{"prompt":"Hello"}' -H "Content-Type: application/json"

Triggering rate limit response:


# After exceeding per-minute threshold

curl -i -H "X-API-Key: abc123" https://api.g0dm0d3.com/v1/chat

Expected response headers and body:


HTTP/1.1 429 Too Many Requests
Content-Type: application/json

{
  "error": "Rate limit exceeded",
  "detail": "Per-minute limit of 60 requests reached"
}

Client handling in Node.js:

async function callChat(payload) {
  const resp = await fetch('https://api.g0dm0d3.com/v1/chat', {
    method: 'POST',
    headers: { 
      'X-API-Key': API_KEY, 
      'Content-Type': 'application/json' 
    },
    body: JSON.stringify(payload),
  })

  if (resp.status === 429) {
    const info = await resp.json()
    console.warn('Rate limit hit:', info.detail)
    // Implement exponential backoff here
    await new Promise(r => setTimeout(r, 60000))
    return callChat(payload)
  }
  return resp.json()
}

Summary

  • The rate limiting system uses an in-memory Map<string, number[]> structure in src/api/middleware/rateLimit.ts to track request timestamps per API key
  • Configuration derives from src/api/lib/tiers.ts with fallbacks to environment variables RATE_LIMIT_TOTAL, RATE_LIMIT_PER_MINUTE, and RATE_LIMIT_PER_DAY
  • Three-tier validation checks per-minute, per-day, and total limits using sliding window calculations on UNIX timestamp arrays
  • Violations return HTTP 429 "Too Many Requests" responses immediately, halting further request processing to conserve compute resources
  • Automatic cleanup triggers when tracking exceeds 10,000 API keys to maintain predictable memory usage
  • Middleware registration in src/api/server.ts ensures consistent enforcement across all protected endpoints, with identical logic present in the Hugging Face variant

Frequently Asked Questions

How are rate limits configured for different user tiers?

Limits are defined in src/api/lib/tiers.ts where each tier object specifies rateLimit.total, rateLimit.perMinute, and rateLimit.perDay values. If a tier configuration omits these properties, the system automatically falls back to the corresponding environment variables, parsing them as integers at runtime.

What happens when a rate limit is exceeded?

The middleware immediately responds with HTTP status 429 "Too Many Requests" and a JSON payload indicating which specific limit triggered the block. The request never reaches route handlers, ensuring zero compute waste for rejected traffic and allowing clients to detect throttling programmatically.

How does the API track usage without a database?

The implementation uses a Map<string, number[]> called rateLimitMap that stores UNIX timestamps in memory. This approach eliminates database latency while providing microsecond-level tracking precision, though data persists only for the server's runtime duration and is cleared on restart.

Is rate limiting identical between the main API and Hugging Face deployment?

Yes. The Hugging Face variant at src/HF/api/middleware/rateLimit.ts contains the same logic as the main implementation in src/api/middleware/rateLimit.ts, ensuring consistent behavior across both hosting environments. Both versions reference the same tier configuration and environment variable fallbacks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →