# How Rate Limiting Works in the G0DM0D3 API: Implementation and Configuration

> Discover how G0DM0D3 API implements rate limiting. Learn about tier-based usage limits, sliding-window checks, and HTTP 429 responses for efficient API management.

- Repository: [pliny/G0DM0D3](https://github.com/elder-plinius/G0DM0D3)
- Tags: how-to-guide
- Published: 2026-07-19

---

**The G0DM0D3 API enforces tier-based usage limits through a middleware that maintains an in-memory `Map<string, number[]>` to track request timestamps per API key, applying sliding-window checks for per-minute, per-day, and total limits before rejecting excess requests with HTTP 429 responses.**

The G0DM0D3 project implements request throttling through a dedicated middleware layer defined in [`src/api/middleware/rateLimit.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/middleware/rateLimit.ts). This system ensures fair resource allocation across different subscription tiers by tracking usage patterns in real-time without requiring external database calls.

## Architecture Overview

The rate limiting implementation relies on a three-layer approach: tier-based configuration, in-memory state management, and sliding-window validation. Each API key is associated with a subscription tier defined in [`src/api/lib/tiers.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/lib/tiers.ts), which determines specific constraints on request volume. The middleware intercepts every authenticated request to validate against these constraints before allowing access to route handlers.

## Tier-Based Configuration System

The foundation of the rate limiting strategy resides in [`src/api/lib/tiers.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/lib/tiers.ts), where each subscription tier (Free, Pro, Enterprise) declares specific constraints:

- **`rateLimit.total`**: Maximum cumulative requests allowed per API key
- **`rateLimit.perMinute`**: Burst capacity for short time windows
- **`rateLimit.perDay`**: Daily allocation limits

When a tier omits specific values, the middleware falls back to environment variables: `RATE_LIMIT_TOTAL`, `RATE_LIMIT_PER_MINUTE`, and `RATE_LIMIT_PER_DAY`. These defaults are parsed as integers at runtime to ensure consistent numeric comparisons.

## In-Memory Request Tracking

Rather than querying a database on every request, the middleware utilizes a high-performance in-memory structure defined in [`src/api/middleware/rateLimit.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/middleware/rateLimit.ts):

```typescript
rateLimitMap: Map<string, number[]>

```

The Map key represents the API key identifier, while the value stores an array of UNIX timestamps marking each request's arrival time. This design enables **O(1) lookups** with minimal latency overhead, eliminating I/O bottlenecks that would occur with persistent storage solutions.

## Sliding Window Enforcement Logic

For every incoming request, the middleware performs three concurrent checks against the timestamp array:

1. **Per-minute validation**: Filters the array to retain only timestamps from the last 60 seconds, comparing the count against `MINUTE_LIMIT`
2. **Per-day validation**: Retains only timestamps from the last 24 hours (86,400 seconds), checked against `DAY_LIMIT`
3. **Total usage validation**: Compares the complete array length against `TOTAL_LIMIT`

If any threshold exceeds its configured limit, the middleware immediately terminates the request chain before route handlers execute.

## HTTP 429 Error Handling

When limits are breached, the API returns a structured JSON response with standard HTTP semantics:

```json
{
  "error": "Rate limit exceeded",
  "detail": "Per-minute limit of 60 requests reached"
}

```

This response includes the **HTTP 429 status code**, preventing unnecessary processing of rejected requests and signaling to clients that they should implement exponential backoff strategies.

## Memory Management and Housekeeping

To prevent unbounded memory growth from abandoned or infrequently used API keys, the implementation includes automatic housekeeping logic. When the `rateLimitMap` exceeds **10,000 distinct keys**, the system purges the oldest entries based on earliest activity timestamps. This mechanism maintains predictable memory consumption in long-running production environments.

## Server Integration and Route Protection

The middleware applies globally to protected routes in [`src/api/server.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/server.ts) using Express.js composition:

```typescript
app.use('/v1/chat', apiKeyAuth, rateLimit, chatRoutes)

```

This declaration ensures every chat completion request passes through authentication followed by rate limiting before reaching business logic. The Hugging Face variant at [`src/HF/api/middleware/rateLimit.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/HF/api/middleware/rateLimit.ts) implements identical logic for the HF-hosted deployment, ensuring consistent enforcement across both hosting environments.

## Practical Usage Examples

Testing within limits:

```bash
curl -H "X-API-Key: abc123" https://api.g0dm0d3.com/v1/chat \
  -d '{"prompt":"Hello"}' -H "Content-Type: application/json"

```

Triggering rate limit response:

```bash

# After exceeding per-minute threshold

curl -i -H "X-API-Key: abc123" https://api.g0dm0d3.com/v1/chat

```

Expected response headers and body:

```

HTTP/1.1 429 Too Many Requests
Content-Type: application/json

{
  "error": "Rate limit exceeded",
  "detail": "Per-minute limit of 60 requests reached"
}

```

Client handling in Node.js:

```javascript
async function callChat(payload) {
  const resp = await fetch('https://api.g0dm0d3.com/v1/chat', {
    method: 'POST',
    headers: { 
      'X-API-Key': API_KEY, 
      'Content-Type': 'application/json' 
    },
    body: JSON.stringify(payload),
  })

  if (resp.status === 429) {
    const info = await resp.json()
    console.warn('Rate limit hit:', info.detail)
    // Implement exponential backoff here
    await new Promise(r => setTimeout(r, 60000))
    return callChat(payload)
  }
  return resp.json()
}

```

## Summary

- The rate limiting system uses an in-memory `Map<string, number[]>` structure in [`src/api/middleware/rateLimit.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/middleware/rateLimit.ts) to track request timestamps per API key
- Configuration derives from [`src/api/lib/tiers.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/lib/tiers.ts) with fallbacks to environment variables `RATE_LIMIT_TOTAL`, `RATE_LIMIT_PER_MINUTE`, and `RATE_LIMIT_PER_DAY`
- Three-tier validation checks per-minute, per-day, and total limits using sliding window calculations on UNIX timestamp arrays
- Violations return HTTP 429 "Too Many Requests" responses immediately, halting further request processing to conserve compute resources
- Automatic cleanup triggers when tracking exceeds 10,000 API keys to maintain predictable memory usage
- Middleware registration in [`src/api/server.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/server.ts) ensures consistent enforcement across all protected endpoints, with identical logic present in the Hugging Face variant

## Frequently Asked Questions

### How are rate limits configured for different user tiers?

Limits are defined in [`src/api/lib/tiers.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/lib/tiers.ts) where each tier object specifies `rateLimit.total`, `rateLimit.perMinute`, and `rateLimit.perDay` values. If a tier configuration omits these properties, the system automatically falls back to the corresponding environment variables, parsing them as integers at runtime.

### What happens when a rate limit is exceeded?

The middleware immediately responds with HTTP status 429 "Too Many Requests" and a JSON payload indicating which specific limit triggered the block. The request never reaches route handlers, ensuring zero compute waste for rejected traffic and allowing clients to detect throttling programmatically.

### How does the API track usage without a database?

The implementation uses a `Map<string, number[]>` called `rateLimitMap` that stores UNIX timestamps in memory. This approach eliminates database latency while providing microsecond-level tracking precision, though data persists only for the server's runtime duration and is cleared on restart.

### Is rate limiting identical between the main API and Hugging Face deployment?

Yes. The Hugging Face variant at [`src/HF/api/middleware/rateLimit.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/HF/api/middleware/rateLimit.ts) contains the same logic as the main implementation in [`src/api/middleware/rateLimit.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/api/middleware/rateLimit.ts), ensuring consistent behavior across both hosting environments. Both versions reference the same tier configuration and environment variable fallbacks.