# How Sub2API Performs Rate Limiting: Architecture and Implementation

> Discover how Sub2API implements rate limiting. Learn about its architecture, the centralized RateLimitService, and how it enforces per-account quotas using X-RateLimit headers.

- Repository: [Wesley Liddick/sub2api](https://github.com/Wei-Shaw/sub2api)
- Tags: architecture
- Published: 2026-08-23

---

**Sub2API enforces rate limiting through a centralized `RateLimitService` that evaluates per-account quotas against usage repositories, applies configurable multipliers, and communicates limits via standard `X-RateLimit-*` HTTP headers.**

The open-source **Sub2API** project (available at `Wei-Shaw/sub2api`) implements a multi-layered rate limiting strategy to protect downstream AI providers like OpenAI and Gemini from abuse. The system coordinates between persistent storage, temporary caching, and HTTP middleware to enforce requests-per-minute (RPM) and tokens-per-minute (TPM) constraints dynamically according to the source code.

## Core Architecture and Components

The rate limiting system relies on a pipeline of specialized components wired together through dependency injection.

### RateLimitService

The **`RateLimitService`** in [`backend/internal/service/ratelimit_service.go`](https://github.com/Wei-Shaw/sub2api/blob/main/backend/internal/service/ratelimit_service.go) serves as the central authority for quota enforcement. It aggregates configuration data, account-level overrides, and real-time usage statistics to determine whether to allow, delay, or reject incoming requests according to the repository structure.

### Supporting Repositories and Caches

- **`AccountRepository`**: Persists per-account limits and custom multipliers (e.g., elevated quotas for specific user tiers).
- **`UsageRepository`**: Tracks real-time request and token counters within the current time window.
- **`TempUnschedulableCache`**: Stores temporary cooldown flags when accounts breach limits, preventing scheduled background jobs from executing during backoff periods.
- **`GeminiQuotaService`**: An optional adapter that handles Google Gemini-specific quota constraints separately from generic account limits.

### Response Header Middleware

The **`ResponseHeaders`** utility in [`backend/internal/util/responseheaders/responseheaders.go`](https://github.com/Wei-Shaw/sub2api/blob/main/backend/internal/util/responseheaders/responseheaders.go) (line 26) defines the standard `X-RateLimit-*` header names, including `x-ratelimit-limit-requests`, `x-ratelimit-remaining-requests`, and `x-ratelimit-reset-requests`, which the middleware attaches to every HTTP response.

## Rate Limit Evaluation Flow

For each incoming request, `RateLimitService` executes the following decision sequence:

1. **Lookup Limits**: Retrieve the base RPM and TPM values from the global configuration and any account-specific overrides stored in `AccountRepository`.
2. **Read Current Usage**: Query `UsageRepository` for the number of requests and tokens already consumed within the current sliding window.
3. **Apply Multipliers**: Scale limits according to the account's `rate_multiplier` field to support tiered pricing or promotional quotas.
4. **Check Thresholds**:
   - If the request would exceed the limit, reject it with HTTP 429.
   - If the request approaches the limit, set a temporary unschedulable flag in `TempUnschedulableCache` to pause background job scheduling for a calculated backoff period.
5. **Send Response Headers**: Populate `X-RateLimit-*` headers so clients can programmatically adjust their request rate.

## HTTP Headers and Client Communication

Sub2API follows industry standards by exposing quota state through response headers. As implemented in [`backend/internal/util/responseheaders/responseheaders.go`](https://github.com/Wei-Shaw/sub2api/blob/main/backend/internal/util/responseheaders/responseheaders.go), the platform returns:

- **`x-ratelimit-limit-requests`**: The maximum number of requests allowed in the current window.
- **`x-ratelimit-remaining-requests`**: The number of requests remaining before hitting the limit.
- **`x-ratelimit-reset-requests`**: Unix timestamp indicating when the request quota resets.
- **Token equivalents**: Corresponding headers for TPM (tokens per minute) limits.

This allows API consumers to implement client-side throttling without hard-coding limit values.

## Handling Model-Specific and Upstream Limits

### Model-Specific Constraints

Certain AI models (e.g., Codex and Gemini variants) carry distinct rate limits separate from account-level quotas. When a request targets one of these models, `RateLimitService` first validates against the model-specific limit before falling back to the account default, as handled by the optional `GeminiQuotaService` integration.

### Upstream Error Translation

When downstream providers return HTTP 429 or 403 rate-limit errors, the service extracts the `Retry-After` header value. The `HandleUpstreamError` method converts this into a platform-wide throttling response while preserving the original cooldown timing, ensuring users receive accurate backoff instructions rather than generic rejections.

## Code Implementation Examples

### Wiring Dependencies

The `ProvideRateLimitService` function in [`backend/internal/service/wire.go`](https://github.com/Wei-Shaw/sub2api/blob/main/backend/internal/service/wire.go) assembles the service with its required dependencies:

```go
func ProvideRateLimitService(
    accountRepo AccountRepository,
    usageRepo UsageRepository,
    cfg *config.Config,
    geminiQuotaService *GeminiQuotaService,
    tempUnschedCache *TempUnschedulableCache,
) *RateLimitService {
    return NewRateLimitService(accountRepo, usageRepo, cfg, geminiQuotaService, tempUnschedCache)
}

```

### Handler Integration

API handlers check quotas before forwarding requests to upstream providers:

```go
func (h *Handler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
    // Extract user/account ID from request context
    if err := h.rateLimitSvc.CheckRequest(ctx, accountID, modelID, tokenCount); err != nil {
        http.Error(w, err.Error(), http.StatusTooManyRequests)
        return
    }

    // Forward to upstream API if limit check passes
    h.proxy.ServeHTTP(w, r)
}

```

### Middleware Header Injection

The response middleware injects rate limiting headers after request processing:

```go
func RateLimitHeaders(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        next.ServeHTTP(w, r)

        // Headers populated by RateLimitService into context
        w.Header().Set("x-ratelimit-limit-requests", ctx.Value("limitRequests").(string))
        w.Header().Set("x-ratelimit-remaining-requests", ctx.Value("remainingRequests").(string))
        w.Header().Set("x-ratelimit-reset-requests", ctx.Value("resetRequests").(string))
    })
}

```

### Upstream Error Handling

Convert provider-specific rate limit errors into platform responses:

```go
func (s *RateLimitService) HandleUpstreamError(err error, accountID string) error {
    if errors.Is(err, upstream.ErrRateLimited) {
        cooldown := parseRetryAfter(err) // Extract from Retry-After header
        s.tempUnschedCache.SetCooldown(accountID, cooldown)
        return fmt.Errorf("rate limit exceeded, retry after %s", cooldown)
    }
    return err
}

```

## Summary

- **Sub2API rate limiting** centers on the `RateLimitService` located in [`backend/internal/service/ratelimit_service.go`](https://github.com/Wei-Shaw/sub2api/blob/main/backend/internal/service/ratelimit_service.go).
- The system uses `AccountRepository` and `UsageRepository` to enforce per-account RPM and TPM quotas with configurable multipliers.
- Temporary cooldowns are managed through `TempUnschedulableCache` to protect scheduled job queues.
- Standard **`X-RateLimit-*`** headers are defined in [`backend/internal/util/responseheaders/responseheaders.go`](https://github.com/Wei-Shaw/sub2api/blob/main/backend/internal/util/responseheaders/responseheaders.go) and injected by middleware.
- Model-specific limits (Gemini, Codex) receive specialized handling through dedicated quota services.
- Upstream provider rate limits trigger automatic cooldown synchronization via `HandleUpstreamError`.

## Frequently Asked Questions

### How does Sub2API handle concurrent requests during rate limiting?

The `RateLimitService` evaluates each request atomically against the `UsageRepository` counters. If multiple simultaneous requests would collectively exceed the quota, the service rejects excess requests with HTTP 429 while allowing compliant requests through, ensuring precise enforcement without overages.

### What happens when a user hits a rate limit on a specific model like Gemini?

When targeting model-specific endpoints, Sub2API first checks the dedicated `GeminiQuotaService` limits before evaluating account-wide quotas. If the model limit is breached, the request is rejected immediately with model-specific error messaging, preventing consumption of the user's general account quota.

### Can rate limits be customized per user account?

Yes. The `AccountRepository` stores per-account overrides including custom `rate_multiplier` values. During the evaluation flow, `RateLimitService` applies these multipliers to the base configuration limits, enabling tiered access levels for different customer segments without code changes.

### Where are the rate limit headers defined in the codebase?

The header constants are defined in [`backend/internal/util/responseheaders/responseheaders.go`](https://github.com/Wei-Shaw/sub2api/blob/main/backend/internal/util/responseheaders/responseheaders.go) at line 26, which specifies keys like `x-ratelimit-limit-requests` and `x-ratelimit-remaining-requests`. The middleware references these constants to ensure consistent header naming across all API responses.