How Sub2API Performs Rate Limiting: Architecture and Implementation

Sub2API enforces rate limiting through a centralized RateLimitService that evaluates per-account quotas against usage repositories, applies configurable multipliers, and communicates limits via standard X-RateLimit-* HTTP headers.

The open-source Sub2API project (available at Wei-Shaw/sub2api) implements a multi-layered rate limiting strategy to protect downstream AI providers like OpenAI and Gemini from abuse. The system coordinates between persistent storage, temporary caching, and HTTP middleware to enforce requests-per-minute (RPM) and tokens-per-minute (TPM) constraints dynamically according to the source code.

Core Architecture and Components

The rate limiting system relies on a pipeline of specialized components wired together through dependency injection.

RateLimitService

The RateLimitService in backend/internal/service/ratelimit_service.go serves as the central authority for quota enforcement. It aggregates configuration data, account-level overrides, and real-time usage statistics to determine whether to allow, delay, or reject incoming requests according to the repository structure.

Supporting Repositories and Caches

  • AccountRepository: Persists per-account limits and custom multipliers (e.g., elevated quotas for specific user tiers).
  • UsageRepository: Tracks real-time request and token counters within the current time window.
  • TempUnschedulableCache: Stores temporary cooldown flags when accounts breach limits, preventing scheduled background jobs from executing during backoff periods.
  • GeminiQuotaService: An optional adapter that handles Google Gemini-specific quota constraints separately from generic account limits.

Response Header Middleware

The ResponseHeaders utility in backend/internal/util/responseheaders/responseheaders.go (line 26) defines the standard X-RateLimit-* header names, including x-ratelimit-limit-requests, x-ratelimit-remaining-requests, and x-ratelimit-reset-requests, which the middleware attaches to every HTTP response.

Rate Limit Evaluation Flow

For each incoming request, RateLimitService executes the following decision sequence:

  1. Lookup Limits: Retrieve the base RPM and TPM values from the global configuration and any account-specific overrides stored in AccountRepository.
  2. Read Current Usage: Query UsageRepository for the number of requests and tokens already consumed within the current sliding window.
  3. Apply Multipliers: Scale limits according to the account's rate_multiplier field to support tiered pricing or promotional quotas.
  4. Check Thresholds:
    • If the request would exceed the limit, reject it with HTTP 429.
    • If the request approaches the limit, set a temporary unschedulable flag in TempUnschedulableCache to pause background job scheduling for a calculated backoff period.
  5. Send Response Headers: Populate X-RateLimit-* headers so clients can programmatically adjust their request rate.

HTTP Headers and Client Communication

Sub2API follows industry standards by exposing quota state through response headers. As implemented in backend/internal/util/responseheaders/responseheaders.go, the platform returns:

  • x-ratelimit-limit-requests: The maximum number of requests allowed in the current window.
  • x-ratelimit-remaining-requests: The number of requests remaining before hitting the limit.
  • x-ratelimit-reset-requests: Unix timestamp indicating when the request quota resets.
  • Token equivalents: Corresponding headers for TPM (tokens per minute) limits.

This allows API consumers to implement client-side throttling without hard-coding limit values.

Handling Model-Specific and Upstream Limits

Model-Specific Constraints

Certain AI models (e.g., Codex and Gemini variants) carry distinct rate limits separate from account-level quotas. When a request targets one of these models, RateLimitService first validates against the model-specific limit before falling back to the account default, as handled by the optional GeminiQuotaService integration.

Upstream Error Translation

When downstream providers return HTTP 429 or 403 rate-limit errors, the service extracts the Retry-After header value. The HandleUpstreamError method converts this into a platform-wide throttling response while preserving the original cooldown timing, ensuring users receive accurate backoff instructions rather than generic rejections.

Code Implementation Examples

Wiring Dependencies

The ProvideRateLimitService function in backend/internal/service/wire.go assembles the service with its required dependencies:

func ProvideRateLimitService(
    accountRepo AccountRepository,
    usageRepo UsageRepository,
    cfg *config.Config,
    geminiQuotaService *GeminiQuotaService,
    tempUnschedCache *TempUnschedulableCache,
) *RateLimitService {
    return NewRateLimitService(accountRepo, usageRepo, cfg, geminiQuotaService, tempUnschedCache)
}

Handler Integration

API handlers check quotas before forwarding requests to upstream providers:

func (h *Handler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
    // Extract user/account ID from request context
    if err := h.rateLimitSvc.CheckRequest(ctx, accountID, modelID, tokenCount); err != nil {
        http.Error(w, err.Error(), http.StatusTooManyRequests)
        return
    }

    // Forward to upstream API if limit check passes
    h.proxy.ServeHTTP(w, r)
}

Middleware Header Injection

The response middleware injects rate limiting headers after request processing:

func RateLimitHeaders(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        next.ServeHTTP(w, r)

        // Headers populated by RateLimitService into context
        w.Header().Set("x-ratelimit-limit-requests", ctx.Value("limitRequests").(string))
        w.Header().Set("x-ratelimit-remaining-requests", ctx.Value("remainingRequests").(string))
        w.Header().Set("x-ratelimit-reset-requests", ctx.Value("resetRequests").(string))
    })
}

Upstream Error Handling

Convert provider-specific rate limit errors into platform responses:

func (s *RateLimitService) HandleUpstreamError(err error, accountID string) error {
    if errors.Is(err, upstream.ErrRateLimited) {
        cooldown := parseRetryAfter(err) // Extract from Retry-After header
        s.tempUnschedCache.SetCooldown(accountID, cooldown)
        return fmt.Errorf("rate limit exceeded, retry after %s", cooldown)
    }
    return err
}

Summary

  • Sub2API rate limiting centers on the RateLimitService located in backend/internal/service/ratelimit_service.go.
  • The system uses AccountRepository and UsageRepository to enforce per-account RPM and TPM quotas with configurable multipliers.
  • Temporary cooldowns are managed through TempUnschedulableCache to protect scheduled job queues.
  • Standard X-RateLimit-* headers are defined in backend/internal/util/responseheaders/responseheaders.go and injected by middleware.
  • Model-specific limits (Gemini, Codex) receive specialized handling through dedicated quota services.
  • Upstream provider rate limits trigger automatic cooldown synchronization via HandleUpstreamError.

Frequently Asked Questions

How does Sub2API handle concurrent requests during rate limiting?

The RateLimitService evaluates each request atomically against the UsageRepository counters. If multiple simultaneous requests would collectively exceed the quota, the service rejects excess requests with HTTP 429 while allowing compliant requests through, ensuring precise enforcement without overages.

What happens when a user hits a rate limit on a specific model like Gemini?

When targeting model-specific endpoints, Sub2API first checks the dedicated GeminiQuotaService limits before evaluating account-wide quotas. If the model limit is breached, the request is rejected immediately with model-specific error messaging, preventing consumption of the user's general account quota.

Can rate limits be customized per user account?

Yes. The AccountRepository stores per-account overrides including custom rate_multiplier values. During the evaluation flow, RateLimitService applies these multipliers to the base configuration limits, enabling tiered access levels for different customer segments without code changes.

Where are the rate limit headers defined in the codebase?

The header constants are defined in backend/internal/util/responseheaders/responseheaders.go at line 26, which specifies keys like x-ratelimit-limit-requests and x-ratelimit-remaining-requests. The middleware references these constants to ensure consistent header naming across all API responses.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →