How Sub2API Performs Rate Limiting: Architecture and Implementation
Sub2API enforces rate limiting through a centralized RateLimitService that evaluates per-account quotas against usage repositories, applies configurable multipliers, and communicates limits via standard X-RateLimit-* HTTP headers.
The open-source Sub2API project (available at Wei-Shaw/sub2api) implements a multi-layered rate limiting strategy to protect downstream AI providers like OpenAI and Gemini from abuse. The system coordinates between persistent storage, temporary caching, and HTTP middleware to enforce requests-per-minute (RPM) and tokens-per-minute (TPM) constraints dynamically according to the source code.
Core Architecture and Components
The rate limiting system relies on a pipeline of specialized components wired together through dependency injection.
RateLimitService
The RateLimitService in backend/internal/service/ratelimit_service.go serves as the central authority for quota enforcement. It aggregates configuration data, account-level overrides, and real-time usage statistics to determine whether to allow, delay, or reject incoming requests according to the repository structure.
Supporting Repositories and Caches
AccountRepository: Persists per-account limits and custom multipliers (e.g., elevated quotas for specific user tiers).UsageRepository: Tracks real-time request and token counters within the current time window.TempUnschedulableCache: Stores temporary cooldown flags when accounts breach limits, preventing scheduled background jobs from executing during backoff periods.GeminiQuotaService: An optional adapter that handles Google Gemini-specific quota constraints separately from generic account limits.
Response Header Middleware
The ResponseHeaders utility in backend/internal/util/responseheaders/responseheaders.go (line 26) defines the standard X-RateLimit-* header names, including x-ratelimit-limit-requests, x-ratelimit-remaining-requests, and x-ratelimit-reset-requests, which the middleware attaches to every HTTP response.
Rate Limit Evaluation Flow
For each incoming request, RateLimitService executes the following decision sequence:
- Lookup Limits: Retrieve the base RPM and TPM values from the global configuration and any account-specific overrides stored in
AccountRepository. - Read Current Usage: Query
UsageRepositoryfor the number of requests and tokens already consumed within the current sliding window. - Apply Multipliers: Scale limits according to the account's
rate_multiplierfield to support tiered pricing or promotional quotas. - Check Thresholds:
- If the request would exceed the limit, reject it with HTTP 429.
- If the request approaches the limit, set a temporary unschedulable flag in
TempUnschedulableCacheto pause background job scheduling for a calculated backoff period.
- Send Response Headers: Populate
X-RateLimit-*headers so clients can programmatically adjust their request rate.
HTTP Headers and Client Communication
Sub2API follows industry standards by exposing quota state through response headers. As implemented in backend/internal/util/responseheaders/responseheaders.go, the platform returns:
x-ratelimit-limit-requests: The maximum number of requests allowed in the current window.x-ratelimit-remaining-requests: The number of requests remaining before hitting the limit.x-ratelimit-reset-requests: Unix timestamp indicating when the request quota resets.- Token equivalents: Corresponding headers for TPM (tokens per minute) limits.
This allows API consumers to implement client-side throttling without hard-coding limit values.
Handling Model-Specific and Upstream Limits
Model-Specific Constraints
Certain AI models (e.g., Codex and Gemini variants) carry distinct rate limits separate from account-level quotas. When a request targets one of these models, RateLimitService first validates against the model-specific limit before falling back to the account default, as handled by the optional GeminiQuotaService integration.
Upstream Error Translation
When downstream providers return HTTP 429 or 403 rate-limit errors, the service extracts the Retry-After header value. The HandleUpstreamError method converts this into a platform-wide throttling response while preserving the original cooldown timing, ensuring users receive accurate backoff instructions rather than generic rejections.
Code Implementation Examples
Wiring Dependencies
The ProvideRateLimitService function in backend/internal/service/wire.go assembles the service with its required dependencies:
func ProvideRateLimitService(
accountRepo AccountRepository,
usageRepo UsageRepository,
cfg *config.Config,
geminiQuotaService *GeminiQuotaService,
tempUnschedCache *TempUnschedulableCache,
) *RateLimitService {
return NewRateLimitService(accountRepo, usageRepo, cfg, geminiQuotaService, tempUnschedCache)
}
Handler Integration
API handlers check quotas before forwarding requests to upstream providers:
func (h *Handler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
// Extract user/account ID from request context
if err := h.rateLimitSvc.CheckRequest(ctx, accountID, modelID, tokenCount); err != nil {
http.Error(w, err.Error(), http.StatusTooManyRequests)
return
}
// Forward to upstream API if limit check passes
h.proxy.ServeHTTP(w, r)
}
Middleware Header Injection
The response middleware injects rate limiting headers after request processing:
func RateLimitHeaders(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
next.ServeHTTP(w, r)
// Headers populated by RateLimitService into context
w.Header().Set("x-ratelimit-limit-requests", ctx.Value("limitRequests").(string))
w.Header().Set("x-ratelimit-remaining-requests", ctx.Value("remainingRequests").(string))
w.Header().Set("x-ratelimit-reset-requests", ctx.Value("resetRequests").(string))
})
}
Upstream Error Handling
Convert provider-specific rate limit errors into platform responses:
func (s *RateLimitService) HandleUpstreamError(err error, accountID string) error {
if errors.Is(err, upstream.ErrRateLimited) {
cooldown := parseRetryAfter(err) // Extract from Retry-After header
s.tempUnschedCache.SetCooldown(accountID, cooldown)
return fmt.Errorf("rate limit exceeded, retry after %s", cooldown)
}
return err
}
Summary
- Sub2API rate limiting centers on the
RateLimitServicelocated inbackend/internal/service/ratelimit_service.go. - The system uses
AccountRepositoryandUsageRepositoryto enforce per-account RPM and TPM quotas with configurable multipliers. - Temporary cooldowns are managed through
TempUnschedulableCacheto protect scheduled job queues. - Standard
X-RateLimit-*headers are defined inbackend/internal/util/responseheaders/responseheaders.goand injected by middleware. - Model-specific limits (Gemini, Codex) receive specialized handling through dedicated quota services.
- Upstream provider rate limits trigger automatic cooldown synchronization via
HandleUpstreamError.
Frequently Asked Questions
How does Sub2API handle concurrent requests during rate limiting?
The RateLimitService evaluates each request atomically against the UsageRepository counters. If multiple simultaneous requests would collectively exceed the quota, the service rejects excess requests with HTTP 429 while allowing compliant requests through, ensuring precise enforcement without overages.
What happens when a user hits a rate limit on a specific model like Gemini?
When targeting model-specific endpoints, Sub2API first checks the dedicated GeminiQuotaService limits before evaluating account-wide quotas. If the model limit is breached, the request is rejected immediately with model-specific error messaging, preventing consumption of the user's general account quota.
Can rate limits be customized per user account?
Yes. The AccountRepository stores per-account overrides including custom rate_multiplier values. During the evaluation flow, RateLimitService applies these multipliers to the base configuration limits, enabling tiered access levels for different customer segments without code changes.
Where are the rate limit headers defined in the codebase?
The header constants are defined in backend/internal/util/responseheaders/responseheaders.go at line 26, which specifies keys like x-ratelimit-limit-requests and x-ratelimit-remaining-requests. The middleware references these constants to ensure consistent header naming across all API responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →