Inference Proxy Rate Limiting and Routing in Thunderbolt: Responsibilities and Implementation
The inference proxy in Thunderbolt handles routing by mapping model names to specific AI providers and enforces per-user rate limits of 20 requests per minute, returning 429 errors with Retry-After headers when limits are exceeded.
The inference proxy serves as the backend layer for the Thunderbird Thunderbolt project, managing all /v1/chat/completions requests. It determines which external AI provider should process each request while ensuring fair usage through strict rate limiting policies.
Routing Responsibilities
The routing system translates client-provided model identifiers into concrete provider connections.
Model-to-Provider Mapping
When a request arrives at POST /v1/chat/completions, the handler in backend/src/inference/routes.ts extracts the model field from the request body. This value is looked up against the supportedModels map to determine the target provider.
The system supports multiple providers including thunderbolt, mistral, and anthropic. For example, a model value of gpt-oss-120b resolves to the thunderbolt provider.
Provider Selection and Logging
After resolving the provider, the system logs the routing decision via console.info, showing exactly which provider receives the request. The handler then obtains the concrete client instance through getInferenceClient(provider) and forwards the request to the provider's SDK using client.chat.completions.create.
Rate Limiting Implementation
Rate limiting occurs before routing, protecting downstream providers from overload and ensuring equitable access.
Per-User Request Quotas
The middleware defined in backend/src/middleware/rate-limit.ts implements user-based limiting through the createInferenceRateLimit function. By default, each user is limited to 20 requests per minute.
The middleware runs during onBeforeHandle, consuming a point from the user's quota identified by user:<id>. If the quota is exhausted, the middleware immediately aborts with a 429 Too Many Requests status and includes a Retry-After header indicating when the client should retry.
Rate Limit Headers and Response Integration
When a request passes the rate limit check, the middleware injects standard RateLimit-* headers into ctx.set.headers. These headers propagate through the request lifecycle.
In backend/src/inference/routes.ts, before building the SSE streaming response, the handler explicitly copies all RateLimit-* headers from the context into the final response header map. This ensures clients receive real-time quota information even during streaming completions.
Request Processing Flow
The complete flow through the inference proxy follows these steps:
- Authentication: The auth macro ensures
ctx.user.idis available for rate limiting. - Rate Limit Check:
createInferenceRateLimitmiddleware verifies the user's quota and sets headers. - Model Resolution: The handler extracts the model name and looks up the provider in
supportedModels. - Provider Routing: The system logs the routing decision and retrieves the appropriate client.
- Request Forwarding: The request is sent to the provider's SDK with
client.chat.completions.create. - Response Streaming: The provider response is converted to an SSE stream with rate limit headers preserved.
- Error Handling: Provider errors are caught, logged, and re-thrown as generic messages to prevent internal leakage.
Practical Implementation Example
// Example client implementation consuming the inference proxy
import { fetch } from 'cross-fetch';
async function streamChatCompletion() {
const response = await fetch('https://api.thunderbolt.dev/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
// Authentication handled via BetterAuth cookie/token
},
body: JSON.stringify({
model: 'gpt-oss-120b',
stream: true,
temperature: 0.7,
messages: [{ role: 'user', content: 'Explain rate limiting' }],
}),
});
// Inspect rate limit headers injected by the proxy
console.log('Rate limit remaining:', response.headers.get('RateLimit-Remaining'));
// Consume the SSE stream
const reader = response.body?.getReader();
// Process chunks as they arrive...
}
This request routes to the Thunderbolt provider based on the model mapping, undergoes the 20 requests per minute quota check, and receives streaming headers indicating the remaining quota.
Key Implementation Files
| File | Purpose |
|---|---|
backend/src/inference/routes.ts |
Defines the /chat/completions endpoint, implements model-to-provider mapping via supportedModels, logs routing decisions, and merges rate limit headers into SSE responses. |
backend/src/middleware/rate-limit.ts |
Implements createInferenceRateLimit for per-user rate limiting (20 req/min), sets RateLimit-* headers, and returns 429 with Retry-After on quota exhaustion. |
backend/src/index.ts |
Registers inference routes with the rate limiting middleware in the application stack. |
Summary
- Routing: The inference proxy maps incoming
modelparameters to specific AI providers (thunderbolt,mistral,anthropic) using thesupportedModelslookup inbackend/src/inference/routes.ts. - Rate Limiting: A per-user limit of 20 requests per minute is enforced by
createInferenceRateLimitmiddleware before routing occurs, returning 429 errors withRetry-Afterheaders when exceeded. - Header Propagation: Rate limit headers are injected into the SSE streaming response so clients can monitor quota usage in real-time.
- Error Isolation: Provider-specific errors are caught and re-thrown as generic messages to prevent internal implementation details from leaking to clients.
Frequently Asked Questions
How does the inference proxy determine which provider handles a request?
The proxy extracts the model field from the request body and performs a lookup against the supportedModels map defined in backend/src/inference/routes.ts. This map associates model identifiers like gpt-oss-120b with specific provider strings such as thunderbolt, mistral, or anthropic, and the proxy instantiates the corresponding client using getInferenceClient(provider).
What happens when a user exceeds the rate limit?
When the per-user quota of 20 requests per minute is exhausted, the createInferenceRateLimit middleware in backend/src/middleware/rate-limit.ts immediately aborts the request with HTTP status 429 Too Many Requests. The response includes a Retry-After header indicating the number of seconds until the quota resets, and the client receives no provider routing or inference processing.
Are rate limit headers available during streaming responses?
Yes. The inference proxy preserves rate limit metadata throughout the request lifecycle. The middleware injects RateLimit-* headers into the context, and before the SSE stream is returned in backend/src/inference/routes.ts, the handler explicitly copies these headers into the final response. This allows streaming clients to monitor remaining quota and reset times in real-time.
Where is the rate limiting middleware registered in the application?
The createInferenceRateLimit middleware is registered in backend/src/index.ts at lines 91-93, where it is applied specifically to the inference routes group. This ensures that rate limiting executes after authentication but before the route handler performs model lookup and provider routing, protecting downstream services from excessive load.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →