# Inference Proxy Rate Limiting and Routing in Thunderbolt: Responsibilities and Implementation

> Learn how Thunderbolt's inference proxy manages rate limiting and routing. Discover its role in mapping models to AI providers and enforcing per-user request limits for optimal performance.

- Repository: [Thunderbird/thunderbolt](https://github.com/thunderbird/thunderbolt)
- Tags: deep-dive
- Published: 2026-04-19

---

**The inference proxy in Thunderbolt handles routing by mapping model names to specific AI providers and enforces per-user rate limits of 20 requests per minute, returning 429 errors with Retry-After headers when limits are exceeded.**

The inference proxy serves as the backend layer for the Thunderbird Thunderbolt project, managing all `/v1/chat/completions` requests. It determines which external AI provider should process each request while ensuring fair usage through strict rate limiting policies.

## Routing Responsibilities

The routing system translates client-provided model identifiers into concrete provider connections.

### Model-to-Provider Mapping

When a request arrives at `POST /v1/chat/completions`, the handler in [`backend/src/inference/routes.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/inference/routes.ts) extracts the `model` field from the request body. This value is looked up against the `supportedModels` map to determine the target provider.

The system supports multiple providers including `thunderbolt`, `mistral`, and `anthropic`. For example, a model value of `gpt-oss-120b` resolves to the `thunderbolt` provider.

### Provider Selection and Logging

After resolving the provider, the system logs the routing decision via `console.info`, showing exactly which provider receives the request. The handler then obtains the concrete client instance through `getInferenceClient(provider)` and forwards the request to the provider's SDK using `client.chat.completions.create`.

## Rate Limiting Implementation

Rate limiting occurs **before** routing, protecting downstream providers from overload and ensuring equitable access.

### Per-User Request Quotas

The middleware defined in [`backend/src/middleware/rate-limit.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/middleware/rate-limit.ts) implements user-based limiting through the `createInferenceRateLimit` function. By default, each user is limited to **20 requests per minute**.

The middleware runs during `onBeforeHandle`, consuming a point from the user's quota identified by `user:<id>`. If the quota is exhausted, the middleware immediately aborts with a **429 Too Many Requests** status and includes a `Retry-After` header indicating when the client should retry.

### Rate Limit Headers and Response Integration

When a request passes the rate limit check, the middleware injects standard `RateLimit-*` headers into `ctx.set.headers`. These headers propagate through the request lifecycle.

In [`backend/src/inference/routes.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/inference/routes.ts), before building the SSE streaming response, the handler explicitly copies all `RateLimit-*` headers from the context into the final response header map. This ensures clients receive real-time quota information even during streaming completions.

## Request Processing Flow

The complete flow through the inference proxy follows these steps:

1. **Authentication**: The auth macro ensures `ctx.user.id` is available for rate limiting.
2. **Rate Limit Check**: `createInferenceRateLimit` middleware verifies the user's quota and sets headers.
3. **Model Resolution**: The handler extracts the model name and looks up the provider in `supportedModels`.
4. **Provider Routing**: The system logs the routing decision and retrieves the appropriate client.
5. **Request Forwarding**: The request is sent to the provider's SDK with `client.chat.completions.create`.
6. **Response Streaming**: The provider response is converted to an SSE stream with rate limit headers preserved.
7. **Error Handling**: Provider errors are caught, logged, and re-thrown as generic messages to prevent internal leakage.

## Practical Implementation Example

```typescript
// Example client implementation consuming the inference proxy
import { fetch } from 'cross-fetch';

async function streamChatCompletion() {
  const response = await fetch('https://api.thunderbolt.dev/v1/chat/completions', {
    method: 'POST',
    headers: {
      'Content-Type': 'application/json',
      // Authentication handled via BetterAuth cookie/token
    },
    body: JSON.stringify({
      model: 'gpt-oss-120b',
      stream: true,
      temperature: 0.7,
      messages: [{ role: 'user', content: 'Explain rate limiting' }],
    }),
  });

  // Inspect rate limit headers injected by the proxy
  console.log('Rate limit remaining:', response.headers.get('RateLimit-Remaining'));

  // Consume the SSE stream
  const reader = response.body?.getReader();
  // Process chunks as they arrive...
}

```

This request routes to the **Thunderbolt** provider based on the model mapping, undergoes the 20 requests per minute quota check, and receives streaming headers indicating the remaining quota.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`backend/src/inference/routes.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/inference/routes.ts) | Defines the `/chat/completions` endpoint, implements model-to-provider mapping via `supportedModels`, logs routing decisions, and merges rate limit headers into SSE responses. |
| [`backend/src/middleware/rate-limit.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/middleware/rate-limit.ts) | Implements `createInferenceRateLimit` for per-user rate limiting (20 req/min), sets `RateLimit-*` headers, and returns 429 with `Retry-After` on quota exhaustion. |
| [`backend/src/index.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/index.ts) | Registers inference routes with the rate limiting middleware in the application stack. |

## Summary

- **Routing**: The inference proxy maps incoming `model` parameters to specific AI providers (`thunderbolt`, `mistral`, `anthropic`) using the `supportedModels` lookup in [`backend/src/inference/routes.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/inference/routes.ts).
- **Rate Limiting**: A per-user limit of 20 requests per minute is enforced by `createInferenceRateLimit` middleware before routing occurs, returning 429 errors with `Retry-After` headers when exceeded.
- **Header Propagation**: Rate limit headers are injected into the SSE streaming response so clients can monitor quota usage in real-time.
- **Error Isolation**: Provider-specific errors are caught and re-thrown as generic messages to prevent internal implementation details from leaking to clients.

## Frequently Asked Questions

### How does the inference proxy determine which provider handles a request?

The proxy extracts the `model` field from the request body and performs a lookup against the `supportedModels` map defined in [`backend/src/inference/routes.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/inference/routes.ts). This map associates model identifiers like `gpt-oss-120b` with specific provider strings such as `thunderbolt`, `mistral`, or `anthropic`, and the proxy instantiates the corresponding client using `getInferenceClient(provider)`.

### What happens when a user exceeds the rate limit?

When the per-user quota of 20 requests per minute is exhausted, the `createInferenceRateLimit` middleware in [`backend/src/middleware/rate-limit.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/middleware/rate-limit.ts) immediately aborts the request with HTTP status **429 Too Many Requests**. The response includes a `Retry-After` header indicating the number of seconds until the quota resets, and the client receives no provider routing or inference processing.

### Are rate limit headers available during streaming responses?

Yes. The inference proxy preserves rate limit metadata throughout the request lifecycle. The middleware injects `RateLimit-*` headers into the context, and before the SSE stream is returned in [`backend/src/inference/routes.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/inference/routes.ts), the handler explicitly copies these headers into the final response. This allows streaming clients to monitor remaining quota and reset times in real-time.

### Where is the rate limiting middleware registered in the application?

The `createInferenceRateLimit` middleware is registered in [`backend/src/index.ts`](https://github.com/thunderbird/thunderbolt/blob/main/backend/src/index.ts) at lines 91-93, where it is applied specifically to the inference routes group. This ensures that rate limiting executes after authentication but before the route handler performs model lookup and provider routing, protecting downstream services from excessive load.