# How to Implement Rate Limiting for MetaMCP Endpoints: User and Client Level Guide

> Master MetaMCP rate limiting with this guide. Learn to implement user and client level token-bucket and sliding-window limits for your endpoints. Boost performance and security.

- Repository: [metatool-ai/metamcp](https://github.com/metatool-ai/metamcp)
- Tags: how-to-guide
- Published: 2026-03-07

---

**MetaMCP uses a dual-layer rate limiting architecture that applies token-bucket limits per user and sliding-window limits per client, configurable per endpoint via middleware flags.**

MetaMCP is an open-source platform for managing Model Context Protocol (MCP) endpoints. To prevent abuse and ensure fair resource allocation, the codebase implements a sophisticated rate limiting system in [`apps/backend/src/lib/rate-limit.ts`](https://github.com/metatool-ai/metamcp/blob/main/apps/backend/src/lib/rate-limit.ts) that operates at both the user level (global per-namespace) and client level (per-IP or API key). This guide explains how to configure and extend these protections.

## Understanding MetaMCP's Dual-Layer Rate Limiting Architecture

The system distinguishes between two distinct protection layers that can be toggled independently for each endpoint.

### User-Level Rate Limiting (Token Bucket)

The **user-level limit** restricts the total requests a user can issue across all clients, identified by the `user_id` associated with the endpoint record.

In [`apps/backend/src/lib/rate-limit.ts`](https://github.com/metatool-ai/metamcp/blob/main/apps/backend/src/lib/rate-limit.ts), the `RateLimiting` class implements a **token-bucket** algorithm. It maintains a bucket per `namespace_uuid`, where each request consumes one token. The bucket refills at `max_rate` tokens every `max_rate_seconds`. If the bucket is empty, the `onRequest` method throws a `RateLimitError`, which the middleware converts into an HTTP **503** response.

### Client-Level Rate Limiting (Sliding Window)

The **client-level limit** restricts requests from an individual client, identified by the `client_max_rate_strategy` (for example, an IP address header).

The `SlidingWindowRateLimiting` class in the same file implements a **sliding-window** algorithm. It maintains a timestamp list per client key, allowing only `client_max_rate` requests within the last `client_max_rate_seconds`. Exceeding this window throws a `RateLimitError`, translated to an HTTP **429** response by the middleware.

## Configuring Rate Limits in MetaMCP Endpoints

Rate limiting is configured per endpoint through the Zod schema defined in [`packages/zod-types/src/endpoints.zod.ts`](https://github.com/metatool-ai/metamcp/blob/main/packages/zod-types/src/endpoints.zod.ts). The following fields control the behavior:

```typescript
enable_max_rate: boolean,
max_rate: number,
max_rate_seconds: number,
enable_client_max_rate: boolean,
client_max_rate: number,
client_max_rate_seconds: number,
client_max_rate_strategy: string,      // e.g., "ip"
client_max_rate_strategy_key: string,  // header name, e.g., "x-forwarded-for"

```

When creating or updating an endpoint via the REST API, include these fields to activate the respective limiters:

```bash
POST /api/v1/endpoints
Content-Type: application/json

{
  "name": "production-mcp",
  "namespaceUuid": "c3f2d1e0-8b9a-4c5d-9e8f-123456789abc",
  "enableApiKeyAuth": true,
  "enableMaxRate": true,
  "maxRate": 500,
  "maxRateSeconds": 60,
  "enableClientMaxRate": true,
  "clientMaxRate": 50,
  "clientMaxRateSeconds": 60,
  "clientMaxRateStrategy": "ip",
  "clientMaxRateStrategyKey": "x-forwarded-for"
}

```

## How the Rate Limiting Middleware Works

The `rateLimitMiddleware` in [`apps/backend/src/middleware/rate-limit.middleware.ts`](https://github.com/metatool-ai/metamcp/blob/main/apps/backend/src/middleware/rate-limit.middleware.ts) orchestrates the limiters based on the endpoint configuration attached to the request.

The middleware reads the `DatabaseEndpoint` object (attached by `lookupEndpoint` middleware) and executes the appropriate logic:

- **Both flags enabled**: Runs both `RateLimiting` (token bucket) and `SlidingWindowRateLimiting` sequentially.
- **Only `enable_client_max_rate`**: Runs only the sliding-window limiter.
- **Only `enable_max_rate`**: Runs only the token-bucket limiter.
- **Neither enabled**: Passes through without rate limiting.

This design allows granular control per endpoint without modifying the route handlers. The middleware is applied to public MCP routes in [`apps/backend/src/routers/public-metamcp/streamable-http.ts`](https://github.com/metatool-ai/metamcp/blob/main/apps/backend/src/routers/public-metamcp/streamable-http.ts) and [`sse.ts`](https://github.com/metatool-ai/metamcp/blob/main/sse.ts):

```typescript
import { rateLimitMiddleware } from "@/middleware/rate-limit.middleware";

router.get(
  "/:endpoint_name/mcp",
  lookupEndpoint,
  authenticateApiKey,
  rateLimitMiddleware,
  async (req, res) => { /* MCP transport logic */ }
);

```

## Implementing Custom Rate Limiting Strategies

The modular design in [`apps/backend/src/lib/rate-limit.ts`](https://github.com/metatool-ai/metamcp/blob/main/apps/backend/src/lib/rate-limit.ts) allows you to extend or replace the default algorithms without touching the middleware logic.

### Customizing Token Bucket Refill Rates

To implement a custom refill algorithm per namespace, extend the `RateLimiting` class:

```typescript
export class CustomRateLimiting extends RateLimiting {
  constructor(private refillRate: number) {
    super();
  }

  async onRequest(context: Context, callNext: CallNext) {
    const { namespace_uuid } = context.req.endpoint;
    let limiter = this.limiters.get(namespace_uuid);
    
    if (!limiter) {
      // Use custom refillRate instead of endpoint.max_rate_seconds
      limiter = new TokenBucketRateLimiter(this.maxRate, this.refillRate);
      this.limiters.set(namespace_uuid, limiter);
    }
    
    // Continue with parent implementation
    return super.onRequest(context, callNext);
  }
}

```

Swap the instance in [`rate-limit.middleware.ts`](https://github.com/metatool-ai/metamcp/blob/main/rate-limit.middleware.ts) to apply the new logic globally:

```typescript
// Replace: const rateLimiter = new RateLimiting();
const rateLimiter = new CustomRateLimiting(0.5);

```

### Switching to Redis for Distributed Rate Limiting

For production deployments across multiple server instances, replace the in-memory `Map` storage with Redis. The limiter state lives in `this.limiters` (a `Map`), which you can replace with a Redis client implementing the same `consume` interface:

```typescript
import { createClient } from "redis";

class RedisTokenBucket {
  private client = createClient({ url: process.env.REDIS_URL });

  async consume(key: string, tokens = 1): Promise<boolean> {
    // Lua script for atomic check-and-update
    const script = `
      local tokens = tonumber(ARGV[1])
      local capacity = tonumber(ARGV[2])
      local refill = tonumber(ARGV[3])
      local now = tonumber(ARGV[4])

      local bucket = redis.call("HMGET", KEYS[1], "tokens", "last")
      local curTokens = tonumber(bucket[1]) or capacity
      local last = tonumber(bucket[2]) or now

      local elapsed = now - last
      curTokens = math.min(capacity, curTokens + elapsed * refill)
      
      if curTokens < tokens then
        return 0
      end
      
      curTokens = curTokens - tokens
      redis.call("HMSET", KEYS[1], "tokens", curTokens, "last", now)
      redis.call("EXPIRE", KEYS[1], math.ceil(capacity / refill))
      return 1
    `;
    
    const now = Date.now() / 1000;
    const result = await this.client.eval(
      script,
      { keys: [`bucket:${key}`], arguments: [tokens, this.capacity, this.refillRate, now] }
    );
    
    return Boolean(result);
  }
}

```

Plug this class into `RateLimiting` by replacing the `TokenBucketRateLimiter` instantiation. The middleware and error handling remain unchanged.

## Summary

- **MetaMCP implements dual-layer rate limiting** through [`apps/backend/src/lib/rate-limit.ts`](https://github.com/metatool-ai/metamcp/blob/main/apps/backend/src/lib/rate-limit.ts), combining token-bucket (user-level) and sliding-window (client-level) algorithms.
- **Configuration is endpoint-specific** via boolean flags (`enable_max_rate`, `enable_client_max_rate`) and numeric parameters defined in [`packages/zod-types/src/endpoints.zod.ts`](https://github.com/metatool-ai/metamcp/blob/main/packages/zod-types/src/endpoints.zod.ts).
- **Middleware orchestration** in [`apps/backend/src/middleware/rate-limit.middleware.ts`](https://github.com/metatool-ai/metamcp/blob/main/apps/backend/src/middleware/rate-limit.middleware.ts) automatically selects the appropriate limiter(s) based on endpoint settings, returning **503** for user limits and **429** for client limits.
- **Extensible architecture** allows customization of refill algorithms or replacement of in-memory storage with Redis for distributed deployments without modifying the core middleware logic.

## Frequently Asked Questions

### What is the difference between user-level and client-level rate limiting in MetaMCP?

**User-level rate limiting** applies to all requests from a specific user across all clients, using a token-bucket algorithm keyed by `namespace_uuid`. It returns HTTP **503** when exceeded. **Client-level rate limiting** applies to individual clients (identified by IP or custom header) using a sliding-window algorithm, returning HTTP **429** when exceeded. Both can be enabled independently per endpoint.

### How do I enable rate limiting for a specific MetaMCP endpoint?

Set the boolean flags `enableMaxRate` and/or `enableClientMaxRate` to `true` when creating or updating the endpoint via the API. Include the corresponding numeric parameters (`maxRate`, `maxRateSeconds`, `clientMaxRate`, `clientMaxRateSeconds`) and strategy fields (`clientMaxRateStrategy`, `clientMaxRateStrategyKey`). The `rateLimitMiddleware` automatically applies the configured limits to all requests hitting that endpoint.

### Can I use Redis instead of in-memory storage for rate limiting?

Yes. The `RateLimiting` and `SlidingWindowRateLimiting` classes use a `Map` for state storage by default. You can replace this with a Redis-backed implementation that exposes the same `consume` or tracking interface. The middleware in `rateLimitMiddleware` will work unchanged because it only depends on the limiter's public API, not the underlying storage mechanism.

### What HTTP status codes does MetaMCP return when rate limits are exceeded?

MetaMCP returns **503 Service Unavailable** when the user-level token-bucket limit is exceeded, indicating the server is temporarily overloaded for that namespace. It returns **429 Too Many Requests** when the client-level sliding-window limit is exceeded, indicating the specific client has sent too many requests. These distinctions help clients implement appropriate backoff strategies.