How to Implement Rate Limiting for MetaMCP Endpoints: User and Client Level Guide

MetaMCP uses a dual-layer rate limiting architecture that applies token-bucket limits per user and sliding-window limits per client, configurable per endpoint via middleware flags.

MetaMCP is an open-source platform for managing Model Context Protocol (MCP) endpoints. To prevent abuse and ensure fair resource allocation, the codebase implements a sophisticated rate limiting system in apps/backend/src/lib/rate-limit.ts that operates at both the user level (global per-namespace) and client level (per-IP or API key). This guide explains how to configure and extend these protections.

Understanding MetaMCP's Dual-Layer Rate Limiting Architecture

The system distinguishes between two distinct protection layers that can be toggled independently for each endpoint.

User-Level Rate Limiting (Token Bucket)

The user-level limit restricts the total requests a user can issue across all clients, identified by the user_id associated with the endpoint record.

In apps/backend/src/lib/rate-limit.ts, the RateLimiting class implements a token-bucket algorithm. It maintains a bucket per namespace_uuid, where each request consumes one token. The bucket refills at max_rate tokens every max_rate_seconds. If the bucket is empty, the onRequest method throws a RateLimitError, which the middleware converts into an HTTP 503 response.

Client-Level Rate Limiting (Sliding Window)

The client-level limit restricts requests from an individual client, identified by the client_max_rate_strategy (for example, an IP address header).

The SlidingWindowRateLimiting class in the same file implements a sliding-window algorithm. It maintains a timestamp list per client key, allowing only client_max_rate requests within the last client_max_rate_seconds. Exceeding this window throws a RateLimitError, translated to an HTTP 429 response by the middleware.

Configuring Rate Limits in MetaMCP Endpoints

Rate limiting is configured per endpoint through the Zod schema defined in packages/zod-types/src/endpoints.zod.ts. The following fields control the behavior:

enable_max_rate: boolean,
max_rate: number,
max_rate_seconds: number,
enable_client_max_rate: boolean,
client_max_rate: number,
client_max_rate_seconds: number,
client_max_rate_strategy: string,      // e.g., "ip"
client_max_rate_strategy_key: string,  // header name, e.g., "x-forwarded-for"

When creating or updating an endpoint via the REST API, include these fields to activate the respective limiters:

POST /api/v1/endpoints
Content-Type: application/json

{
  "name": "production-mcp",
  "namespaceUuid": "c3f2d1e0-8b9a-4c5d-9e8f-123456789abc",
  "enableApiKeyAuth": true,
  "enableMaxRate": true,
  "maxRate": 500,
  "maxRateSeconds": 60,
  "enableClientMaxRate": true,
  "clientMaxRate": 50,
  "clientMaxRateSeconds": 60,
  "clientMaxRateStrategy": "ip",
  "clientMaxRateStrategyKey": "x-forwarded-for"
}

How the Rate Limiting Middleware Works

The rateLimitMiddleware in apps/backend/src/middleware/rate-limit.middleware.ts orchestrates the limiters based on the endpoint configuration attached to the request.

The middleware reads the DatabaseEndpoint object (attached by lookupEndpoint middleware) and executes the appropriate logic:

  • Both flags enabled: Runs both RateLimiting (token bucket) and SlidingWindowRateLimiting sequentially.
  • Only enable_client_max_rate: Runs only the sliding-window limiter.
  • Only enable_max_rate: Runs only the token-bucket limiter.
  • Neither enabled: Passes through without rate limiting.

This design allows granular control per endpoint without modifying the route handlers. The middleware is applied to public MCP routes in apps/backend/src/routers/public-metamcp/streamable-http.ts and sse.ts:

import { rateLimitMiddleware } from "@/middleware/rate-limit.middleware";

router.get(
  "/:endpoint_name/mcp",
  lookupEndpoint,
  authenticateApiKey,
  rateLimitMiddleware,
  async (req, res) => { /* MCP transport logic */ }
);

Implementing Custom Rate Limiting Strategies

The modular design in apps/backend/src/lib/rate-limit.ts allows you to extend or replace the default algorithms without touching the middleware logic.

Customizing Token Bucket Refill Rates

To implement a custom refill algorithm per namespace, extend the RateLimiting class:

export class CustomRateLimiting extends RateLimiting {
  constructor(private refillRate: number) {
    super();
  }

  async onRequest(context: Context, callNext: CallNext) {
    const { namespace_uuid } = context.req.endpoint;
    let limiter = this.limiters.get(namespace_uuid);
    
    if (!limiter) {
      // Use custom refillRate instead of endpoint.max_rate_seconds
      limiter = new TokenBucketRateLimiter(this.maxRate, this.refillRate);
      this.limiters.set(namespace_uuid, limiter);
    }
    
    // Continue with parent implementation
    return super.onRequest(context, callNext);
  }
}

Swap the instance in rate-limit.middleware.ts to apply the new logic globally:

// Replace: const rateLimiter = new RateLimiting();
const rateLimiter = new CustomRateLimiting(0.5);

Switching to Redis for Distributed Rate Limiting

For production deployments across multiple server instances, replace the in-memory Map storage with Redis. The limiter state lives in this.limiters (a Map), which you can replace with a Redis client implementing the same consume interface:

import { createClient } from "redis";

class RedisTokenBucket {
  private client = createClient({ url: process.env.REDIS_URL });

  async consume(key: string, tokens = 1): Promise<boolean> {
    // Lua script for atomic check-and-update
    const script = `
      local tokens = tonumber(ARGV[1])
      local capacity = tonumber(ARGV[2])
      local refill = tonumber(ARGV[3])
      local now = tonumber(ARGV[4])

      local bucket = redis.call("HMGET", KEYS[1], "tokens", "last")
      local curTokens = tonumber(bucket[1]) or capacity
      local last = tonumber(bucket[2]) or now

      local elapsed = now - last
      curTokens = math.min(capacity, curTokens + elapsed * refill)
      
      if curTokens < tokens then
        return 0
      end
      
      curTokens = curTokens - tokens
      redis.call("HMSET", KEYS[1], "tokens", curTokens, "last", now)
      redis.call("EXPIRE", KEYS[1], math.ceil(capacity / refill))
      return 1
    `;
    
    const now = Date.now() / 1000;
    const result = await this.client.eval(
      script,
      { keys: [`bucket:${key}`], arguments: [tokens, this.capacity, this.refillRate, now] }
    );
    
    return Boolean(result);
  }
}

Plug this class into RateLimiting by replacing the TokenBucketRateLimiter instantiation. The middleware and error handling remain unchanged.

Summary

  • MetaMCP implements dual-layer rate limiting through apps/backend/src/lib/rate-limit.ts, combining token-bucket (user-level) and sliding-window (client-level) algorithms.
  • Configuration is endpoint-specific via boolean flags (enable_max_rate, enable_client_max_rate) and numeric parameters defined in packages/zod-types/src/endpoints.zod.ts.
  • Middleware orchestration in apps/backend/src/middleware/rate-limit.middleware.ts automatically selects the appropriate limiter(s) based on endpoint settings, returning 503 for user limits and 429 for client limits.
  • Extensible architecture allows customization of refill algorithms or replacement of in-memory storage with Redis for distributed deployments without modifying the core middleware logic.

Frequently Asked Questions

What is the difference between user-level and client-level rate limiting in MetaMCP?

User-level rate limiting applies to all requests from a specific user across all clients, using a token-bucket algorithm keyed by namespace_uuid. It returns HTTP 503 when exceeded. Client-level rate limiting applies to individual clients (identified by IP or custom header) using a sliding-window algorithm, returning HTTP 429 when exceeded. Both can be enabled independently per endpoint.

How do I enable rate limiting for a specific MetaMCP endpoint?

Set the boolean flags enableMaxRate and/or enableClientMaxRate to true when creating or updating the endpoint via the API. Include the corresponding numeric parameters (maxRate, maxRateSeconds, clientMaxRate, clientMaxRateSeconds) and strategy fields (clientMaxRateStrategy, clientMaxRateStrategyKey). The rateLimitMiddleware automatically applies the configured limits to all requests hitting that endpoint.

Can I use Redis instead of in-memory storage for rate limiting?

Yes. The RateLimiting and SlidingWindowRateLimiting classes use a Map for state storage by default. You can replace this with a Redis-backed implementation that exposes the same consume or tracking interface. The middleware in rateLimitMiddleware will work unchanged because it only depends on the limiter's public API, not the underlying storage mechanism.

What HTTP status codes does MetaMCP return when rate limits are exceeded?

MetaMCP returns 503 Service Unavailable when the user-level token-bucket limit is exceeded, indicating the server is temporarily overloaded for that namespace. It returns 429 Too Many Requests when the client-level sliding-window limit is exceeded, indicating the specific client has sent too many requests. These distinctions help clients implement appropriate backoff strategies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →