How to Configure Rate Limiting for API Endpoints in AutoGPT

AutoGPT protects its platform API using a per-API-key rate limiter implemented as FastAPI middleware that stores rolling-window request timestamps in Redis and enforces a configurable requests-per-minute quota.

The Significant-Gravitas/AutoGPT repository includes a production-ready rate limiting system for its platform backend. You can configure rate limiting for API endpoints in AutoGPT by adjusting environment variables that control Redis connection details and per-minute request quotas, or by modifying the underlying Python classes directly.

Architecture Overview

The rate limiting system consists of four core components that work together to enforce quotas and provide observability.

RateLimitSettings Configuration

The RateLimitSettings class defined in autogpt_platform/autogpt_libs/autogpt_libs/rate_limit/config.py (lines 7-28) is a Pydantic settings object that holds all configurable values. It exposes the following environment variable overrides:

  • REDIS_HOST – Redis server hostname
  • REDIS_PORT – Redis server port
  • REDIS_PASSWORD – Redis authentication password
  • RATE_LIMIT_REQUESTS_PER_MINUTE – Maximum requests allowed per API key per minute (default: 60)

RateLimiter Core Logic

The RateLimiter class in autogpt_platform/autogpt_libs/autogpt_libs/rate_limit/limiter.py (lines 9-26) connects to Redis and maintains a sorted set (ZSET) per API key. It implements a sliding window algorithm:

  1. Removes timestamps older than the 60-second window from the Redis key ratelimit:<key>:1min
  2. Adds the current request timestamp to the ZSET
  3. Counts remaining slots against the configured maximum
  4. Returns an allowance status determining if the request should proceed

FastAPI Middleware Integration

The rate_limit_middleware function in autogpt_platform/autogpt_libs/autogpt_libs/rate_limit/middleware.py (lines 7-31) runs on every request whose path starts with /api. It extracts the bearer token from the Authorization header, calls RateLimiter.check_rate_limit(key), and aborts with 429 Too Many Requests if the quota is exceeded. The middleware automatically adds standard X-RateLimit-* response headers including X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset.

Observability Instrumentation

The system records rate limit hits using a Prometheus counter defined in autogpt_platform/backend/backend/monitoring/instrumentation.py (lines 95-103). The metric autogpt_rate_limit_hits_total increments every time an API key hits its quota, enabling monitoring and alerting in production environments.

Configuring the Rate Limit

The default policy allows 60 requests per minute per API key. You can modify this quota using one of three methods:

Set RATE_LIMIT_REQUESTS_PER_MINUTE before starting the platform service:

export RATE_LIMIT_REQUESTS_PER_MINUTE=120

This method works best for containerized deployments using Docker or Kubernetes.

.env File Configuration

Add the variable to your .env file at the repository root:

RATE_LIMIT_REQUESTS_PER_MINUTE=120
REDIS_HOST=redis://my-redis:6379
REDIS_PASSWORD=securepassword

The Pydantic settings loader automatically picks up these values on service startup.

Direct Code Modification

Edit the default value in RateLimitSettings.requests_per_minute at lines 24-27 of config.py. This requires rebuilding the Docker image or restarting the Python service to take effect.

Customising the Window Duration

The limiter currently uses a fixed 60-second sliding window defined in autogpt_platform/autogpt_libs/autogpt_libs/rate_limit/limiter.py at line 23:

self.window = 60  # seconds

To implement a 30-second window instead, modify line 23:

self.window = 30  # new window in seconds

After changing the code, restart the backend service. Note that no environment variable currently exists for the window duration; exposing this setting would require a pull request to the repository.

Handling Rate Limits in Client Code

Client applications should respect the 429 status code and X-RateLimit-Reset header to implement exponential backoff. Here is a Python example demonstrating proper client-side handling:

import requests
import time

api_key = "my-api-key"
headers = {"Authorization": f"Bearer {api_key}"}
url = "https://auto-gpt.example.com/api/v1/agents"

while True:
    resp = requests.get(url, headers=headers)
    
    if resp.status_code == 429:
        reset = int(resp.headers.get("X-RateLimit-Reset", time.time() + 60))
        wait = max(0, reset - time.time())
        print(f"Rate limited – waiting {wait:.1f}s")
        time.sleep(wait)
        continue
    
    # Process successful response

    print(f"Remaining quota: {resp.headers.get('X-RateLimit-Remaining')}")
    break

Summary

  • AutoGPT implements per-API-key rate limiting using a Redis-backed sliding window algorithm in the RateLimiter class.
  • Configuration occurs via environment variables: Set RATE_LIMIT_REQUESTS_PER_MINUTE to adjust the default 60-requests-per-minute quota.
  • Middleware enforces limits on /api/* routes and returns standard X-RateLimit-* headers for client-side throttling.
  • Observability is built-in through the autogpt_rate_limit_hits_total Prometheus metric defined in the instrumentation module.
  • Window duration requires code changes: Edit self.window in limiter.py line 23 to modify the 60-second fixed window.

Frequently Asked Questions

What Redis data structure does AutoGPT use for rate limiting?

AutoGPT uses a Redis sorted set (ZSET) named ratelimit:<key>:1min for each API key. The RateLimiter class adds request timestamps to this set and removes entries older than the window duration to calculate remaining quota.

How do I know when my rate limit will reset?

The rate_limit_middleware adds three headers to every response: X-RateLimit-Limit (total quota), X-RateLimit-Remaining (current remaining requests), and X-RateLimit-Reset (Unix timestamp when the window expires). When you receive a 429 status, check X-RateLimit-Reset to determine exactly when to retry.

Can I set different rate limits for different API endpoints?

The current implementation in middleware.py applies a single global quota per API key to all /api/* routes. To implement endpoint-specific limits, you would need to modify the rate_limit_middleware logic to check request.url.path and instantiate multiple RateLimiter instances with different Redis key prefixes or quota settings.

Is the rate limit global or per API key?

The limit is per API key, not global. Each unique bearer token extracted from the Authorization header receives its own independent quota tracked in separate Redis keys. This ensures that one heavy user cannot exhaust the quota for other API consumers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →