Mem0 Hosted API vs Self-Hosted: Complete Guide to Limits and Rate Limits

Mem0's hosted platform enforces managed rate limits and memory quotas per workspace that raise specific RateLimitError and MemoryQuotaExceededError exceptions, while the self-hosted OSS version imposes no built-in limits—restrictions arise solely from your underlying LLM providers, vector databases, and hardware capacity.

The mem0ai/mem0 repository offers two distinct deployment modes with fundamentally different approaches to resource constraints. Understanding Mem0 hosted API vs self-hosted limits and rate limits is critical for system architecture, as the hosted service provides automatic quota management and graceful error handling, whereas the open-source version delegates all boundary enforcement to your chosen infrastructure providers.

Managed Rate Limits in the Mem0 Hosted API

The hosted platform implements managed quotas at the workspace level, controlling requests per second, burst capacity, and total monthly usage through centralized infrastructure.

Workspace Quotas and Exception Handling

When you exceed configured thresholds, the MemoryClient raises specific exceptions defined in mem0/exceptions.py. The platform returns structured error responses containing retry_after metadata for exponential backoff implementation.

from mem0 import MemoryClient
from mem0.exceptions import RateLimitError, MemoryQuotaExceededError
import time

client = MemoryClient(api_key="YOUR_MEM0_API_KEY")

def safe_add(messages, user_id):
    try:
        client.add(messages, user_id=user_id)
    except RateLimitError as e:
        # Exponential back-off based on hosted platform's retry guidance

        wait = e.debug_info.get("retry_after", 30)
        print(f"Rate limit hit – retrying in {wait}s")
        time.sleep(wait)
        return safe_add(messages, user_id)
    except MemoryQuotaExceededError as e:
        print("Memory quota exhausted – upgrade your plan or prune old memories")
        raise

Source: The MemoryClient.add() method in mem0/client/main.py catches HTTP 429 responses and wraps them in RateLimitError, while mem0/exceptions.py defines the exception hierarchy.

Real-Time Quota Visibility

Hosted users monitor consumption through the Mem0 dashboard or programmatically via the users() endpoint, which returns current request_quota and memory_quota statistics.

from mem0 import MemoryClient

client = MemoryClient(api_key="YOUR_MEM0_API_KEY")
usage = client.users()  # Returns per-workspace quota consumption

print(usage)  # Contains request rates, storage utilization, and limits

Rate Limits in Self-Hosted Mem0 (OSS)

The self-hosted version (Memory class) operates as a thin orchestration layer over your chosen LLM, embedding provider, and vector store. It enforces no intrinsic rate limits or memory quotas.

Provider-Imposed Constraints

Rate limiting in OSS deployments originates entirely from upstream services. For example, OpenAI typically enforces 60 requests per second (RPS) for GPT models, while Anthropic Claude APIs may limit you to 5 RPS. The SDK forwards these provider-specific errors directly to your application.

from mem0 import Memory
from mem0.exceptions import RateLimitError
import time

mem = Memory()  # Uses locally configured LLM/embedder/vector store

def add_with_provider_backoff(messages, user_id):
    try:
        mem.add(messages, user_id=user_id)
    except RateLimitError as e:
        # Error originates from OpenAI, Anthropic, or other providers, not Mem0

        wait = e.debug_info.get("retry_after", 20)
        print(f"Provider rate-limit – sleeping {wait}s")
        time.sleep(wait)
        return add_with_provider_backoff(messages, user_id)

Source: mem0/__init__.py exports the Memory class, which routes requests to configured providers in mem0/vector_stores/*.py and LLM clients without intermediary rate checking.

Infrastructure and Storage Boundaries

Self-hosted storage limits depend on your vector database allocation (Qdrant, Weaviate, Chroma, etc.) and available disk/memory resources. Unlike the hosted platform, no MemoryQuotaExceededError originates from Mem0 itself—your vector store returns connection errors or storage exhaustion signals instead.


# Example: Monitoring Qdrant vector store capacity

from qdrant_client import QdrantClient

qdrant = QdrantClient(host="localhost", port=6333)
stats = qdrant.get_collection("mem0_vectors").vectors_count
print(f"Stored vectors: {stats} / your_hardware_limit")

Memory Quotas: Platform vs. OSS

Hosted environments allocate fixed vector storage per workspace (e.g., X GiB of vectors, Y million memories). Exceeding these triggers MemoryQuotaExceededError with actionable upgrade paths.

Self-hosted deployments offer unlimited capacity constrained only by your provisioned hardware, cloud database tiers, and network bandwidth. You must implement your own pruning strategies or storage alerts outside the Mem0 SDK.

Scaling and Observability Differences

Automatic Scaling (Hosted)

The Mem0 platform provides automatic auto-scaling—you never provision additional nodes or shard vector indexes manually. Rate limits adjust according to your subscription tier, with explicit documentation available in docs/platform/platform-vs-oss.mdx.

Manual Instrumentation (Self-Hosted)

OSS deployments require you to size and scale your vector DB, LLM inference servers, and graph stores independently. Monitor request rates using Prometheus, Datadog, or provider-specific dashboards rather than Mem0-native tooling.

Summary

  • Hosted API: Enforces managed per-workspace rate limits and memory quotas, surfacing RateLimitError and MemoryQuotaExceededError through mem0/client/main.py with automatic scaling and dashboard visibility.
  • Self-Hosted: Imposes no built-in limits; restrictions come from underlying providers (OpenAI, Anthropic) and your hardware capacity, requiring manual monitoring and infrastructure scaling.
  • Error Handling: Both modes use the same exception classes from mem0/exceptions.py, but hosted errors originate from Mem0 infrastructure while OSS errors bubble up from external services.
  • Storage: Hosted provides managed quotas with upgrade paths; OSS provides unlimited storage bounded only by your vector database and disk allocation.

Frequently Asked Questions

What happens when I exceed rate limits on the Mem0 hosted API?

The hosted API returns a RateLimitError containing retry_after metadata in the debug_info dictionary. The SDK raises this exception in mem0/client/main.py immediately upon receiving HTTP 429 responses, allowing you to implement exponential backoff or queue strategies.

Does self-hosted Mem0 have any built-in rate limiting?

No. The self-hosted Memory class forwards all requests directly to your configured LLM and vector store providers. Any RateLimitError exceptions originate from services like OpenAI or Anthropic, not from Mem0 itself, as the OSS codebase contains no request throttling logic.

How do I check my current quota usage on the hosted platform?

Call the users() method on your MemoryClient instance to retrieve real-time quota consumption, or view the metrics in the Mem0 dashboard. This endpoint returns structured data including request_quota and memory_quota fields showing current utilization against your plan limits.

Can I encounter memory quota errors in self-hosted Mem0?

No. The self-hosted version does not enforce MemoryQuotaExceededError. Storage exhaustion manifests as database connection errors or disk-space warnings from your chosen vector store (Qdrant, Weaviate, etc.), requiring you to monitor storage independently through database-specific tools or cloud provider metrics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →