How the Adaptive Rate Limiting System Works in Local Deep Research: Architecture and Tuning Guide

The adaptive rate limiting system in Local Deep Research uses a singleton AdaptiveRateLimitTracker that learns optimal wait times per search engine through an exploration-exploitation algorithm, storing median-based estimates with configurable exponential decay.

The learningcircuit/local-deep-research repository implements an intelligent request pacing mechanism that automatically adjusts to each search engine's tolerance. This adaptive rate limiting system eliminates manual guesswork by analyzing real-world success and failure patterns, persisting learned values across sessions when database access is available.

Architecture of the Adaptive Rate Limiting System

The core implementation resides in src/local_deep_research/web_search_engines/rate_limiting/tracker.py. The system centers on the AdaptiveRateLimitTracker class (L50-L102), which maintains per-engine state and applies machine learning concepts to HTTP request pacing.

The AdaptiveRateLimitTracker Class

The tracker operates as a singleton via the get_tracker() factory function. It initializes with three critical data structures: a settings snapshot capturing runtime configuration, a recent_attempts deque (size controlled by memory_window) storing the last N outcomes per engine, and current_estimates holding learned base, minimum, and maximum wait times.

When instantiated, the constructor calls _apply_profile() (L12-L35) to map user-selected behavioral presets—conservative, balanced, or aggressive—to specific numeric values for exploration probability and learning velocity.

Exploration vs. Exploitation Strategy

The get_wait_time() method (L78-L92) implements a stochastic selection algorithm. With probability exploration_rate, the system selects a faster wait time than currently estimated to probe for more aggressive limits. Otherwise, it exploits existing knowledge by applying small jitter adjustments to the learned base wait.

This dual-mode approach prevents the tracker from stagnating at suboptimal rates while maintaining safety through enforced min/max boundaries. The method returns the calculated delay, which the calling engine applies via time.sleep() before executing HTTP requests.

Learning Algorithm and Persistence

After each request completes, engines call record_outcome() with success status and metadata. This triggers _update_estimate() (L84-L155), which computes a new base wait time using the median of successful waits from the recent history deque. If fewer than three attempts exist or failures dominate, the algorithm falls back to a scaled maximum of observed failure waits.

The system applies an exponential moving average with weight learning_rate to smooth transitions between old and new estimates. Database persistence occurs through RateLimitEstimate rows (defined in src/local_deep_research/database/models.py), with _ensure_estimates_loaded() (L70-L88) applying time-based decay using decay_per_day to discount stale historical data.

How to Tune the Adaptive Rate Limiting System for Different Search Engines

All configuration parameters live under the rate_limiting namespace in the settings system. You can adjust these via environment variables or programmatic snapshots to match engine-specific strictness levels.

Understanding Configuration Settings

Setting Function Default Recommended Range
rate_limiting.enabled Global toggle True true/false
rate_limiting.memory_window Size of per-engine history deque 100 30–500
rate_limiting.exploration_rate Probability of testing faster rates 0.1 0.0–0.3
rate_limiting.learning_rate EMA weight for new estimates 0.3 0.1–0.5
rate_limiting.decay_per_day Daily discount factor for old DB estimates 0.95 0.9–0.99
rate_limiting.profile Preset behavior multiplier balanced conservative/balanced/aggressive

Smaller memory_window values enable quicker adaptation to changing conditions but increase volatility. Lower learning_rate values produce stable, conservative pacing ideal for strict APIs like Google, while higher values suit resilient self-hosted engines such as SearXNG.

Using Predefined Profiles

The _apply_profile() method (L12-L35) implements three optimization presets:

  • Conservative: Caps exploration at 5% and learning rate at 0.2. Use for rate-strict commercial APIs where violations risk account suspension.
  • Balanced: Respects explicit individual settings. Ideal for general research tasks with unknown engine characteristics.
  • Aggressive: Allows up to 20% exploration and 0.5 learning rate. Optimal for high-throughput scenarios where occasional rate limit errors are acceptable.

Configure via environment:

RATE_LIMITING_PROFILE=conservative
RATE_LIMITING_MEMORY_WINDOW=50
RATE_LIMITING_EXPLORATION_RATE=0.02

Per-Engine Customization

For engines requiring unique handling, instantiate a custom tracker with an overridden settings snapshot:

from local_deep_research.web_search_engines.rate_limiting import AdaptiveRateLimitTracker

# Specialized configuration for high-volume Bing usage

bing_snapshot = {
    "rate_limiting.memory_window": 200,
    "rate_limiting.exploration_rate": 0.25,
    "rate_limiting.learning_rate": 0.45,
    "rate_limiting.profile": "balanced",
}

custom_tracker = AdaptiveRateLimitTracker(settings_snapshot=bing_snapshot)

This approach bypasses global defaults without affecting other search engines in your pipeline.

Implementation Guide for Developers

Integrating the adaptive rate limiting system requires calling three primary methods from the singleton tracker instance.

Basic Integration Pattern

Standard usage follows a pre-request delay and post-request feedback pattern:

from local_deep_research.web_search_engines.rate_limiting import get_tracker

tracker = get_tracker()  # Returns initialized singleton

engine_type = "GoogleSearchEngine"

# Apply learned delay before HTTP request

wait_time = tracker.apply_rate_limit(engine_type)

# ... execute actual search request ...

# Record outcome to refine future estimates

tracker.record_outcome(
    engine_type=engine_type,
    wait_time=wait_time,
    success=True,  # Set False if rate limit encountered

    retry_count=0,
    error_type=None,
    search_result_count=15,
)

The apply_rate_limit() method internally calls get_wait_time() and handles the sleep operation, returning the actual delay applied for logging purposes.

Resetting and Inspecting State

To clear learned data for a specific engine and restart the learning process:

tracker.reset_engine("BingSearchEngine")

For debugging or monitoring convergence, retrieve current statistics:

stats = tracker.get_stats("DuckDuckGoSearchEngine")
for engine, base, mn, mx, timestamp, attempts, confidence in stats:
    print(f"{engine}: base={base:.2f}s, confidence={confidence:.2f}, n={attempts}")

When operating outside user contexts (e.g., CI pipelines), disable database persistence:

tracker = AdaptiveRateLimitTracker(programmatic_mode=True)

Summary

  • The adaptive rate limiting system uses a singleton AdaptiveRateLimitTracker that learns per-engine wait times through median-based statistical analysis and exponential moving averages.
  • Exploration vs. exploitation logic randomly tests faster rates while maintaining safe defaults, controlled by the exploration_rate parameter.
  • Persistence optionally stores estimates in SQLAlchemy RateLimitEstimate rows with time-decay weighting via decay_per_day.
  • Tuning occurs through environment variables or programmatic snapshots, using profiles (conservative, balanced, aggressive) or individual settings for memory_window, learning_rate, and exploration probability.
  • Integration requires calling apply_rate_limit() before requests and record_outcome() after completion to close the feedback loop.

Frequently Asked Questions

How does the tracker decide between exploration and exploitation?

The get_wait_time() method in src/local_deep_research/web_search_engines/rate_limiting/tracker.py generates a random value and compares it against the exploration_rate setting (default 0.1). If the random value falls below this threshold, the system selects a wait time faster than the current estimate to test boundary limits. Otherwise, it jitters the learned base estimate slightly to exploit known-good values while adding minor variance to prevent synchronization patterns.

Can I use different rate limiting strategies for different search engines simultaneously?

Yes, while the default get_tracker() returns a global singleton, you can instantiate separate AdaptiveRateLimitTracker instances with distinct settings_snapshot dictionaries for specific engine classes. However, most users achieve per-engine variation through the tracker's internal mapping, which maintains isolated recent_attempts deques and current_estimates for each engine_type string within a single instance.

What happens when the database is unavailable or I'm running in a script?

When initialized with programmatic_mode=True or when user context (username/password) is absent, the tracker skips _load_estimates() and _update_estimate() database operations. It operates purely in-memory using the deque-based history and computed estimates, meaning learned values persist only for the process lifetime. This mode is ideal for serverless functions or batch processing scripts.

How quickly does the system adapt to a search engine changing its rate limits?

Adaptation speed depends on the learning_rate and memory_window settings. With default values (learning_rate=0.3, memory_window=100), the system requires at least three new attempts before updating estimates via _update_estimate(), and the exponential moving average weights new observations at 30% versus 70% historical data. Decreasing memory_window to 30-50 and increasing learning_rate to 0.4-0.5 enables detection of limit changes within 5-10 requests rather than 20-30.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →