# How the Adaptive Rate Limiting System Works in Local Deep Research: Architecture and Tuning Guide

> Discover how the adaptive rate limiting system in Local Deep Research works. Learn to tune wait times per search engine using an exploration-exploitation algorithm and exponential decay.

- Repository: [learningcircuit/local-deep-research](https://github.com/learningcircuit/local-deep-research)
- Tags: architecture
- Published: 2026-03-05

---

**The adaptive rate limiting system in Local Deep Research uses a singleton `AdaptiveRateLimitTracker` that learns optimal wait times per search engine through an exploration-exploitation algorithm, storing median-based estimates with configurable exponential decay.**

The `learningcircuit/local-deep-research` repository implements an intelligent request pacing mechanism that automatically adjusts to each search engine's tolerance. This adaptive rate limiting system eliminates manual guesswork by analyzing real-world success and failure patterns, persisting learned values across sessions when database access is available.

## Architecture of the Adaptive Rate Limiting System

The core implementation resides in [`src/local_deep_research/web_search_engines/rate_limiting/tracker.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/rate_limiting/tracker.py). The system centers on the **`AdaptiveRateLimitTracker`** class (L50-L102), which maintains per-engine state and applies machine learning concepts to HTTP request pacing.

### The AdaptiveRateLimitTracker Class

The tracker operates as a **singleton** via the `get_tracker()` factory function. It initializes with three critical data structures: a settings snapshot capturing runtime configuration, a `recent_attempts` deque (size controlled by `memory_window`) storing the last N outcomes per engine, and `current_estimates` holding learned base, minimum, and maximum wait times.

When instantiated, the constructor calls `_apply_profile()` (L12-L35) to map user-selected behavioral presets—**conservative**, **balanced**, or **aggressive**—to specific numeric values for exploration probability and learning velocity.

### Exploration vs. Exploitation Strategy

The `get_wait_time()` method (L78-L92) implements a stochastic selection algorithm. With probability `exploration_rate`, the system selects a **faster** wait time than currently estimated to probe for more aggressive limits. Otherwise, it **exploits** existing knowledge by applying small jitter adjustments to the learned base wait.

This dual-mode approach prevents the tracker from stagnating at suboptimal rates while maintaining safety through enforced `min/max` boundaries. The method returns the calculated delay, which the calling engine applies via `time.sleep()` before executing HTTP requests.

### Learning Algorithm and Persistence

After each request completes, engines call `record_outcome()` with success status and metadata. This triggers `_update_estimate()` (L84-L155), which computes a new base wait time using the **median of successful waits** from the recent history deque. If fewer than three attempts exist or failures dominate, the algorithm falls back to a scaled maximum of observed failure waits.

The system applies an **exponential moving average** with weight `learning_rate` to smooth transitions between old and new estimates. Database persistence occurs through `RateLimitEstimate` rows (defined in [`src/local_deep_research/database/models.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/database/models.py)), with `_ensure_estimates_loaded()` (L70-L88) applying time-based decay using `decay_per_day` to discount stale historical data.

## How to Tune the Adaptive Rate Limiting System for Different Search Engines

All configuration parameters live under the `rate_limiting` namespace in the settings system. You can adjust these via environment variables or programmatic snapshots to match engine-specific strictness levels.

### Understanding Configuration Settings

| Setting | Function | Default | Recommended Range |
|---------|----------|---------|-------------------|
| `rate_limiting.enabled` | Global toggle | `True` | `true`/`false` |
| `rate_limiting.memory_window` | Size of per-engine history deque | `100` | `30`–`500` |
| `rate_limiting.exploration_rate` | Probability of testing faster rates | `0.1` | `0.0`–`0.3` |
| `rate_limiting.learning_rate` | EMA weight for new estimates | `0.3` | `0.1`–`0.5` |
| `rate_limiting.decay_per_day` | Daily discount factor for old DB estimates | `0.95` | `0.9`–`0.99` |
| `rate_limiting.profile` | Preset behavior multiplier | `balanced` | `conservative`/`balanced`/`aggressive` |

Smaller `memory_window` values enable quicker adaptation to changing conditions but increase volatility. Lower `learning_rate` values produce stable, conservative pacing ideal for strict APIs like Google, while higher values suit resilient self-hosted engines such as SearXNG.

### Using Predefined Profiles

The `_apply_profile()` method (L12-L35) implements three optimization presets:

- **Conservative**: Caps exploration at 5% and learning rate at 0.2. Use for rate-strict commercial APIs where violations risk account suspension.
- **Balanced**: Respects explicit individual settings. Ideal for general research tasks with unknown engine characteristics.
- **Aggressive**: Allows up to 20% exploration and 0.5 learning rate. Optimal for high-throughput scenarios where occasional rate limit errors are acceptable.

Configure via environment:

```bash
RATE_LIMITING_PROFILE=conservative
RATE_LIMITING_MEMORY_WINDOW=50
RATE_LIMITING_EXPLORATION_RATE=0.02

```

### Per-Engine Customization

For engines requiring unique handling, instantiate a custom tracker with an overridden settings snapshot:

```python
from local_deep_research.web_search_engines.rate_limiting import AdaptiveRateLimitTracker

# Specialized configuration for high-volume Bing usage

bing_snapshot = {
    "rate_limiting.memory_window": 200,
    "rate_limiting.exploration_rate": 0.25,
    "rate_limiting.learning_rate": 0.45,
    "rate_limiting.profile": "balanced",
}

custom_tracker = AdaptiveRateLimitTracker(settings_snapshot=bing_snapshot)

```

This approach bypasses global defaults without affecting other search engines in your pipeline.

## Implementation Guide for Developers

Integrating the adaptive rate limiting system requires calling three primary methods from the singleton tracker instance.

### Basic Integration Pattern

Standard usage follows a pre-request delay and post-request feedback pattern:

```python
from local_deep_research.web_search_engines.rate_limiting import get_tracker

tracker = get_tracker()  # Returns initialized singleton

engine_type = "GoogleSearchEngine"

# Apply learned delay before HTTP request

wait_time = tracker.apply_rate_limit(engine_type)

# ... execute actual search request ...

# Record outcome to refine future estimates

tracker.record_outcome(
    engine_type=engine_type,
    wait_time=wait_time,
    success=True,  # Set False if rate limit encountered

    retry_count=0,
    error_type=None,
    search_result_count=15,
)

```

The `apply_rate_limit()` method internally calls `get_wait_time()` and handles the sleep operation, returning the actual delay applied for logging purposes.

### Resetting and Inspecting State

To clear learned data for a specific engine and restart the learning process:

```python
tracker.reset_engine("BingSearchEngine")

```

For debugging or monitoring convergence, retrieve current statistics:

```python
stats = tracker.get_stats("DuckDuckGoSearchEngine")
for engine, base, mn, mx, timestamp, attempts, confidence in stats:
    print(f"{engine}: base={base:.2f}s, confidence={confidence:.2f}, n={attempts}")

```

When operating outside user contexts (e.g., CI pipelines), disable database persistence:

```python
tracker = AdaptiveRateLimitTracker(programmatic_mode=True)

```

## Summary

- The **adaptive rate limiting system** uses a singleton `AdaptiveRateLimitTracker` that learns per-engine wait times through median-based statistical analysis and exponential moving averages.
- **Exploration vs. exploitation** logic randomly tests faster rates while maintaining safe defaults, controlled by the `exploration_rate` parameter.
- **Persistence** optionally stores estimates in SQLAlchemy `RateLimitEstimate` rows with time-decay weighting via `decay_per_day`.
- **Tuning** occurs through environment variables or programmatic snapshots, using profiles (`conservative`, `balanced`, `aggressive`) or individual settings for `memory_window`, `learning_rate`, and exploration probability.
- **Integration** requires calling `apply_rate_limit()` before requests and `record_outcome()` after completion to close the feedback loop.

## Frequently Asked Questions

### How does the tracker decide between exploration and exploitation?

The `get_wait_time()` method in [`src/local_deep_research/web_search_engines/rate_limiting/tracker.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/rate_limiting/tracker.py) generates a random value and compares it against the `exploration_rate` setting (default 0.1). If the random value falls below this threshold, the system selects a wait time faster than the current estimate to test boundary limits. Otherwise, it jitters the learned base estimate slightly to exploit known-good values while adding minor variance to prevent synchronization patterns.

### Can I use different rate limiting strategies for different search engines simultaneously?

Yes, while the default `get_tracker()` returns a global singleton, you can instantiate separate `AdaptiveRateLimitTracker` instances with distinct `settings_snapshot` dictionaries for specific engine classes. However, most users achieve per-engine variation through the tracker's internal mapping, which maintains isolated `recent_attempts` deques and `current_estimates` for each `engine_type` string within a single instance.

### What happens when the database is unavailable or I'm running in a script?

When initialized with `programmatic_mode=True` or when user context (username/password) is absent, the tracker skips `_load_estimates()` and `_update_estimate()` database operations. It operates purely in-memory using the deque-based history and computed estimates, meaning learned values persist only for the process lifetime. This mode is ideal for serverless functions or batch processing scripts.

### How quickly does the system adapt to a search engine changing its rate limits?

Adaptation speed depends on the `learning_rate` and `memory_window` settings. With default values (learning_rate=0.3, memory_window=100), the system requires at least three new attempts before updating estimates via `_update_estimate()`, and the exponential moving average weights new observations at 30% versus 70% historical data. Decreasing `memory_window` to 30-50 and increasing `learning_rate` to 0.4-0.5 enables detection of limit changes within 5-10 requests rather than 20-30.