# Troubleshooting Search Engine Rate Limits and API Errors in local-deep-research

> Resolve search engine rate limit and API errors in local deep research. Bypass throttling, reset wait times, and verify API keys with these troubleshooting steps.

- Repository: [learningcircuit/local-deep-research](https://github.com/learningcircuit/local-deep-research)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Set `DISABLE_RATE_LIMITING=true` to bypass throttling immediately, use the CLI reset command to clear corrupted wait times, and verify API keys match your tier limits to resolve most search engine rate limit and authentication errors.**

The `learningcircuit/local-deep-research` repository implements an adaptive throttling mechanism to prevent search engine bans, but aggressive scraping or misconfigured credentials frequently trigger HTTP 429 errors and API rejections. Understanding the interaction between the `AdaptiveRateLimitTracker`, environment variables, and engine-specific settings is essential for maintaining reliable research pipelines. This guide provides actionable troubleshooting steps derived directly from the source code to diagnose and fix these common issues.

## How the Rate-Limiting Architecture Works

### The AdaptiveRateLimitTracker Core

Located in [[`src/local_deep_research/web_search_engines/rate_limiting/tracker.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/rate_limiting/tracker.py)](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/rate_limiting/tracker.py), the **AdaptiveRateLimitTracker** class implements an exploration-exploitation algorithm that learns optimal wait times per engine. It maintains an in-memory queue of recent attempts and adjusts a "base wait" interval based on success or failure feedback. When a search engine raises a `RateLimitError` (defined in [[`exceptions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/exceptions.py)](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/rate_limiting/exceptions.py)), the calling engine queries `tracker.get_wait_time(engine)`, sleeps for the recommended duration, and records the outcome via `record_outcome()`.

### Configuration and Environment Controls

The [[`env_registry.py`](https://github.com/learningcircuit/local-deep-research/blob/main/env_registry.py)](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/settings/env_registry.py) file manages the **DISABLE_RATE_LIMITING** flag. When set to `true`, `1`, or `yes`, the function `is_rate_limiting_enabled()` returns `False`, causing the tracker to return minimal wait times (0.1s) and treating rate-limit errors as regular failures rather than retry triggers. Additional settings like `rate_limiting.profile` (aggressive/conservative), `rate_limiting.learning_rate`, and `rate_limiting.memory_window` control the adaptive behavior.

## Diagnosing Common Rate-Limit and API Error Symptoms

### "Rate Limit Exceeded" from OpenAI, DuckDuckGo, or SearXNG

This occurs when the remote service rejects requests due to excessive frequency or insufficient account quotas. Verify the issue by checking logs in [`local_deep_research/web/routes/api_routes.py`](https://github.com/learningcircuit/local-deep-research/blob/main/local_deep_research/web/routes/api_routes.py) or running the CLI status command. Resolve by switching to a **conservative** profile in settings, reducing `questions_per_iteration` or `search.iterations` to lower request volume, or upgrading your API key tier.

### Empty Results Despite Successful Connections

If DuckDuckGo or SearXNG return empty payloads without throwing exceptions, the engine is likely being silently throttled. Test connectivity manually with `curl` to the endpoint, then reset the engine's learned limits via the CLI or temporarily disable rate limiting with the environment variable to confirm the diagnosis.

### API Key Format and Permission Errors

OpenAI and OpenRouter errors often stem from malformed keys (missing `sk-` prefix), extra whitespace, or insufficient model permissions. Consult the "OpenAI API Errors" section in the troubleshooting documentation and ensure your key supports the requested model (e.g., `gpt-4` vs `gpt-3.5-turbo`).

### Database Persistence Warnings

When running in CI environments or without a logged-in user context, you may see "Skipping database..." messages. This happens because the encrypted user database is unavailable or the `user_password` is not set in the thread context. For automated scripts, rely on the in-memory tracker by ensuring `programmatic_mode=True` is set when constructing custom tracker instances, or simply accept that limits won't persist between runs.

## Troubleshooting Steps and Solutions

### Disable Rate Limiting Globally

For debugging or CI pipelines, set the environment variable to bypass all throttling:

```bash
export DISABLE_RATE_LIMITING=true
python -m local_deep_research.run_benchmark

```

This effectively disables the tracker consultation, preventing back-off loops but removing protection from IP bans.

### Reset Corrupted Engine Statistics

If the tracker has learned excessively long wait times due to transient network issues, reset specific engines using the CLI defined in [[`cli.py`](https://github.com/learningcircuit/local-deep-research/blob/main/cli.py)](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/rate_limiting/cli.py):

```bash
python -m local_deep_research.web_search_engines.rate_limiting.cli reset \
    --engine DuckDuckGoSearchEngine

```

### Adjust Throttling Aggressiveness

Modify the profile setting to `conservative` for longer, safer intervals, or `aggressive` for faster but riskier scraping. Access the tracker programmatically to force profile changes during debugging:

```python
from local_deep_research.web_search_engines.rate_limiting.tracker import get_tracker

tracker = get_tracker()
stats = tracker.get_stats("DuckDuckGoSearchEngine")
print(stats)  # View current base_wait and success rate

# Force conservative profile (debugging only)

tracker._apply_profile("conservative")

```

### Export Data for Analysis

Investigate patterns across multiple engines by exporting the tracker's state:

```bash
python -m local_deep_research.web_search_engines.rate_limiting.cli export \
    --format csv > rate_limits.csv

```

## Handling RateLimitError in Custom Implementations

When building custom downloaders or search wrappers, catch `RateLimitError` and invoke the tracker's sleep mechanism:

```python
from local_deep_research.web_search_engines.rate_limiting import RateLimitError, get_tracker

class CustomDownloader:
    def fetch(self, query, engine_name="CustomEngine"):
        try:
            return self._make_request(query)
        except RateLimitError:
            # Retrieves recommended wait time and sleeps automatically

            get_tracker().apply_rate_limit(engine_name)
            return self.fetch(query, engine_name)  # Retry once

```

## Summary

- **The `AdaptiveRateLimitTracker`** in [`tracker.py`](https://github.com/learningcircuit/local-deep-research/blob/main/tracker.py) learns optimal delays through feedback loops, storing data in-memory with optional encrypted database persistence when a user context is available.
- **Set `DISABLE_RATE_LIMITING=true`** to instantly bypass throttling for debugging, though this risks IP bans and removes automatic retry logic.
- **Use the CLI** (`reset`, `export`, `status`) to manage corrupted statistics or analyze throttling patterns across engines like DuckDuckGo and SearXNG.
- **Verify API keys** for correct prefixes, whitespace, and tier permissions when encountering authentication errors from OpenAI or OpenRouter.
- **Handle `RateLimitError`** programmatically by calling `apply_rate_limit()` to respect the tracker's recommended delays and avoid hard-coded sleep intervals.

## Frequently Asked Questions

### How do I completely disable rate limiting for a test environment?

Set the `DISABLE_RATE_LIMITING` environment variable to `true`, `1`, or `yes` before running your script. According to [[`env_registry.py`](https://github.com/learningcircuit/local-deep-research/blob/main/env_registry.py)](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/settings/env_registry.py), this forces `is_rate_limiting_enabled()` to return `False`, causing the tracker to return minimal wait times and preventing automatic retries on 429 errors.

### Why does my search return empty results even though the API is reachable?

Silent throttling often causes engines like DuckDuckGo or SearXNG to return empty payloads instead of explicit errors when the `AdaptiveRateLimitTracker` has increased wait times beyond your timeout threshold. Run a manual `curl` request to verify connectivity, then reset the engine's statistics using the CLI reset command or temporarily disable rate limiting to confirm if throttling is the cause.

### What should I do when I see "Skipping database" warnings in the logs?

This indicates the tracker is operating without the encrypted user database, typically in CI environments or when running outside a normal `LDRClient` session where the `user_password` is not set in the thread context. The system falls back to in-memory storage automatically. For scripts, ensure you instantiate the tracker with `programmatic_mode=True` to suppress warnings, or simply note that rate-limit data will not persist between process restarts.

### How can I programmatically adjust how aggressive the throttling is?

Access the tracker via `get_tracker()` from [`tracker.py`](https://github.com/learningcircuit/local-deep-research/blob/main/tracker.py) and inspect current statistics with `get_stats(engine_name)`. While the `_apply_profile()` method exists for debugging, production adjustments should be made via the `rate_limiting.profile` setting (set to `conservative` or `aggressive`) in your configuration files to control the exploration-exploitation trade-off safely without modifying private methods.