How to Configure Multiple Search Engines for Meta-Search in Local Deep Research

Configure multiple search engines for meta-search by setting use_in_auto_search to true for each desired engine in a settings snapshot, selecting the "meta" search tool, and optionally passing a meta_search_config dictionary to control aggregation and deduplication.

Local Deep Research ships with a powerful meta-search engine that dynamically combines web search providers like SearXNG, Wikipedia, arXiv, and PubMed into unified research results. When you configure multiple search engines for meta-search, you create a runtime configuration that controls engine discovery, ranking heuristics, and result aggregation. The system uses these settings to automatically select the best engines for each query and merge their outputs into a single coherent response.

How the Meta-Search Engine Discovers Available Engines

The MetaSearchEngine class, located in [src/local_deep_research/web_search_engines/engines/meta_search_engine.py](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/engines/meta_search_engine.py), dynamically discovers available engines during initialization by calling self._get_available_engines().

This method evaluates every configured engine against three strict criteria:

  1. Auto-search enabled: The setting search.engine.web.<engine>.use_in_auto_search must be true
  2. API key validated: If the engine requires an API key (requires_api_key), the caller must set use_api_key_services=True and provide a valid api_key
  3. Internal exclusion: Engine names meta and auto are filtered out (these are reserved internal selectors)

If no engine satisfies these criteria, the system raises a RuntimeError with the message "No search engines enabled for auto search…".

Query Analysis and Engine Selection

Once configured, the meta-search engine determines which engines to invoke for a specific query through the analyze_query() method (lines 64-140).

Domain-specific heuristics take precedence. If your query contains keywords like "arxiv", "pubmed", or "github", the engine immediately returns a short list of matching available engines.

When no heuristic triggers, the system:

  • Prioritizes SearXNG as the general-purpose aggregator
  • Orders remaining engines by their reliability value from the configuration (higher values rank first)

If an LLM is available, the engine constructs a detailed prompt (starting at line 78) listing each engine's description, strengths, and weaknesses, requesting a comma-separated list of 1-3 engine names. The response is validated against available_engines, with SearXNG appended as a fallback if validation fails.

Configuring Engines via Settings Snapshots

To enable specific engines for meta-search, create a settings snapshot using create_settings_snapshot from local_deep_research.api.settings_utils. This function accepts a dictionary matching the structure defined in docs/CONFIGURATION.md.

from local_deep_research.api.settings_utils import create_settings_snapshot

settings = create_settings_snapshot({
    # Enable engines for meta-search discovery

    "search.engine.web.wikipedia.use_in_auto_search": {"value": True},
    "search.engine.web.arxiv.use_in_auto_search": {"value": True},
    "search.engine.web.pubmed.use_in_auto_search": {"value": True},
    # Optional: configure reliability scores for ranking

    "search.engine.web.wikipedia.reliability": {"value": 80},
    "search.engine.web.arxiv.reliability": {"value": 90},
    "search.engine.web.pubmed.reliability": {"value": 85},
})

Keys follow the pattern search.engine.web.<engine_name>.<property>. The use_in_auto_search boolean flag is mandatory for inclusion in the meta-search pool.

Runtime Configuration with meta_search_config

Fine-tune meta-search behavior by passing a meta_search_config dictionary to the high-level API. This configuration is forwarded unmodified to MetaSearchEngine and can override auto-discovery settings.

Available parameters include:

  • engines: Explicit list of engine names to include (overrides auto-discovery)
  • aggregate: Boolean controlling whether to merge results from all engines
  • deduplicate: Boolean enabling content-hash deduplication across sources
  • max_results_per_engine: Integer capping contributions per engine

Reference the demonstration in examples/api_usage/programmatic/hybrid_search_example.py (function demonstrate_meta_search_config, lines 15-45) for production usage patterns.

End-to-End Configuration Example

The following example configures Wikipedia, arXiv, and PubMed for meta-search, explicitly selects two engines, and enables result aggregation:

from local_deep_research.api import quick_summary
from local_deep_research.api.settings_utils import create_settings_snapshot

# 1. Build settings snapshot enabling desired engines

settings = create_settings_snapshot({
    "search.engine.web.wikipedia.use_in_auto_search": {"value": True},
    "search.engine.web.arxiv.use_in_auto_search": {"value": True},
    "search.engine.web.pubmed.use_in_auto_search": {"value": True},
    "search.engine.web.arxiv.reliability": {"value": 95},  # Prefer arXiv

})

# 2. Execute meta-search with custom configuration

result = quick_summary(
    query="Latest breakthroughs in quantum error correction",
    settings_snapshot=settings,
    search_tool="meta",                       # Activate meta-search

    meta_search_config={
        "engines": ["arxiv", "wikipedia"],   # Explicit selection

        "aggregate": True,                   # Merge all results

        "deduplicate": True,                 # Remove duplicates

        "max_results_per_engine": 5,
    },
    iterations=2,
    questions_per_iteration=3,
    programmatic_mode=True,
)

print(f"Summary: {result['summary']}")
print(f"Sources: {len(result.get('sources', []))}")

Under the Hood: Meta-Search Execution Flow

Understanding the internal pipeline helps debug configuration issues:

Step Component Action
A create_settings_snapshot Converts your configuration dict into the internal snapshot format expected by the engine factory
B SearchToolFactory Instantiates MetaSearchEngine when search_tool="meta" is specified
C MetaSearchEngine.__init__ Calls _get_available_engines() to build the whitelist from your snapshot
D MetaSearchEngine.analyze_query Ranks engines using heuristics, reliability scores, or LLM prompts
E _get_previews Creates engine instances on-demand and collects preliminary results
F _get_full_content Retrieves complete content when search.snippets_only is false
G Aggregation Applies aggregate and deduplicate flags from meta_search_config

Summary

  • Enable engines for meta-search by setting search.engine.web.<name>.use_in_auto_search to true in a settings snapshot created via create_settings_snapshot
  • Activate meta-search mode by passing search_tool="meta" to the API or setting search.tool to "meta"
  • Control engine ranking using the reliability configuration value (higher values = higher priority)
  • Override auto-discovery and tune aggregation using the meta_search_config dictionary with keys like engines, aggregate, deduplicate, and max_results_per_engine
  • The MetaSearchEngine class in meta_search_engine.py automatically handles engine discovery, query analysis, and result merging according to your configuration

Frequently Asked Questions

You must create a settings snapshot that sets use_in_auto_search to true for at least one engine, then pass search_tool="meta" to the API. For example, enabling only Wikipedia requires: {"search.engine.web.wikipedia.use_in_auto_search": {"value": True}}. Without this flag, MetaSearchEngine._get_available_engines() finds no valid engines and raises a RuntimeError.

How does Local Deep Research prioritize which search engine to use first?

The system uses a three-tier ranking in analyze_query(): first, domain-specific heuristics (e.g., "arxiv" in the query boosts the arXiv engine); second, reliability scores from your configuration; third, LLM-based selection when a language model is available. SearXNG serves as the universal fallback when no specific heuristic matches.

Yes, but you must set use_api_key_services=True and ensure the engine configuration includes a valid api_key. The MetaSearchEngine validates both the flag and key presence during _get_available_engines(). If the key is missing or the flag is false, the engine is excluded from the available pool even if use_in_auto_search is enabled.

How do I troubleshoot when no search engines are available?

The error "No search engines enabled for auto search" indicates that MetaSearchEngine._get_available_engines() found no engines matching the criteria. Verify that: (1) your settings snapshot includes use_in_auto_search: true for desired engines, (2) API-key engines have valid credentials if use_api_key_services is enabled, and (3) you are not accidentally filtering out all engines with an explicit engines list in meta_search_config.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →