How the Candidate Exploration System Works with Exploration Strategies in Local Deep Research

The candidate exploration system in learningcircuit/local-deep-research employs a pluggable architecture where an abstract BaseCandidateExplorer class provides shared utilities for search execution and candidate deduplication, while concrete strategy implementations like AdaptiveExplorer and ParallelExplorer execute distinct exploration policies selected via the ExplorationStrategy enum.

The learningcircuit/local-deep-research repository implements a modular candidate exploration system that decouples search strategy selection from execution mechanics. This design enables the high-level search pipeline to switch dynamically between adaptive, diversity-focused, and constraint-guided approaches while reusing common infrastructure for relevance ranking and result deduplication across all exploration modes.

Core Architecture of the Candidate Exploration System

The exploration framework centers on two foundational components: an abstract base class that enforces a uniform interface and an enumeration that defines the supported strategy types.

BaseCandidateExplorer Abstract Interface

The BaseCandidateExplorer class, defined in src/local_deep_research/advanced_search_system/candidate_exploration/base_explorer.py (lines 43–86), serves as the backbone for all exploration strategies. It supplies shared utilities including search execution, candidate extraction, deduplication, and relevance ranking. Concrete subclasses must implement two abstract methods: explore() to execute the search loop and generate_exploration_queries() to produce query variations. The base class also provides helper methods such as _should_continue_exploration(), _deduplicate_candidates(), and _rank_candidates_by_relevance() (lines 69–87) that ensure consistent behavior across diverse exploration algorithms.

ExplorationStrategy Enum

The ExplorationStrategy enum, located in the same base file (lines 20–28), enumerates the five built-in strategies: BREADTH_FIRST, DEPTH_FIRST, CONSTRAINT_GUIDED, DIVERSITY_FOCUSED, and ADAPTIVE. This enum acts as metadata for tracking which strategy was exercised during a search operation and provides type safety when configuring the exploration pipeline.

Built-In Exploration Strategies

Four concrete explorer implementations ship with the repository, each encoding a distinct search philosophy while inheriting common infrastructure from the base class.

AdaptiveExplorer for Dynamic Heuristic Switching

The AdaptiveExplorer (source: src/local_deep_research/advanced_search_system/candidate_exploration/adaptive_explorer.py, lines 23–86) implements a self-adapting search policy that dynamically switches among a set of heuristics—such as direct_search and synonym_expansion—based on measured success rates. After a configurable number of searches (the adaptation_threshold), the explorer evaluates which heuristics yield the highest quality candidates and adjusts its query generation strategy accordingly, optimizing for both coverage and precision without manual tuning.

DiversityExplorer for Semantic Coverage

The DiversityExplorer (source: src/local_deep_research/advanced_search_system/candidate_exploration/diversity_explorer.py, lines 23–33) focuses on maintaining a candidate set spread across distinct semantic categories. It employs a Shannon-entropy-like scoring mechanism to measure the information diversity of retrieved results, continuing exploration until a configurable diversity_factor threshold is satisfied. This prevents the system from over-concentrating on a single topical cluster when researching broad or multi-faceted queries.

The ParallelExplorer (source: src/local_deep_research/advanced_search_system/candidate_exploration/parallel_explorer.py, lines 23–51) executes multiple queries simultaneously using a ThreadPoolExecutor, favoring a breadth-first coverage of the query space. By default configured with max_workers=4, this strategy is ideal for time-sensitive searches requiring rapid scanning of multiple information sources or query variations in parallel.

ConstraintGuidedExplorer for Targeted Retrieval

The ConstraintGuidedExplorer (source: src/local_deep_research/advanced_search_system/candidate_exploration/constraint_guided_explorer.py, lines 22–30) builds queries that directly target the most important constraints in a research question. It validates candidates early against constraint-specific patterns and ranks results by their alignment with those constraints, ensuring that retrieved information directly supports the specific requirements of the research task rather than general topical relevance.

Strategy Selection in the Modular Pipeline

The high-level ModularStrategy class orchestrates which concrete explorer to instantiate based on the exploration_strategy argument supplied at construction. The factory method _create_candidate_explorer() in src/local_deep_research/advanced_search_system/strategies/modular_strategy.py (lines 61–87) performs the selection:

def _create_candidate_explorer(self, strategy_type: str):
    if strategy_type == "parallel":
        return ParallelExplorer(search_engine=self.search_engine,
                                model=self.model,
                                max_workers=4)
    elif strategy_type == "adaptive":
        return AdaptiveExplorer(search_engine=self.search_engine,
                                model=self.model,
                                learning_rate=0.1)
    elif strategy_type == "constraint_guided":
        return ConstraintGuidedExplorer(search_engine=self.search_engine,
                                        model=self.model)
    elif strategy_type == "diversity":
        return DiversityExplorer(search_engine=self.search_engine,
                                 model=self.model,
                                 diversity_factor=0.3)
    else:
        raise ValueError(f"Unknown exploration strategy: {strategy_type}")

During execution, ModularStrategy.search() builds a list of queries—including original and LLM-generated variations—and delegates them to the selected explorer. The explorer runs its internal loop (or parallel batch) and returns an ExplorationResult containing the final candidate set, meta-information about the exploration path, and the ExplorationStrategy that was exercised (source: modular_strategy.py, lines 146–166). This result is then passed to downstream components for constraint checking and answer synthesis.

Extending the System with Custom Explorers

Adding a new exploration style requires only three steps:

  1. Subclass BaseCandidateExplorer and implement explore() and generate_exploration_queries().
  2. Optionally add a corresponding entry in the ExplorationStrategy enum for metadata consistency.
  3. Update ModularStrategy._create_candidate_explorer to instantiate the new class when its strategy key is provided.

All pipeline components—including constraint checking, early rejection, and result ranking—automatically reuse the shared utilities from the base class, ensuring that custom strategies benefit from robust deduplication and relevance ranking without additional implementation overhead.

Example: Direct Explorer Usage

You can instantiate and run explorers directly for fine-grained control:

from local_deep_research.advanced_search_system.candidate_exploration.adaptive_explorer import AdaptiveExplorer
from local_deep_research.web_search_engines.engines.search_engine_wikipedia import SearchEngineWikipedia
from langchain_core.language_models import ChatOpenAI  # placeholder model

model = ChatOpenAI(model="gpt-4")
search_engine = SearchEngineWikipedia()

explorer = AdaptiveExplorer(
    model=model,
    search_engine=search_engine,
    max_candidates=30,
    max_search_time=45.0,
    adaptation_threshold=4,
)

result = explorer.explore(
    initial_query="renewable energy storage technologies",
    constraints=None,
    entity_type="technology"
)

print("Top candidates:", [c.name for c in result.candidates])

Example: High-Level Strategy Selection

For most use cases, switch strategies via the client API:

from local_deep_research.api.client import LocalDeepResearchClient

client = LocalDeepResearchClient(
    model="gpt-4",
    search_engine="wikipedia",
    exploration_strategy="diversity",   # can be "parallel", "adaptive", "constraint_guided"

    constraint_checker_type="strict"
)

answer, metadata = client.analyze_topic(
    "historical examples of volcanic eruptions that caused global cooling"
)

print(metadata["strategy"], metadata["exploration_strategy"])
print("Best candidate:", answer)

Example: Custom Explorer Skeleton

Implement a new strategy by subclassing the base:

from .base_explorer import BaseCandidateExplorer, ExplorationResult, ExplorationStrategy

class MyCustomExplorer(BaseCandidateExplorer):
    def explore(self, initial_query, constraints=None, entity_type=None):
        # Custom loop – e.g. use a knowledge-graph first, then fallback to web search

        # Return an ExplorationResult just like the built-ins

        ...

    def generate_exploration_queries(self, base_query, found_candidates, constraints=None):
        # Produce new queries based on graph neighbours, etc.

        ...

Summary

  • The candidate exploration system relies on BaseCandidateExplorer in base_explorer.py to provide shared search utilities and enforce a uniform interface across all strategies.
  • The ExplorationStrategy enum defines five built-in modes: BREADTH_FIRST, DEPTH_FIRST, CONSTRAINT_GUIDED, DIVERSITY_FOCUSED, and ADAPTIVE.
  • Concrete implementations include AdaptiveExplorer (heuristic switching), DiversityExplorer (entropy-based scoring), ParallelExplorer (concurrent execution), and ConstraintGuidedExplorer (constraint alignment).
  • ModularStrategy selects the appropriate explorer via _create_candidate_explorer() in modular_strategy.py (lines 61–87) based on the exploration_strategy parameter.
  • Extending the system requires only implementing two abstract methods; all deduplication and ranking logic remains inherited from the base class.

Frequently Asked Questions

How does the candidate exploration system handle duplicate candidates across multiple queries?

The BaseCandidateExplorer class provides the _deduplicate_candidates() helper method (source: base_explorer.py, lines 69–87), which is invoked uniformly across all concrete strategies. This ensures that whether you use the ParallelExplorer or AdaptiveExplorer, identical results from different query variations are merged before ranking and final selection.

Can I combine multiple exploration strategies in a single research session?

While the current ModularStrategy implementation selects a single explorer via _create_candidate_explorer(), the architecture supports sequential or hierarchical composition by manually instantiating explorers and feeding the ExplorationResult from one into the query generation phase of another. Future extensions could implement a meta-explorer that delegates to multiple strategies internally.

What is the difference between the BREADTH_FIRST enum value and the ParallelExplorer?

The BREADTH_FIRST value in the ExplorationStrategy enum serves as metadata indicating a broad, shallow search pattern, while the ParallelExplorer is the concrete implementation that realizes this philosophy using a ThreadPoolExecutor to execute queries concurrently (source: parallel_explorer.py, lines 23–51). The enum tracks the strategy type in the ExplorationResult, while the explorer class contains the actual execution logic.

How does the AdaptiveExplorer determine when to switch heuristics?

The AdaptiveExplorer monitors search success rates across a sliding window defined by the adaptation_threshold parameter (defaulting to 4 searches). After each threshold is reached, it evaluates which heuristics (such as direct_search or synonym_expansion) yielded the highest quality candidates and adjusts the probability distribution for future query generation accordingly (source: adaptive_explorer.py, lines 23–86).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →