How to Integrate Custom LangChain Retrievers as Search Sources in Local Deep Research

Local Deep Research treats any LangChain BaseRetriever as a first-class search engine through a thread-safe global registry, enabling you to query private vector stores and databases alongside web sources using the retrievers and search_tool parameters.

The learningcircuit/local-deep-research framework extends beyond traditional web search by accepting custom LangChain retrievers as native search sources. This architecture allows you to integrate proprietary vector databases, SQL stores, or internal document repositories without modifying core library code. By leveraging the global Retriever Registry, you can combine private knowledge bases with public search engines in a single unified research workflow.

Understanding the Retriever Registry Architecture

The integration operates through three coordinated components that bridge LangChain retrievers with the research pipeline.

The Global Retriever Registry

The Retriever Registry is a thread-safe global singleton defined in src/local_deep_research/web_search_engines/retriever_registry.py. This registry stores name-to-retriever mappings and exposes a simple API including register, register_multiple, get, unregister, and list_registered. Because the registry persists at the process level, you can register retrievers once and reuse them across multiple research calls without reinitialization overhead.

Research Function Integration

High-level API functions—including quick_summary, detailed_research, and generate_report—accept an optional retrievers dictionary parameter. When provided, these functions automatically import the registry and invoke retriever_registry.register_multiple(...) to populate the global store, as implemented in src/local_deep_research/api/research_functions.py at lines 46-52.

Search System Construction

The internal _init_search_system function constructs an AdvancedSearchSystem that resolves retriever names to concrete instances. The system reads the search_tool argument from kwargs and queries the registry to retrieve the matching object, executing retriever.get_relevant_documents(query) to fetch context. This design ensures retrievers receive identical treatment to built-in web search engines, supporting pagination, iterative questioning, and meta-search strategies without modification.

Registering and Using Custom Retrievers

You can integrate custom retrievers by passing them directly to research functions or by manipulating the registry for advanced use cases.

Vector Store Retrievers (FAISS Example)

To use a FAISS vector store as a search source, instantiate your retriever and pass it via the retrievers dictionary, then select it using search_tool:

from langchain.vectorstores import FAISS
from langchain.embeddings import OpenAIEmbeddings
from local_deep_research.api import quick_summary

# Create a FAISS retriever

embeddings = OpenAIEmbeddings()
vectorstore = FAISS.from_documents(documents, embeddings)
retriever = vectorstore.as_retriever()

# Use it as a search source

result = quick_summary(
    query="What are our deployment procedures?",
    retrievers={"company_kb": retriever},
    search_tool="company_kb",
)
print(result["summary"])

Combining Multiple Retrievers

For hybrid retrieval across different data types, register multiple retrievers simultaneously and set search_tool="auto" to query all sources:

from langchain.vectorstores import Chroma
from my_sql_retriever import SqlRetriever
from local_deep_research.api import detailed_research

vector_retriever = Chroma(...).as_retriever()
sql_retriever = SqlRetriever(connection_string="postgres://...")

result = detailed_research(
    query="Explain the latest safety regulations",
    retrievers={
        "vector_db": vector_retriever,
        "sql_db": sql_retriever,
    },
    search_tool="auto",
)
print(result["findings"])

Implementing Custom BaseRetriever Subclasses

For proprietary data stores or specialized retrieval logic, subclass BaseRetriever and implement the required methods:

from langchain.schema import Document, BaseRetriever
from local_deep_research.api import quick_summary

class TestRetriever(BaseRetriever):
    def get_relevant_documents(self, query: str):
        return [Document(page_content=f"Test doc about {query}")]

    async def aget_relevant_documents(self, query: str):
        return self.get_relevant_documents(query)

# Register and use

result = quick_summary(
    query="test integration",
    retrievers={"test": TestRetriever()},
    search_tool="test",
)
print(result["summary"])

Hybrid Search with Web Engines

Combine custom retrievers with traditional web search engines by listing them in the search_engines parameter:

result = quick_summary(
    query="Compare internal policies with public standards",
    retrievers={"internal": internal_retriever},
    search_tool="auto",
    search_engines=["internal", "wikipedia", "searxng"],
)

This configuration queries your internal retriever alongside Wikipedia and SearXNG, enabling comprehensive cross-source research without coupling your retrieval logic to the LDR codebase.

Summary

  • The Retriever Registry in src/local_deep_research/web_search_engines/retriever_registry.py provides a thread-safe global store for LangChain BaseRetriever instances.
  • Research functions automatically register provided retrievers via register_multiple before executing queries.
  • Use the search_tool parameter to select specific retrievers by name, or specify "auto" to aggregate results from all registered sources.
  • Custom retrievers support full LDR capabilities including pagination, iterative questioning, and meta-search strategies.
  • Hybrid configurations allow combining private vector stores, SQL databases, and web search engines in unified research queries.

Frequently Asked Questions

Can I register retrievers globally instead of passing them to every function call?

Yes. While research functions accept the retrievers parameter for convenience, you can import the global registry directly from src/local_deep_research/web_search_engines/retriever_registry.py and call register(name, retriever) once during application initialization. The registry is thread-safe and persists across all subsequent research calls in the same process.

Does Local Deep Research support asynchronous retrievers?

Yes. The AdvancedSearchSystem checks for aget_relevant_documents methods on retriever instances and handles async retrieval appropriately. When implementing custom BaseRetriever subclasses, define both get_relevant_documents for synchronous operations and aget_relevant_documents for async contexts to ensure full compatibility with the research pipeline.

How do I unregister or replace a retriever during runtime?

The registry exposes an unregister(name) method that removes existing retrievers by name. Subsequent calls to register with the same name will replace the previous instance. You can also call list_registered() to inspect current registrations before modifying the registry state.

Can I use custom retrievers with detailed report generation?

Absolutely. The detailed_research and generate_report functions support the identical retrievers and search_tool parameters as quick_summary. According to the source code in src/local_deep_research/api/research_functions.py, all high-level API functions forward these arguments to _init_search_system, ensuring consistent behavior across simple summaries and comprehensive multi-source reports.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →