How to Set Up Elasticsearch as a Search Engine Backend in Local-Deep-Research
You can configure Elasticsearch as a drop-in search backend by instantiating ElasticsearchManager for indexing and ElasticsearchSearchEngine for querying, connecting to any local or remote Elasticsearch 8.x cluster.
Local-Deep-Research provides a complete Elasticsearch integration that acts as a pluggable replacement for web search engines. The implementation centers on two Python classes that handle index management and search operations, allowing you to index custom documents and perform advanced queries using DSL or query-string syntax.
Core Components Overview
The Elasticsearch backend consists of two primary classes working in tandem.
ElasticsearchManager (Index Operations)
The ElasticsearchManager utility class in src/local_deep_research/utilities/es_utils.py handles low-level index administration. It creates indices with sensible default mappings, bulk-indexes documents, and provides connection management. Key methods include create_index() (lines 95-115) for schema setup and bulk_index_documents() (lines 25-48) for data ingestion.
ElasticsearchSearchEngine (Query Interface)
The ElasticsearchSearchEngine class in src/local_deep_research/web_search_engines/engines/search_engine_elasticsearch.py implements the BaseSearchEngine interface. It manages connection validation (lines 84-95), executes preview-only or full-content retrieval, and supports advanced search syntax. This class integrates directly with the framework's two-phase search flow defined in the abstract base class.
Step-by-Step Configuration
1. Launch Elasticsearch and Install Dependencies
Start an Elasticsearch 8.x instance and install the required Python packages. The repository declares elasticsearch in pyproject.toml, but you must ensure the client version matches your server.
# Run Elasticsearch locally via Docker
docker run -p 9200:9200 -e "discovery.type=single-node" elasticsearch:8.12.0
# Install Python dependencies
pip install elasticsearch==8.12.0 unstructured langchain-community
2. Create an Index with Custom Mappings
Initialize the manager and create an index. If you omit custom mappings, ElasticsearchManager.create_index builds a default schema with fields for title, content, URL, and timestamps.
from src.local_deep_research.utilities.es_utils import ElasticsearchManager
es = ElasticsearchManager(hosts=["http://localhost:9200"])
es.create_index("documents") # Uses default mapping
3. Bulk Index Your Documents
Index documents using the bulk API for efficient ingestion. The refresh=True parameter makes documents immediately searchable.
sample_docs = [
{"title": "Elasticsearch Intro", "content": "Elasticsearch is a distributed search engine…"},
{"title": "Python Basics", "content": "Python is an interpreted high-level language…"}
]
es.bulk_index_documents("documents", sample_docs, refresh=True)
4. Initialize the Search Engine
Instantiate ElasticsearchSearchEngine with your cluster hosts and index name. The constructor validates the connection and raises an error if the cluster is unreachable.
from src.local_deep_research.web_search_engines.engines.search_engine_elasticsearch import ElasticsearchSearchEngine
search_engine = ElasticsearchSearchEngine(
hosts=["http://localhost:9200"],
index_name="documents",
max_results=10
)
5. Execute Basic and Advanced Queries
Run a basic search that returns title and snippet data:
results = search_engine.run("elasticsearch")
for r in results:
print(r["title"], r["snippet"])
For advanced queries, use query-string syntax:
qs_results = search_engine.search_by_query_string(
"content:deep learning OR title:elasticsearch"
)
Or execute raw DSL queries for complex filtering:
dsl = {
"query": {
"bool": {
"must": {"match": {"content": "artificial intelligence"}},
"filter": {"term": {"category.keyword": "AI"}}
}
}
}
dsl_results = search_engine.search_by_dsl(dsl)
Advanced search methods are implemented starting at lines 59 (search_by_query_string) and 100 (search_by_dsl) in search_engine_elasticsearch.py.
Authentication and Production Deployments
Both classes support authentication via username/password, API keys, or Elastic Cloud deployment IDs. Pass these credentials during instantiation:
es = ElasticsearchManager(
hosts=["https://my-cloud-es.es.amazonaws.com"],
api_key="my-api-key"
)
engine = ElasticsearchSearchEngine(
hosts=["https://my-cloud-es.es.amazonaws.com"],
api_key="my-api-key",
index_name="production-docs"
)
Authentication handling is implemented in the constructors (lines 70-80 of both es_utils.py and search_engine_elasticsearch.py).
Complete Working Example
The repository includes a runnable demonstration at examples/elasticsearch/search_example.py that ties together index creation, document ingestion, and query execution:
python examples/elasticsearch/search_example.py
This script demonstrates end-to-end functionality including error handling, logging configuration, and both basic and advanced search patterns.
Summary
- ElasticsearchManager (
es_utils.py) handles index creation and bulk indexing throughcreate_index()andbulk_index_documents(). - ElasticsearchSearchEngine (
search_engine_elasticsearch.py) implements the search interface with support for basic queries, query-string syntax, and raw DSL. - Connection validation occurs automatically during class instantiation, ensuring cluster availability before operations proceed.
- Both classes accept standard Elasticsearch authentication parameters for secure production deployments.
- The example script at
examples/elasticsearch/search_example.pyprovides a complete reference implementation.
Frequently Asked Questions
How do I verify my Elasticsearch connection is working?
The ElasticsearchSearchEngine constructor validates connectivity immediately upon instantiation. If the cluster is unreachable or authentication fails, it raises a connection error before any search operations execute. You can also test manually by calling es.ping() on an ElasticsearchManager instance.
Can I use an existing Elasticsearch index with custom mappings?
Yes. When calling ElasticsearchManager.create_index(), pass ignore_existing=True or simply skip index creation if your index already exists. The search engine reads field data dynamically, though optimal snippet extraction requires title and content fields (configurable in the mapping).
What Elasticsearch versions are supported?
The implementation targets Elasticsearch 8.x. Ensure your Python elasticsearch client version matches your server version (e.g., pip install elasticsearch==8.12.0). Version mismatches between client and server can cause protocol errors or authentication failures.
How do I switch from the default web search to Elasticsearch?
Replace your current search engine instantiation with ElasticsearchSearchEngine, passing the appropriate hosts and index_name. The class inherits from BaseSearchEngine, so it integrates seamlessly with existing research pipelines that expect the standard run() method interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →