How Document Analysis Mode Searches Private Documents with AI in Local Deep Research
Document analysis mode uses the analyze_documents() orchestrator to connect to private document collections by selectively disabling SSRF protections via allow_private_ips=True, retrieving content from internal services like Paperless-ngx, and synthesizing AI-generated summaries through a local language model.
The learningcircuit/local-deep-research repository enables secure AI analysis of sensitive documents stored on private networks. Unlike public web search modes, document analysis specifically targets local collections—such as Paperless-ngx instances or Elasticsearch indexes—while maintaining security boundaries through controlled network access and verified file operations.
The Orchestration Entry Point: analyze_documents()
The document analysis workflow centers on analyze_documents() in src/local_deep_research/api/research_functions.py. This function serves as the primary API entry point, coordinating language model initialization, private search engine instantiation, and secure result persistence. It accepts parameters including query, collection_name, max_results, temperature, and optional output_file to control the analysis behavior.
Step-by-Step Private Document Search Flow
1. Language Model Initialization
The process begins by creating an LLM instance via get_llm(temperature=...), typically found around lines 65-66. This configured model handles the final summarization step, with temperature controls allowing adjustment of creative versus deterministic output.
2. Private Search Engine Selection
Next, get_search(collection_name, llm_instance=llm) instantiates a search engine bound to the specified private collection, as implemented around lines 68-70. The collection_name parameter directs the engine to communicate with local services rather than public APIs, selecting implementations like Paperless-ngx or Library-RAG connectors.
3. Bypassing SSRF Protections for Internal Networks
Private document engines require access to internal IP ranges (192.168.*, ::1, localhost). Each engine uses safe_get()—a security wrapper around requests.get—with explicit flags allow_private_ips=True and often allow_localhost=True. This temporarily disables default SSRF protection for that specific request while maintaining audit logging. The PaperlessSearchEngine implementation in src/local_deep_research/web_search_engines/engines/search_engine_paperless.py demonstrates this pattern around lines 29-38.
4. Executing the Document Retrieval
The search.run(query) method dispatches the request to the private service, returning a structured list of dictionaries containing title, link, and content/snippet fields (lines 78-82). This raw retrieval happens entirely within the private network boundary.
5. Context Window Management
To respect LLM token limits, the system truncates the first five results to 1,000 characters each. These excerpts are concatenated into a single prompt instructing the model to "Analyze these document excerpts … Provide a concise summary" (lines 91-104).
6. AI Summarization
The constructed prompt is sent to the LLM via llm.invoke(summary_prompt). The raw response is processed to remove any "think" tags or artifacts, producing the final summary output (lines 13-22).
7. Secure Output Handling
Results are returned as a dictionary containing the summary, raw documents list, and metadata (collection, document_count). When output_file is specified, write_file_verified() from src/local_deep_research/security/file_write_verifier.py safely persists the report to disk, enforcing repository security policies (lines 23-30).
Practical Implementation Examples
Query a Paperless-ngx collection containing company policies:
from local_deep_research.api.research_functions import analyze_documents
result = analyze_documents(
query="Company policy on remote work",
collection_name="paperless_my_company",
max_results=8,
temperature=0.5,
output_file="remote_work_summary.md",
)
print(result["summary"])
Analyze security guidelines from a local Elasticsearch index:
from local_deep_research.api.research_functions import analyze_documents
summary = analyze_documents(
query="Latest security guidelines",
collection_name="my_elastic_index",
max_results=5,
temperature=0.7,
).get("summary")
print("AI-generated summary:", summary)
Key Source Files and Architecture
src/local_deep_research/api/research_functions.py– Containsanalyze_documents(), orchestrating LLM creation, engine selection, search execution, and summarization.src/local_deep_research/web_search_engines/search_engine_base.py– Defines theBaseSearchEngineabstract class establishing therun()contract for all engine implementations.src/local_deep_research/web_search_engines/engines/search_engine_paperless.py– Reference implementation showingsafe_get(..., allow_private_ips=True)usage for internal services.src/local_deep_research/security/ssrf_validator.py– Provides the SSRF guard andallow_private_ipsflag logic for secure private network access.src/local_deep_research/config/search_config.py– Implementsget_search(), constructing parameter dictionaries and factory dispatch for private engines.src/local_deep_research/security/file_write_verifier.py– Houseswrite_file_verified(), ensuring optional report output respects security boundaries.
Summary
- Document analysis mode targets private collections via configurable search engines bound to local services like Paperless-ngx or Elasticsearch.
- SSRF protection is selectively disabled using
allow_private_ips=Truewithinsafe_get()wrappers, enabling access to192.168.x.xand localhost while maintaining request logging. - The
analyze_documents()orchestrator manages the full lifecycle: LLM initialization, private search execution, context truncation (max 5 documents, 1,000 characters each), and AI summarization. - Security persists through all stages, from controlled network access to verified file writes via
write_file_verified().
Frequently Asked Questions
How does document analysis mode differ from public web search?
Public web search uses engines like DuckDuckGo or Google that respect SSRF protections and query external APIs. Document analysis mode explicitly uses get_search() with a collection_name parameter to instantiate private engines that target internal services, coupled with allow_private_ips=True to reach documents on private subnets or localhost.
What security measures protect against malicious private IP exploitation?
The safe_get() wrapper in src/local_deep_research/security/ssrf_validator.py provides granular control. While allow_private_ips=True permits access to RFC 1918 ranges, the flag is set explicitly per engine instance and requests are logged for audit. Additionally, write_file_verified() ensures that any persisted output undergoes security validation before disk write.
Can I analyze documents from multiple private collections simultaneously?
The current analyze_documents() implementation accepts a single collection_name per invocation. To analyze multiple collections, invoke the function separately for each collection and aggregate results programmatically, or extend the orchestrator to accept multiple engine instances.
Which LLM providers are compatible with document analysis mode?
Any provider supported by the repository's get_llm() factory function works with document analysis mode. The LLM is instantiated independently of the search engine, allowing you to use local models via Ollama, LM Studio, or remote APIs depending on your configuration, with temperature and other parameters passed directly to control generation behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →