# How Document Analysis Mode Searches Private Documents with AI in Local Deep Research

> Discover how Document Analysis Mode uses AI to search private documents. Learn how it bypasses protections to access internal services and generate AI summaries locally.

- Repository: [learningcircuit/local-deep-research](https://github.com/learningcircuit/local-deep-research)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Document analysis mode uses the `analyze_documents()` orchestrator to connect to private document collections by selectively disabling SSRF protections via `allow_private_ips=True`, retrieving content from internal services like Paperless-ngx, and synthesizing AI-generated summaries through a local language model.**

The `learningcircuit/local-deep-research` repository enables secure AI analysis of sensitive documents stored on private networks. Unlike public web search modes, document analysis specifically targets local collections—such as Paperless-ngx instances or Elasticsearch indexes—while maintaining security boundaries through controlled network access and verified file operations.

## The Orchestration Entry Point: analyze_documents()

The document analysis workflow centers on `analyze_documents()` in [`src/local_deep_research/api/research_functions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/api/research_functions.py). This function serves as the primary API entry point, coordinating language model initialization, private search engine instantiation, and secure result persistence. It accepts parameters including `query`, `collection_name`, `max_results`, `temperature`, and optional `output_file` to control the analysis behavior.

## Step-by-Step Private Document Search Flow

### 1. Language Model Initialization

The process begins by creating an LLM instance via `get_llm(temperature=...)`, typically found around lines 65-66. This configured model handles the final summarization step, with temperature controls allowing adjustment of creative versus deterministic output.

### 2. Private Search Engine Selection

Next, `get_search(collection_name, llm_instance=llm)` instantiates a search engine bound to the specified private collection, as implemented around lines 68-70. The `collection_name` parameter directs the engine to communicate with local services rather than public APIs, selecting implementations like **Paperless-ngx** or **Library-RAG** connectors.

### 3. Bypassing SSRF Protections for Internal Networks

Private document engines require access to internal IP ranges (`192.168.*`, `::1`, localhost). Each engine uses `safe_get()`—a security wrapper around `requests.get`—with explicit flags `allow_private_ips=True` and often `allow_localhost=True`. This temporarily disables default SSRF protection for that specific request while maintaining audit logging. The `PaperlessSearchEngine` implementation in [`src/local_deep_research/web_search_engines/engines/search_engine_paperless.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/engines/search_engine_paperless.py) demonstrates this pattern around lines 29-38.

### 4. Executing the Document Retrieval

The `search.run(query)` method dispatches the request to the private service, returning a structured list of dictionaries containing `title`, `link`, and `content`/`snippet` fields (lines 78-82). This raw retrieval happens entirely within the private network boundary.

### 5. Context Window Management

To respect LLM token limits, the system truncates the first **five results** to **1,000 characters** each. These excerpts are concatenated into a single prompt instructing the model to "Analyze these document excerpts … Provide a concise summary" (lines 91-104).

### 6. AI Summarization

The constructed prompt is sent to the LLM via `llm.invoke(summary_prompt)`. The raw response is processed to remove any "think" tags or artifacts, producing the final summary output (lines 13-22).

### 7. Secure Output Handling

Results are returned as a dictionary containing the `summary`, raw `documents` list, and metadata (`collection`, `document_count`). When `output_file` is specified, `write_file_verified()` from [`src/local_deep_research/security/file_write_verifier.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/security/file_write_verifier.py) safely persists the report to disk, enforcing repository security policies (lines 23-30).

## Practical Implementation Examples

Query a Paperless-ngx collection containing company policies:

```python
from local_deep_research.api.research_functions import analyze_documents

result = analyze_documents(
    query="Company policy on remote work",
    collection_name="paperless_my_company",
    max_results=8,
    temperature=0.5,
    output_file="remote_work_summary.md",
)

print(result["summary"])

```

Analyze security guidelines from a local Elasticsearch index:

```python
from local_deep_research.api.research_functions import analyze_documents

summary = analyze_documents(
    query="Latest security guidelines",
    collection_name="my_elastic_index",
    max_results=5,
    temperature=0.7,
).get("summary")

print("AI-generated summary:", summary)

```

## Key Source Files and Architecture

- **[`src/local_deep_research/api/research_functions.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/api/research_functions.py)** – Contains `analyze_documents()`, orchestrating LLM creation, engine selection, search execution, and summarization.
- **[`src/local_deep_research/web_search_engines/search_engine_base.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/search_engine_base.py)** – Defines the `BaseSearchEngine` abstract class establishing the `run()` contract for all engine implementations.
- **[`src/local_deep_research/web_search_engines/engines/search_engine_paperless.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/web_search_engines/engines/search_engine_paperless.py)** – Reference implementation showing `safe_get(..., allow_private_ips=True)` usage for internal services.
- **[`src/local_deep_research/security/ssrf_validator.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/security/ssrf_validator.py)** – Provides the SSRF guard and `allow_private_ips` flag logic for secure private network access.
- **[`src/local_deep_research/config/search_config.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/config/search_config.py)** – Implements `get_search()`, constructing parameter dictionaries and factory dispatch for private engines.
- **[`src/local_deep_research/security/file_write_verifier.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/security/file_write_verifier.py)** – Houses `write_file_verified()`, ensuring optional report output respects security boundaries.

## Summary

- **Document analysis mode** targets private collections via configurable search engines bound to local services like Paperless-ngx or Elasticsearch.
- **SSRF protection is selectively disabled** using `allow_private_ips=True` within `safe_get()` wrappers, enabling access to `192.168.x.x` and localhost while maintaining request logging.
- **The `analyze_documents()` orchestrator** manages the full lifecycle: LLM initialization, private search execution, context truncation (max 5 documents, 1,000 characters each), and AI summarization.
- **Security persists through all stages**, from controlled network access to verified file writes via `write_file_verified()`.

## Frequently Asked Questions

### How does document analysis mode differ from public web search?

Public web search uses engines like DuckDuckGo or Google that respect SSRF protections and query external APIs. Document analysis mode explicitly uses `get_search()` with a `collection_name` parameter to instantiate private engines that target internal services, coupled with `allow_private_ips=True` to reach documents on private subnets or localhost.

### What security measures protect against malicious private IP exploitation?

The `safe_get()` wrapper in [`src/local_deep_research/security/ssrf_validator.py`](https://github.com/learningcircuit/local-deep-research/blob/main/src/local_deep_research/security/ssrf_validator.py) provides granular control. While `allow_private_ips=True` permits access to RFC 1918 ranges, the flag is set explicitly per engine instance and requests are logged for audit. Additionally, `write_file_verified()` ensures that any persisted output undergoes security validation before disk write.

### Can I analyze documents from multiple private collections simultaneously?

The current `analyze_documents()` implementation accepts a single `collection_name` per invocation. To analyze multiple collections, invoke the function separately for each collection and aggregate results programmatically, or extend the orchestrator to accept multiple engine instances.

### Which LLM providers are compatible with document analysis mode?

Any provider supported by the repository's `get_llm()` factory function works with document analysis mode. The LLM is instantiated independently of the search engine, allowing you to use local models via Ollama, LM Studio, or remote APIs depending on your configuration, with `temperature` and other parameters passed directly to control generation behavior.