# How the Document Tree Retriever Handles Hierarchical Document Structures in OpenDerisk

> Discover how the document tree retriever in OpenRisk maps H1-H6 levels to TreeNode objects for section-aware RAG retrieval by building a navigable tree from document hierarchies.

- Repository: [derisk-ai/openderisk](https://github.com/derisk-ai/openderisk)
- Tags: internals
- Published: 2026-02-28

---

**The document tree retriever constructs a navigable tree representation that mirrors document heading hierarchies by mapping H1-H6 levels to parent-child relationships in `TreeNode` objects, enabling section-aware retrieval in RAG pipelines.**

The document tree retriever within the derisk-ai/openderisk repository transforms flat document chunks into structured hierarchical representations. This component preserves the original organization of markdown documents, allowing downstream LLM components to consume context with proper parent-child relationships intact.

## Core Architecture: The TreeNode Data Structure

The foundation of the hierarchy system resides in the `TreeNode` class defined in [`packages/derisk-ext/src/derisk_ext/rag/retriever/doc_tree.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-ext/src/derisk_ext/rag/retriever/doc_tree.py) (lines 26-35). Each node captures a specific heading level and its associated content.

A `TreeNode` stores five critical attributes:

- **node_id**: A unique identifier for the tree node
- **title**: The heading text extracted from the document
- **level**: Numeric depth indicator (0 for document title, 1 for H1/`#`, 2 for H2/`##`, continuing through H6)
- **body_content**: Text content belonging to the heading, spanning from the header to the next heading of equal or higher level
- **children**: A list of nested `TreeNode` objects representing subsections

This structure enables the retriever to maintain the semantic relationships between document sections, ensuring that nested content remains logically grouped during retrieval operations.

## Building Document Hierarchies from Chunked Content

When processing pre-chunked documents, the `DocTreeIndex.add_nodes` method (lines 100-140 of [`doc_tree.py`](https://github.com/derisk-ai/openderisk/blob/main/doc_tree.py)) constructs the tree by iterating over document chunks and their metadata.

The algorithm processes standard metadata keys including `title`, `Header1` through `Header6` to establish the hierarchy:

1. The document title becomes a level-0 root node
2. `Header1` metadata creates level-1 children
3. Successive headers (Header2-Header6) nest as children of their immediate predecessors
4. Each node captures the body content associated with its heading level

This approach allows the document tree retriever to reconstruct the original document outline even when the source material has been split into separate `Chunk` objects for vector storage.

## Parsing Raw Markdown Documents

For scenarios requiring direct markdown processing without pre-existing chunks, the `DocumentOutlineParser.get_outlines_from_body` method (lines 68-98 of [`doc_tree.py`](https://github.com/derisk-ai/openderisk/blob/main/doc_tree.py)) parses raw markdown strings into identical `TreeNode` hierarchies.

The parser implements a stack-based algorithm that:

- Skips code blocks to avoid false header detection
- Extracts header lines using regex pattern matching
- Determines heading levels by counting `#` symbols (1-6)
- Slices interstitial text as `body_content` between consecutive headers
- Maintains parent-child relationships using a node stack

This functionality enables the document tree retriever to ingest markdown documents directly from source files or user inputs without requiring pre-processing into the chunk format.

## Keyword-Based Tree Retrieval

The retrieval mechanism centers on `DocTreeRetriever._aretrieve` (lines 86-99 of [`doc_tree.py`](https://github.com/derisk-ai/openderisk/blob/main/doc_tree.py)), which executes depth-first searches across the document forest.

When processing a query:

1. The system extracts keywords (optionally via a `keywords_extractor` callable)
2. For each keyword, `search_keywords` traverses every `DocTreeIndex` depth-first
3. The matcher compares keywords against node `title` attributes
4. Upon matching, the retriever returns the complete `TreeNode` including its entire subtree

This design ensures that retrieving a parent section automatically includes all nested subsections and their content, preserving the full context for LLM consumption.

## Visualizing Hierarchical Structures

For debugging and verification, the retriever provides tree visualization through `DocTreeIndex.display_tree` (lines 55-80 of [`doc_tree.py`](https://github.com/derisk-ai/openderisk/blob/main/doc_tree.py)). When `show_tree=True` during retrieval, the system prints the matched hierarchy with visual connectors (└──, ├──) illustrating the parent-child relationships.

This visualization helps developers verify that the document tree retriever correctly interprets complex nested structures before deploying to production RAG pipelines.

## Integration into the RAG Pipeline

The retriever integrates into the broader OpenDerisk system through [`knowledge_space.py`](https://github.com/derisk-ai/openderisk/blob/main/knowledge_space.py) (lines 351-353 of [`packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py)), which imports `DocTreeRetriever` and feeds document collections to the retrieval engine.

### Basic Usage Example

```python
from derisk.core import Document, Chunk
from derisk_ext.rag.retriever.doc_tree import DocTreeRetriever

# Build Document objects with hierarchical metadata

doc = Document(
    doc_id="doc1",
    chunks=[
        Chunk(
            chunk_id="c1",
            content="Full text of the chapter",
            metadata={"title": "Chapter One", "Header1": "Background", "Header2": "Motivation"},
        ),
        Chunk(
            chunk_id="c2",
            content="Details of the method",
            metadata={"title": "Chapter One", "Header1": "Methodology"},
        ),
    ],
)

# Initialize the retriever

retriever = DocTreeRetriever(docs=[doc])

# Asynchronously retrieve matching tree nodes

import asyncio

async def demo():
    nodes = await retriever._aretrieve("Motivation")
    for node in nodes:
        print(f"Matched node: {node.title}")
        print(f"Content: {node.body_content}")

asyncio.run(demo())

```

### Parsing Raw Markdown Example

```python
from derisk_ext.rag.retriever.doc_tree import DocumentOutlineParser

markdown = """

# Project Overview

Some intro text.

## Data Sources

Description of data.

### API

Details about the API.

## Model Architecture

Explanation of the model.
"""

parser = DocumentOutlineParser()
outlines = parser.get_outlines_from_body(markdown)

# Display visual tree representation

print(parser.display(outlines))

```

Output:

```

└── Project Overview (content: Some intro text.)
    ├── Data Sources (content: Description of data.)
    │   └── API (content: Details about the API.)
    └── Model Architecture (content: Explanation of the model.)

```

## Summary

- The **document tree retriever** maps document heading hierarchies (H1-H6) to `TreeNode` parent-child relationships, storing each node's title, level, body content, and children in [`doc_tree.py`](https://github.com/derisk-ai/openderisk/blob/main/doc_tree.py).
- **DocTreeIndex.add_nodes** constructs trees from chunked documents using `title` and `Header1`-`Header6` metadata, while **DocumentOutlineParser.get_outlines_from_body** parses raw markdown directly.
- Retrieval executes via **DocTreeRetriever._aretrieve**, performing depth-first keyword matching against node titles and returning complete subtrees to preserve nested context.
- The system integrates into OpenDerisk RAG pipelines through [`knowledge_space.py`](https://github.com/derisk-ai/openderisk/blob/main/knowledge_space.py), with optional tree visualization available via `display_tree`.

## Frequently Asked Questions

### How does the document tree retriever represent different heading levels?

The retriever stores heading levels as numeric values in the `TreeNode.level` attribute, where 0 represents the document title, 1 corresponds to H1 (`#`), 2 to H2 (`##`), and continuing through 6 for H6 (`######`). This numeric mapping enables the tree builder in `DocTreeIndex.add_nodes` to correctly nest children under their appropriate parents when reconstructing the document hierarchy from chunked content.

### Can the retriever process raw markdown strings without pre-chunked documents?

Yes, the `DocumentOutlineParser.get_outlines_from_body` method (lines 68-98 of [`doc_tree.py`](https://github.com/derisk-ai/openderisk/blob/main/doc_tree.py)) parses raw markdown strings directly into `TreeNode` hierarchies. This parser skips code blocks, extracts headers via regex, counts `#` symbols to determine levels, and uses a stack to maintain parent-child relationships, producing the same tree structure as the chunked document processor.

### What happens when a keyword matches a parent node in the hierarchy?

When `DocTreeRetriever._aretrieve` finds a keyword match in a parent node's title, it returns the complete `TreeNode` object including its entire subtree of children. This ensures that retrieving a high-level section automatically includes all nested subsections and their body content, preserving the full hierarchical context for downstream LLM processing rather than returning isolated fragments.

### How is the document tree retriever integrated into the OpenDerisk RAG pipeline?

According to the source code in [`packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py) (lines 351-353), the retriever is instantiated by importing `DocTreeRetriever` from the extension package and feeding it document collections. This integration allows the RAG system to leverage hierarchical document structures during knowledge retrieval operations.