How the Document Tree Retriever Handles Hierarchical Document Structures in OpenDerisk
The document tree retriever constructs a navigable tree representation that mirrors document heading hierarchies by mapping H1-H6 levels to parent-child relationships in TreeNode objects, enabling section-aware retrieval in RAG pipelines.
The document tree retriever within the derisk-ai/openderisk repository transforms flat document chunks into structured hierarchical representations. This component preserves the original organization of markdown documents, allowing downstream LLM components to consume context with proper parent-child relationships intact.
Core Architecture: The TreeNode Data Structure
The foundation of the hierarchy system resides in the TreeNode class defined in packages/derisk-ext/src/derisk_ext/rag/retriever/doc_tree.py (lines 26-35). Each node captures a specific heading level and its associated content.
A TreeNode stores five critical attributes:
- node_id: A unique identifier for the tree node
- title: The heading text extracted from the document
- level: Numeric depth indicator (0 for document title, 1 for H1/
#, 2 for H2/##, continuing through H6) - body_content: Text content belonging to the heading, spanning from the header to the next heading of equal or higher level
- children: A list of nested
TreeNodeobjects representing subsections
This structure enables the retriever to maintain the semantic relationships between document sections, ensuring that nested content remains logically grouped during retrieval operations.
Building Document Hierarchies from Chunked Content
When processing pre-chunked documents, the DocTreeIndex.add_nodes method (lines 100-140 of doc_tree.py) constructs the tree by iterating over document chunks and their metadata.
The algorithm processes standard metadata keys including title, Header1 through Header6 to establish the hierarchy:
- The document title becomes a level-0 root node
Header1metadata creates level-1 children- Successive headers (Header2-Header6) nest as children of their immediate predecessors
- Each node captures the body content associated with its heading level
This approach allows the document tree retriever to reconstruct the original document outline even when the source material has been split into separate Chunk objects for vector storage.
Parsing Raw Markdown Documents
For scenarios requiring direct markdown processing without pre-existing chunks, the DocumentOutlineParser.get_outlines_from_body method (lines 68-98 of doc_tree.py) parses raw markdown strings into identical TreeNode hierarchies.
The parser implements a stack-based algorithm that:
- Skips code blocks to avoid false header detection
- Extracts header lines using regex pattern matching
- Determines heading levels by counting
#symbols (1-6) - Slices interstitial text as
body_contentbetween consecutive headers - Maintains parent-child relationships using a node stack
This functionality enables the document tree retriever to ingest markdown documents directly from source files or user inputs without requiring pre-processing into the chunk format.
Keyword-Based Tree Retrieval
The retrieval mechanism centers on DocTreeRetriever._aretrieve (lines 86-99 of doc_tree.py), which executes depth-first searches across the document forest.
When processing a query:
- The system extracts keywords (optionally via a
keywords_extractorcallable) - For each keyword,
search_keywordstraverses everyDocTreeIndexdepth-first - The matcher compares keywords against node
titleattributes - Upon matching, the retriever returns the complete
TreeNodeincluding its entire subtree
This design ensures that retrieving a parent section automatically includes all nested subsections and their content, preserving the full context for LLM consumption.
Visualizing Hierarchical Structures
For debugging and verification, the retriever provides tree visualization through DocTreeIndex.display_tree (lines 55-80 of doc_tree.py). When show_tree=True during retrieval, the system prints the matched hierarchy with visual connectors (└──, ├──) illustrating the parent-child relationships.
This visualization helps developers verify that the document tree retriever correctly interprets complex nested structures before deploying to production RAG pipelines.
Integration into the RAG Pipeline
The retriever integrates into the broader OpenDerisk system through knowledge_space.py (lines 351-353 of packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py), which imports DocTreeRetriever and feeds document collections to the retrieval engine.
Basic Usage Example
from derisk.core import Document, Chunk
from derisk_ext.rag.retriever.doc_tree import DocTreeRetriever
# Build Document objects with hierarchical metadata
doc = Document(
doc_id="doc1",
chunks=[
Chunk(
chunk_id="c1",
content="Full text of the chapter",
metadata={"title": "Chapter One", "Header1": "Background", "Header2": "Motivation"},
),
Chunk(
chunk_id="c2",
content="Details of the method",
metadata={"title": "Chapter One", "Header1": "Methodology"},
),
],
)
# Initialize the retriever
retriever = DocTreeRetriever(docs=[doc])
# Asynchronously retrieve matching tree nodes
import asyncio
async def demo():
nodes = await retriever._aretrieve("Motivation")
for node in nodes:
print(f"Matched node: {node.title}")
print(f"Content: {node.body_content}")
asyncio.run(demo())
Parsing Raw Markdown Example
from derisk_ext.rag.retriever.doc_tree import DocumentOutlineParser
markdown = """
# Project Overview
Some intro text.
## Data Sources
Description of data.
### API
Details about the API.
## Model Architecture
Explanation of the model.
"""
parser = DocumentOutlineParser()
outlines = parser.get_outlines_from_body(markdown)
# Display visual tree representation
print(parser.display(outlines))
Output:
└── Project Overview (content: Some intro text.)
├── Data Sources (content: Description of data.)
│ └── API (content: Details about the API.)
└── Model Architecture (content: Explanation of the model.)
Summary
- The document tree retriever maps document heading hierarchies (H1-H6) to
TreeNodeparent-child relationships, storing each node's title, level, body content, and children indoc_tree.py. - DocTreeIndex.add_nodes constructs trees from chunked documents using
titleandHeader1-Header6metadata, while DocumentOutlineParser.get_outlines_from_body parses raw markdown directly. - Retrieval executes via DocTreeRetriever._aretrieve, performing depth-first keyword matching against node titles and returning complete subtrees to preserve nested context.
- The system integrates into OpenDerisk RAG pipelines through
knowledge_space.py, with optional tree visualization available viadisplay_tree.
Frequently Asked Questions
How does the document tree retriever represent different heading levels?
The retriever stores heading levels as numeric values in the TreeNode.level attribute, where 0 represents the document title, 1 corresponds to H1 (#), 2 to H2 (##), and continuing through 6 for H6 (######). This numeric mapping enables the tree builder in DocTreeIndex.add_nodes to correctly nest children under their appropriate parents when reconstructing the document hierarchy from chunked content.
Can the retriever process raw markdown strings without pre-chunked documents?
Yes, the DocumentOutlineParser.get_outlines_from_body method (lines 68-98 of doc_tree.py) parses raw markdown strings directly into TreeNode hierarchies. This parser skips code blocks, extracts headers via regex, counts # symbols to determine levels, and uses a stack to maintain parent-child relationships, producing the same tree structure as the chunked document processor.
What happens when a keyword matches a parent node in the hierarchy?
When DocTreeRetriever._aretrieve finds a keyword match in a parent node's title, it returns the complete TreeNode object including its entire subtree of children. This ensures that retrieving a high-level section automatically includes all nested subsections and their body content, preserving the full hierarchical context for downstream LLM processing rather than returning isolated fragments.
How is the document tree retriever integrated into the OpenDerisk RAG pipeline?
According to the source code in packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py (lines 351-353), the retriever is instantiated by importing DocTreeRetriever from the extension package and feeding it document collections. This integration allows the RAG system to leverage hierarchical document structures during knowledge retrieval operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →