What is HyDE and How to Implement It in RAG Systems
HyDE (Hypothetical Document Embedding) improves retrieval accuracy by using an LLM to generate a synthetic "hypothetical" answer to the user's query, embedding that synthetic document, and searching the vector store with this richer representation instead of the original short query.
HyDE is a powerful query enhancement technique for retrieval-augmented generation that bridges the semantic gap between terse user questions and lengthy source documents. This guide explains how HyDE works and provides a complete implementation using the NirDiamant/RAG_Techniques open-source repository, demonstrating how to integrate Hypothetical Document Embedding into your RAG pipeline.
Understanding HyDE (Hypothetical Document Embedding)
HyDE addresses a fundamental mismatch in standard RAG systems: user queries are often short and ambiguous, while the indexed documents are long and detailed. Instead of embedding the raw query directly, HyDE first prompts a large language model to hallucinate a perfect answer—a "hypothetical document"—and then uses that synthetic text as the search vector.
Why HyDE Works
| Benefit | Explanation |
|---|---|
| Improved relevance | By expanding a terse query into a detailed, context‑rich passage, the embedding captures more semantic nuances, increasing the chances of retrieving truly relevant chunks. |
| Better alignment | The vector store contains long document chunks; feeding it a similarly long representation reduces the distribution mismatch that typically hurts similarity search. |
| Model‑agnostic | Any embedding model (e.g., OpenAI embeddings) can be reused; only the LLM that creates the hypothetical document needs to be capable of generation. |
HyDE Architecture and Implementation in RAG_Techniques
The NirDiamant/RAG_Techniques repository provides a production-ready implementation in all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py. The core logic is encapsulated in the HyDERetriever class, which orchestrates document ingestion, hypothetical generation, and similarity search.
Core Components in HyDERetriever
According to the source code in all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py (lines 18‑41), the HyDERetriever class initializes:
- An LLM (
self.llm) for generating hypothetical documents. - An embedding model (
self.embedding_model) shared with the vector store. - A PromptTemplate (
self.hyde_prompt) that instructs the LLM to produce a full answer of lengthself.chunk_size. - A FAISS vector store (
self.vectorstore) built viaencode_pdffromhelper_functions.py.
The helper function encode_pdf (defined in helper_functions.py, lines 48‑57) handles PDF loading, text splitting into overlapping chunks, and embedding into the FAISS index.
The Four-Step Retrieval Flow
When you call retrieve(query, k), the HyDERetriever executes this pipeline:
- Document Ingestion –
encode_pdfprocesses the source PDF into chunked embeddings stored in FAISS. - Query Expansion –
generate_hypothetical_documentfeeds the user query intoself.hyde_prompt, prompting the LLM to write a synthetic answer. - Embedding the Synthetic Answer – The generated text is automatically embedded using
self.embedding_model. - Similarity Search – The embedding queries
self.vectorstore.similarity_searchto return the k most relevant chunks.
The retrieved chunks, along with the hypothetical document, can then be passed to downstream answer-generation chains such as create_question_answer_from_context_chain defined in helper_functions.py.
Practical Implementation: Code Examples
Minimal HyDE RAG Script
The following runnable script demonstrates the complete HyDE workflow using the repository's implementation:
# HyDE RAG example – run as a script
# File: all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py
# (see full source: https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py)
from helper_functions import *
from pathlib import Path
# ----------------------------------------------------------------------
# 1️⃣ Build the retriever (PDF → vector store)
# ----------------------------------------------------------------------
pdf_path = Path("../data/Understanding_Climate_Change.pdf")
retriever = HyDERetriever(str(pdf_path))
# ----------------------------------------------------------------------
# 2️⃣ Expand the query to a hypothetical document
# ----------------------------------------------------------------------
query = "What are the main causes of climate change?"
hypothetical_doc, results = retriever.retrieve(query, k=3)
# ----------------------------------------------------------------------
# 3️⃣ Show the synthetic answer and the retrieved chunks
# ----------------------------------------------------------------------
print("\n=== Hypothetical Document ===\n")
print(text_wrap(hypothetical_doc))
print("\n=== Retrieved Chunks ===\n")
for i, doc in enumerate(results, 1):
print(f"--- Chunk {i} ---")
print(text_wrap(doc.page_content))
print()
Integrating HyDE with Answer Generation Chains
To complete the RAG pipeline, combine HyDE retrieval with a question-answering chain using create_question_answer_from_context_chain from helper_functions.py:
# Integrate HyDE with a context‑answering chain
# (uses create_question_answer_from_context_chain from helper_functions.py)
from helper_functions import *
from langchain_openai import ChatOpenAI
# Build the HyDE retriever
retriever = HyDERetriever("../data/Understanding_Climate_Change.pdf")
# Generate hypothesis and retrieve chunks
query = "How does deforestation affect carbon cycles?"
hypo_doc, docs = retriever.retrieve(query, k=4)
# Prepare context (concatenate retrieved chunks)
context = " ".join([d.page_content for d in docs])
# Set up LLM for answer generation
llm = ChatOpenAI(model_name="gpt-4o-mini", temperature=0)
# Build the Q&A chain (source: helper_functions.py lines 62‑84)
qa_chain = create_question_answer_from_context_chain(llm)
# Get the final answer
answer = qa_chain.invoke({"context": context, "question": query})
print("\n**Answer:**", answer.answer_based_on_content)
Interactive Jupyter Notebook
For a visual walkthrough of the synthetic document generation and retrieval results, explore the interactive notebook at all_rag_techniques/HyDe_Hypothetical_Document_Embedding.ipynb. The notebook illustrates the same pipeline with plots comparing the hypothetical document against retrieved chunks.
Key Files and Resources
| File | Role | Link |
|---|---|---|
all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py |
Full HyDE implementation – retriever class, prompt, and CLI entry point | GitHub |
helper_functions.py |
Utility functions for PDF loading, chunking, embedding, and QA chain (used by HyDE) | GitHub |
all_rag_techniques/HyDe_Hypothetical_Document_Embedding.ipynb |
Notebook illustration with plots of the synthetic document and retrieved chunks | GitHub |
README.md (section “Query Enhancement”) |
High‑level description of HyDE within the catalogue of RAG techniques | GitHub |
evaluation/evalute_rag.py |
Helper for evaluating retrieval performance; can be used to benchmark HyDE vs. baseline | GitHub |
Summary
- HyDE bridges the semantic gap between short user queries and long document chunks by generating a synthetic "hypothetical" answer before retrieval.
- Implementation resides in
all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py, specifically within theHyDERetrieverclass (lines 18‑41). - The pipeline uses
helper_functions.py(lines 48‑57) for PDF ingestion viaencode_pdfand supports downstream answer generation throughcreate_question_answer_from_context_chain. - HyDE is model-agnostic regarding embeddings—any vector model works—but requires an LLM capable of generating coherent hypothetical documents.
- Evaluation can be performed using
evaluation/evalute_rag.pyto measure retrieval improvements over baseline vector search.
Frequently Asked Questions
What is the difference between HyDE and standard RAG retrieval?
Standard RAG embeds the raw user query and searches the vector store directly, which often fails when queries are short or ambiguous. HyDE first prompts an LLM to generate a detailed hypothetical document that answers the query, then embeds this synthetic text. Because the hypothetical document matches the length and style of the indexed chunks, similarity search returns more relevant results.
Does HyDE work with any embedding model?
Yes, HyDE is embedding-model agnostic. The HyDERetriever class in all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py accepts any LangChain-compatible embedding model (e.g., OpenAI, HuggingFace, Cohere) via the embedding_model parameter. The only requirement is that the same model must be used for both indexing the corpus and embedding the hypothetical document to ensure vector space consistency.
How does the hypothetical document length affect retrieval quality?
The hypothetical document length is controlled by the chunk_size parameter passed to HyDERetriever. According to the implementation, the prompt template (self.hyde_prompt) instructs the LLM to generate a response of approximately this length. If the synthetic document is too short, it may lack the semantic richness needed to match detailed source chunks; if too long, it may introduce irrelevant noise. The repository defaults typically align chunk_size with the chunking strategy used during document ingestion (e.g., 512-1024 tokens) to maximize alignment between the hypothetical and real document distributions.
Can HyDE be combined with other query enhancement techniques?
Absolutely. HyDE can be chained with techniques like query rewriting, expansion, or multi-query generation. For example, you could first apply a rewriter to clarify the user intent, then feed the refined query into HyDERetriever.generate_hypothetical_document, and finally combine HyDE retrieval with keyword-based BM25 retrieval for hybrid search. The modular design of the HyDERetriever class in all_rag_techniques_runnable_scripts/HyDe_Hypothetical_Document_Embedding.py makes it straightforward to insert into broader RAG orchestration frameworks like LangChain or LlamaIndex.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →