Advanced RAG Techniques for Improved Retrieval: A Comprehensive Implementation Guide

Advanced RAG techniques for improved retrieval combine hybrid fusion search, intelligent cross-encoder reranking, hierarchical indexing, and synthetic query enhancement to overcome the limitations of basic vector similarity search.

The NirDiamant/RAG_Techniques repository on GitHub provides a production-ready collection of retrieval strategies that extend far beyond simple vector search pipelines. These advanced RAG techniques for improved retrieval address critical challenges including semantic drift, context fragmentation, and domain shift through modular, composable components that can be mixed to match specific latency and accuracy requirements.

Hybrid Fusion and Ensemble Retrieval

Basic dense retrieval often misses exact lexical matches, while keyword search lacks semantic understanding. The repository solves this through Fusion Retrieval and Ensemble Retrieval strategies that aggregate multiple signal types.

Fusion Retrieval (BM25 + Dense Vectors)

Implemented in all_rag_techniques/fusion_retrieval.ipynb, this technique merges keyword-based BM25 results with dense vector matches using a weighted scoring function. The architecture runs both retrievers in parallel, concatenates the result lists, applies a dynamic weighting factor α to each score, and passes the fused list to a downstream reranker for final ordering.

Ensemble Retrieval

As documented in the README section on ensemble retrieval, this approach runs several independent retrievers (BM25, DPR, ColPali image retriever) in parallel and aggregates outputs via voting or weighted averaging. Each retriever produces a normalized score vector, summed with per-retriever weights that can be learned or hand-tuned for domain-specific robustness.

Intelligent Reranking and Multi-Faceted Filtering

Retrieved candidates often contain noise and redundancy. The repository implements sophisticated post-retrieval processing to maximize relevance and diversity before context is sent to the generator.

Cross-Encoder Reranking

The all_rag_techniques/reranking.ipynb notebook demonstrates Intelligent Reranking using cross-encoders or LLM-based scorers. Unlike bi-encoders, cross-encoders jointly encode the query and candidate, producing richer relevance signals. The implementation optionally blends LLM-generated relevance scores with metadata weights (date, source, confidence) for context-aware ordering.

Multi-Faceted Filtering

Before final scoring, the system applies a cascade of filters including metadata constraints (e.g., doc_date > 2020), similarity thresholds (score > 0.7), and diversity heuristics (max-Jaccard similarity < 0.5). These cheap linear passes dramatically reduce noisy results without expensive LLM calls.

Hierarchical and Adaptive Indexing

For large corpora, flat indexing creates excessive search latency and retrieval of irrelevant contexts. These techniques structure the index to enable fast, targeted retrieval.

Hierarchical Indices

The all_rag_techniques/hierarchical_indices.ipynb implementation builds a two-level index: a coarse-grained summary index for fast lookup and a fine-grained chunk index for detailed retrieval. At query time, the system first retrieves the top-N summaries, then drills down to child chunks only from those summaries, reducing the search space by approximately 90%.

Adaptive Retrieval

The all_rag_techniques/adaptive_retrieval.ipynb notebook implements query classification that dynamically selects the optimal pipeline. Incoming queries are classified by type (factoid, reasoning, code, multi-modal), routing to semantic search, BM25, hybrid fusion, or image retrieval as appropriate.

Query Enhancement Strategies

Query-document mismatch represents a major failure mode in RAG systems. The repository provides two complementary synthetic embedding approaches that bridge the semantic gap between how users ask questions and how information is stored.

HyDE (Hypothetical Document Embedding)

Implemented in all_rag_techniques/HyDe_Hypothetical_Document_Embedding.ipynb, HyDE generates synthetic "hypothetical" answers to the query, embeds these answers, and matches against them rather than raw documents. The pipeline follows: query → LLM → synthetic doc → embed → compare, improving recall on sparse corpora where terminology differs between queries and source text.

HyPE (Hypothetical Prompt Embedding)

The all_rag_techniques/HyPE_Hypothetical_Prompt_Embedding.ipynb technique moves synthetic generation to index time. Each chunk is paired with multiple hypothetical prompts, turning retrieval into a question-question matching problem. This eliminates runtime LLM calls, providing a 2-3× speed-up with comparable or higher precision.

Context Enrichment and Semantic Chunking

How documents are divided and annotated significantly impacts retrieval quality. These techniques optimize the granularity and metadata of indexed chunks.

Semantic Chunking and Windows

The all_rag_techniques/semantic_chunking.ipynb notebook uses language models to split documents at natural topic boundaries rather than fixed token limits. Complementing this, the all_rag_techniques/context_enrichment_window_around_chunk.ipynb implementation retrieves configurable windows of surrounding sentences around matched content, preserving local context without increasing embedding count.

Contextual Chunk Headers

The all_rag_techniques/contextual_chunk_headers.ipynb approach prepends informative headers (document title, section name) to each chunk before embedding. This injects higher-level semantic cues during indexing: header + chunk → embed.

Relevant Segment Extraction (RSE)

After initial retrieval, all_rag_techniques/relevant_segment_extraction.ipynb expands chosen chunks into the longest contiguous segment remaining relevant, providing richer generator context through a lightweight classifier over neighboring chunks.

Advanced Scoring and Iterative Retrieval

These techniques implement feedback mechanisms and sophisticated scoring functions that go beyond simple similarity metrics.

Dartboard Retrieval

The all_rag_techniques/dartboard.ipynb implementation optimizes for both relevance and information-gain via a scoring function: Score = λ·relevance + (1-λ)·gain. Gain estimates the marginal improvement in LLM answer quality when adding the candidate, particularly effective for dense corpora where simple relevance plateaus.

Retrieval with Feedback Loop

Implemented in all_rag_techniques/retrieval_with_feedback_loop.ipynb, this technique extracts uncertainty signals from initial LLM answers (missing citations, low confidence) and automatically issues follow-up sub-queries to fill gaps. The self-critique module generates clarifying questions, re-runs retrieval, and regenerates the final answer.

Multi-Modal Retrieval Systems

Extending RAG beyond text requires handling visual and document formats natively.

Caption-Based and ColPali Retrieval

For non-textual media, all_rag_techniques/multi_model_rag_with_captioning.ipynb transforms PDFs, PPTs, and images into captions using vision-LLMs, storing these as text chunks. Alternatively, all_rag_techniques/multi_model_rag_with_colpali.ipynb indexes raw images using the ColPali vision encoder for direct image-to-image similarity matching, feeding top images to a vision-LLM for synthesis.

Implementation Examples

The following examples demonstrate composing these advanced RAG techniques for improved retrieval using LangChain patterns from the repository.

Fusion Retrieval with HyPE and Reranking

from langchain.vectorstores import FAISS
from langchain.embeddings import OpenAIEmbeddings
from langchain.retrievers import BM25Retriever
from langchain.llms import OpenAI
from langchain.chains import RetrievalQA
from pathlib import Path
import json, os

# ---------- 1️⃣ Load embeddings ----------

embed = OpenAIEmbeddings(model="text-embedding-ada-002")

# ---------- 2️⃣ Build two stores ----------

# Dense vector store (Faiss)

faiss = FAISS.from_documents(
    docs,  # `docs` = list of Document objects (loaded elsewhere)

    embedding=embed,
)

# BM25 lexical store (simple in‑memory)

bm25 = BM25Retriever.from_documents(docs)

# ---------- 3️⃣ Fusion layer ----------

def fuse(query, k=10, alpha=0.6):
    # dense scores

    dense_res = faiss.similarity_search_with_score(query, k=k)
    # bm25 scores

    bm25_res = bm25.get_relevant_documents(query)[:k]

    # simple linear interpolation of scores

    fused = []
    for d, s in dense_res:
        bm25_score = next((1.0 for doc in bm25_res if doc.page_content == d.page_content), 0.0)
        fused_score = alpha * s + (1 - alpha) * bm25_score
        fused.append((d, fused_score))
    fused.sort(key=lambda x: x[1], reverse=True)
    return [doc for doc, _ in fused[:k]]

# ---------- 4️⃣ HyPE index (pre‑computed prompts) ----------

# Assume `hype_index.json` maps chunk_id → list of hypothetical prompts

with open(Path("hype_index.json")) as f:
    hype_index = json.load(f)

def hype_search(query, k=5):
    # embed the query

    q_vec = embed.embed_query(query)
    # compare against all stored hypothetical prompts (vectorized offline)

    scores = []
    for chunk_id, prompts in hype_index.items():
        for p in prompts:
            score = cosine_similarity(q_vec, p["embedding"])
            scores.append((chunk_id, score))
    scores.sort(key=lambda x: x[1], reverse=True)
    top_ids = [cid for cid, _ in scores[:k]]
    return [faiss.docstore.search(cid) for cid in top_ids]

# ---------- 5️⃣ Reranker (cross‑encoder) ----------

cross_encoder = OpenAI(model_name="gpt-4")  # placeholder; replace with actual cross‑encoder

def rerank(docs, query):
    scored = []
    for doc in docs:
        prompt = f"Query: {query}\nDocument: {doc.page_content}\nRate relevance 0‑10:"
        score = int(cross_encoder(prompt).strip())
        scored.append((doc, score))
    scored.sort(key=lambda x: x[1], reverse=True)
    return [doc for doc, _ in scored][:5]

# ---------- 6️⃣ Build the QA chain ----------

retriever = lambda q: rerank(fuse(q) + hype_search(q), q)
qa = RetrievalQA.from_chain_type(
    llm=OpenAI(),
    retriever=retriever,
    return_source_documents=True,
)

# ---------- 7️⃣ Ask a question ----------

answer = qa.run("What are the latest trends in vector search?")
print(answer)

Hierarchical Index with Adaptive Routing

from langchain.document_loaders import TextLoader
from langchain.vectorstores import FAISS
from langchain.embeddings import OpenAIEmbeddings
from sklearn.metrics.pairwise import cosine_similarity

# --- Load documents ---

loader = TextLoader("data/nike_2023_annual_report.txt")
docs = loader.load()

# --- Build summary (coarse) embeddings ---

summary_texts = [doc.page_content[:200] + "…" for doc in docs]
summary_embeds = OpenAIEmbeddings().embed_documents(summary_texts)
summary_index = FAISS.from_embeddings(summary_embeds, summary_texts)

# --- Build fine‑grained chunk embeddings (store under summary id) ---

chunk_store = {}
for i, doc in enumerate(docs):
    chunks = split_into_chunks(doc.page_content)  # helper: semantic chunker

    chunk_embeds = OpenAIEmbeddings().embed_documents(chunks)
    chunk_store[i] = {"chunks": chunks, "embeds": chunk_embeds}

def hierarchical_retrieve(query, top_summaries=3, top_chunks=5):
    # 1️⃣ Retrieve best summaries

    sum_res = summary_index.similarity_search_with_score(query, k=top_summaries)
    results = []
    for summary_doc, _ in sum_res:
        sid = summary_doc.metadata["index"]
        # 2️⃣ Retrieve chunks within that summary

        chunk_vecs = chunk_store[sid]["embeds"]
        q_vec = OpenAIEmbeddings().embed_query(query)
        sims = cosine_similarity([q_vec], chunk_vecs)[0]
        top_idxs = sims.argsort()[-top_chunks:][::-1]
        for idx in top_idxs:
            results.append(chunk_store[sid]["chunks"][idx])
    return results

# Adaptive router – choose hierarchical for long queries

def adaptive_router(query):
    if len(query.split()) > 8:
        return hierarchical_retrieve(query)
    else:
        return FAISS.from_documents(docs).similarity_search(query, k=5)

# Example usage

for chunk in adaptive_router("Explain Nike's sustainability initiatives in 2023"):
    print("- ", chunk[:150], "...\n")

Self-Correcting Retrieval with Feedback

from langchain.chains import LLMChain
from langchain.prompts import PromptTemplate

# Basic retriever (could be any of the above)

retriever = ...

# LLM for answering

answer_llm = OpenAI(model_name="gpt-4")

# Prompt that asks the LLM to self‑criticise

critique_template = PromptTemplate(
    input_variables=["question", "answer"],
    template="""
You are given a question and an answer generated from retrieved documents.
Identify any missing facts, hallucinations, or unclear statements.
If you find an issue, output a short follow‑up sub‑question that would retrieve the missing information.
If the answer looks complete, output ONLY "OK".
Question: {question}
Answer: {answer}
""",
)

critique_chain = LLMChain(llm=answer_llm, prompt=critique_template)

def rag_with_feedback(query):
    context = retriever.get_relevant_documents(query)
    answer = answer_llm.run(f"Answer the question using this context:\n{context}")
    feedback = critique_chain.run(question=query, answer=answer).strip()
    if feedback != "OK":
        # treat feedback as a new query and retrieve again

        extra_context = retriever.get_relevant_documents(feedback)
        answer = answer_llm.run(
            f"Refine the previous answer using these additional documents:\n{extra_context}"
        )
    return answer

print(rag_with_feedback("What were Nike's revenue streams in FY2023?"))

Summary

  • Hybrid Fusion Retrieval (all_rag_techniques/fusion_retrieval.ipynb) combines BM25 lexical matching with dense vector search using weighted score interpolation to capture both exact matches and semantic similarity.
  • Intelligent Reranking leverages cross-encoders and metadata blending in all_rag_techniques/reranking.ipynb to refine candidate ordering beyond initial retrieval scores.
  • Hierarchical Indices implement two-level summary-to-chunk indexing in all_rag_techniques/hierarchical_indices.ipynb that reduces search space by ~90% for large corpora.
  • HyPE and HyDE provide query enhancement through synthetic document or prompt embeddings, with HyPE offering 2-3× latency improvements via index-time generation.
  • Feedback Loop Retrieval enables self-correcting RAG systems in all_rag_techniques/retrieval_with_feedback_loop.ipynb that identify information gaps and automatically issue follow-up queries.
  • Multi-Modal Retrieval extends RAG to images and documents through all_rag_techniques/multi_model_rag_with_captioning.ipynb and all_rag_techniques/multi_model_rag_with_colpali.ipynb using caption generation or ColPali vision encoders.

Frequently Asked Questions

What is the difference between Fusion Retrieval and Ensemble Retrieval?

Fusion Retrieval specifically combines dense vector and BM25 lexical results through a weighted scoring function (alpha blending), typically using the same document corpus. Ensemble Retrieval runs multiple independent retrievers (which may include different modalities like image or code search) in parallel and aggregates their outputs through voting or weighted averaging, making it more robust to domain shift across heterogeneous data sources.

How does HyPE improve upon HyDE for production RAG systems?

HyPE (Hypothetical Prompt Embedding) moves the expensive LLM generation step to index time, creating multiple hypothetical prompts per chunk that are embedded and stored. At query time, only a dense similarity search is required, eliminating runtime LLM latency. This provides a 2-3× speed-up compared to HyDE, which generates synthetic documents at query time, while maintaining comparable or higher precision.

When should I use Hierarchical Indices versus standard flat indexing?

Use Hierarchical Indices (all_rag_techniques/hierarchical_indices.ipynb) when working with large document corpora where flat search becomes computationally expensive or returns noisy results. The two-level approach first filters at the document/summary level, then searches only within relevant child chunks, reducing latency and improving precision. For small collections (< 10k chunks), standard flat indexing with reranking is typically sufficient.

What file contains the implementation for multi-modal image retrieval?

The repository provides two approaches: all_rag_techniques/multi_model_rag_with_captioning.ipynb converts images to text captions using vision-LLMs, while all_rag_techniques/multi_model_rag_with_colpali.ipynb implements direct image-to-image similarity using the ColPali vision encoder. The ColPali approach is preferred when visual layout and structure carry semantic meaning beyond what captions capture.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →