How to Use Relevant Segment Extraction in RAG: A Complete Implementation Guide
Relevant Segment Extraction (RSE) is a post-retrieval technique that converts scattered relevant chunks into contiguous ordered segments using a maximum-sum sub-array algorithm, improving context coherence for large language models.
Relevant Segment Extraction solves the fragmentation problem in Retrieval-Augmented Generation systems where vector search returns isolated chunks that lose narrative flow. This guide explains the complete implementation of RSE as found in the NirDiamant/RAG_Techniques repository, transforming disjointed retrieval results into coherent text segments ready for LLM consumption.
How RSE Works: The Architecture Pipeline
RSE operates as a post-retrieval processing layer that sits between your vector database and the LLM prompt. According to the implementation in all_rag_techniques/relevant_segment_extraction.ipynb, the pipeline follows six distinct stages to reconstruct meaningful context from retrieved fragments.
Chunk Store and Initial Retrieval
The foundation of RSE is a key-value chunk store where each document is split into non-overlapping chunks and indexed by doc_id combined with chunk_index. This structure enables fast random access when rebuilding contiguous segments. The initial retrieval phase uses vector similarity search (optionally boosted by keyword search) to fetch candidate chunks, which may then pass through a cross-encoder reranker such as Cohere Rerank (rerank-english-v3.0).
Scoring Fusion and Decay
For every retrieved chunk, the system computes a combined relevance value through two transformations:
- Absolute relevance: Raw similarity scores undergo a β-CDF transformation (
a=b=0.4) to spread values uniformly across the 0-1 range - Decayed score:
absolute_relevance × exp(-rank/decay_rate)wheredecay_ratedefaults to 30, ensuring higher-ranked chunks contribute more weight to the final segment value
Maximum-Sum Sub-Array Search
Each chunk receives its decayed score as its chunk value. The algorithm then treats the document as an array of these values and executes a brute-force maximum-sum sub-array search (with early-stop heuristics) to locate the contiguous range with the highest total value. This operation completes in approximately 5-10ms per document, making it suitable for real-time RAG pipelines.
Sandwich Expansion
When the optimal segment skips over low-relevance chunks, RSE optionally includes the intervening "sandwich" chunks to preserve narrative continuity. This ensures the LLM receives full context even when relevance scores dip temporarily between highly relevant passages.
Implementing RSE in Python
The reference implementation in NirDiamant/RAG_Techniques provides a complete, runnable pipeline. Below is the essential code structure adapted from all_rag_techniques/relevant_segment_extraction.ipynb.
Prerequisites and Setup
!pip install langchain-cohere python-dotenv scipy matplotlib
import os
import numpy as np
from typing import List
from scipy.stats import beta
import cohere
from dotenv import load_dotenv
from langchain_text_splitters import RecursiveCharacterTextSplitter
load_dotenv()
os.environ["CO_API_KEY"] = os.getenv("CO_API_KEY")
Document Chunking Strategy
The split_into_chunks function creates the non-overlapping storage units required by the chunk store. It uses LangChain's RecursiveCharacterTextSplitter with a default size of 800 tokens and zero overlap to ensure clean segment boundaries.
def split_into_chunks(text: str, chunk_size: int = 800) -> List[str]:
splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size,
chunk_overlap=0,
length_function=len
)
docs = splitter.create_documents([text])
return [doc.page_content for doc in docs]
Relevance Scoring with Beta-CDF Transformation
The transform function applies the β-CDF spread to raw reranker scores, while rerank_chunks handles the API call and exponential decay application:
def transform(x: float) -> float:
a, b = 0.4, 0.4 # spreads scores uniformly across 0-1
return beta.cdf(x, a, b)
def rerank_chunks(query: str, chunks: List[str]):
client = cohere.Client(api_key=os.environ["CO_API_KEY"])
decay_rate = 30
resp = client.rerank(
model="rerank-english-v3.0",
query=query,
documents=chunks
)
# Restore original order and compute decayed values
similarity = [0] * len(chunks)
values = [0] * len(chunks)
for rank, result in enumerate(resp.results):
idx = result.index
raw = result.relevance_score
transformed = transform(raw)
similarity[idx] = transformed
values[idx] = np.exp(-rank/decay_rate) * transformed
return similarity, values
Extracting the Optimal Segment
The core RSE algorithm implements a brute-force search to find the contiguous segment with maximum summed value:
def find_best_segment(chunk_values: List[float]):
"""Brute-force max-sum sub-array with early break heuristics."""
best_sum, best_start, best_end = -float("inf"), 0, 0
n = len(chunk_values)
for i in range(n):
cur = 0
for j in range(i, n):
cur += chunk_values[j]
if cur > best_sum:
best_sum, best_start, best_end = cur, i, j
return best_start, best_end, best_sum
# Usage example
document = """Your long document text here..."""
query = "What is the main contribution of the paper?"
chunks = split_into_chunks(document, chunk_size=800)
_, chunk_vals = rerank_chunks(query, chunks)
start, end, score = find_best_segment(chunk_vals)
relevant_segment = " ".join(chunks[start : end + 1])
print(f"Segment [{start}-{end}] (score {score:.2f}):\n{relevant_segment[:1000]}...")
Key Functions and Source Files
The RSE implementation spans multiple files in the repository:
all_rag_techniques/relevant_segment_extraction.ipynb: Contains the full pipeline includingplot_relevance_scoresfor visualization, scoring logic, and evaluation metricshelper_functions.py: Provides shared utilities for data loading and plotting functions used across the RSE notebook- Scoring functions:
transform()applies the Beta-CDF witha=b=0.4, while the decay calculation usesdecay_rate=30to weight rank position
The notebook also includes visualization tools to debug relevance distributions across chunk sequences before segment extraction.
Summary
- Relevant Segment Extraction converts isolated retrieved chunks into ordered, contiguous text segments that preserve document flow for LLM prompting
- The algorithm uses β-CDF transformation (
a=b=0.4) and exponential decay (decay_rate=30) to weight chunks by both absolute relevance and retrieval rank - A maximum-sum sub-array search identifies the optimal contiguous segment in 5-10ms per document, enabling real-time RAG applications
- Sandwich expansion optionally includes low-relevance intervening chunks to maintain narrative coherence when context temporarily dips between high-value passages
- The complete implementation resides in
NirDiamant/RAG_Techniques, specificallyall_rag_techniques/relevant_segment_extraction.ipynb
Frequently Asked Questions
How does Relevant Segment Extraction differ from standard chunk retrieval?
Standard retrieval returns individual chunks ranked by similarity, often breaking narrative flow when relevant passages span multiple chunks. RSE reconstructs contiguous multi-chunk segments that preserve original document order, ensuring the LLM receives coherent context rather than scattered excerpts. This addresses the "lost in the middle" problem where isolated chunks lack surrounding context.
What is the "sandwich expansion" feature in RSE?
Sandwich expansion occurs when the algorithm detects a high-value segment that skips over one or more low-relevance chunks. Rather than returning disjointed fragments, RSE includes the intervening "sandwich" chunks between high-value sections. This ensures the LLM receives complete narrative flow even when relevance scores temporarily dip, preventing loss of critical transitional context.
How does the scoring fusion work in RSE?
Scoring fusion combines two metrics: absolute relevance (transformed via β-CDF with parameters a=b=0.4 to normalize raw similarity scores) and positional decay (calculated as exp(-rank/decay_rate) where decay_rate=30). The final chunk value equals the product of these values, ensuring that highly ranked chunks contribute disproportionately to segment scores while maintaining score distribution across the full 0-1 range.
What are the performance characteristics of RSE?
The maximum-sum sub-array search runs in approximately 5-10ms per document, making it suitable for production RAG pipelines. The algorithm uses brute-force search with early-stop heuristics rather than complex dynamic programming, balancing implementation simplicity with speed. This latency is negligible compared to vector search and LLM inference times, allowing RSE to operate as a real-time post-processing layer without impacting user experience.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →