# How to Use Relevant Segment Extraction in RAG: A Complete Implementation Guide

> Master Relevant Segment Extraction in RAG. This guide explains how RSE orders scattered chunks into cohesive segments, boosting LLM context coherence. Implement RAG techniques effectively.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: how-to-guide
- Published: 2026-02-19

---

**Relevant Segment Extraction (RSE) is a post-retrieval technique that converts scattered relevant chunks into contiguous ordered segments using a maximum-sum sub-array algorithm, improving context coherence for large language models.**

Relevant Segment Extraction solves the fragmentation problem in Retrieval-Augmented Generation systems where vector search returns isolated chunks that lose narrative flow. This guide explains the complete implementation of RSE as found in the `NirDiamant/RAG_Techniques` repository, transforming disjointed retrieval results into coherent text segments ready for LLM consumption.

## How RSE Works: The Architecture Pipeline

RSE operates as a **post-retrieval processing layer** that sits between your vector database and the LLM prompt. According to the implementation in `all_rag_techniques/relevant_segment_extraction.ipynb`, the pipeline follows six distinct stages to reconstruct meaningful context from retrieved fragments.

### Chunk Store and Initial Retrieval

The foundation of RSE is a **key-value chunk store** where each document is split into non-overlapping chunks and indexed by `doc_id` combined with `chunk_index`. This structure enables fast random access when rebuilding contiguous segments. The initial retrieval phase uses vector similarity search (optionally boosted by keyword search) to fetch candidate chunks, which may then pass through a cross-encoder reranker such as **Cohere Rerank** (`rerank-english-v3.0`).

### Scoring Fusion and Decay

For every retrieved chunk, the system computes a **combined relevance value** through two transformations:

- **Absolute relevance**: Raw similarity scores undergo a β-CDF transformation (`a=b=0.4`) to spread values uniformly across the 0-1 range
- **Decayed score**: `absolute_relevance × exp(-rank/decay_rate)` where `decay_rate` defaults to 30, ensuring higher-ranked chunks contribute more weight to the final segment value

### Maximum-Sum Sub-Array Search

Each chunk receives its decayed score as its *chunk value*. The algorithm then treats the document as an array of these values and executes a **brute-force maximum-sum sub-array search** (with early-stop heuristics) to locate the contiguous range with the highest total value. This operation completes in approximately **5-10ms per document**, making it suitable for real-time RAG pipelines.

### Sandwich Expansion

When the optimal segment skips over low-relevance chunks, RSE optionally **includes the intervening "sandwich" chunks** to preserve narrative continuity. This ensures the LLM receives full context even when relevance scores dip temporarily between highly relevant passages.

## Implementing RSE in Python

The reference implementation in `NirDiamant/RAG_Techniques` provides a complete, runnable pipeline. Below is the essential code structure adapted from `all_rag_techniques/relevant_segment_extraction.ipynb`.

### Prerequisites and Setup

```python
!pip install langchain-cohere python-dotenv scipy matplotlib

import os
import numpy as np
from typing import List
from scipy.stats import beta
import cohere
from dotenv import load_dotenv
from langchain_text_splitters import RecursiveCharacterTextSplitter

load_dotenv()
os.environ["CO_API_KEY"] = os.getenv("CO_API_KEY")

```

### Document Chunking Strategy

The `split_into_chunks` function creates the non-overlapping storage units required by the chunk store. It uses LangChain's `RecursiveCharacterTextSplitter` with a default size of 800 tokens and zero overlap to ensure clean segment boundaries.

```python
def split_into_chunks(text: str, chunk_size: int = 800) -> List[str]:
    splitter = RecursiveCharacterTextSplitter(
        chunk_size=chunk_size,
        chunk_overlap=0,
        length_function=len
    )
    docs = splitter.create_documents([text])
    return [doc.page_content for doc in docs]

```

### Relevance Scoring with Beta-CDF Transformation

The `transform` function applies the β-CDF spread to raw reranker scores, while `rerank_chunks` handles the API call and exponential decay application:

```python
def transform(x: float) -> float:
    a, b = 0.4, 0.4  # spreads scores uniformly across 0-1

    return beta.cdf(x, a, b)

def rerank_chunks(query: str, chunks: List[str]):
    client = cohere.Client(api_key=os.environ["CO_API_KEY"])
    decay_rate = 30
    
    resp = client.rerank(
        model="rerank-english-v3.0",
        query=query,
        documents=chunks
    )
    
    # Restore original order and compute decayed values

    similarity = [0] * len(chunks)
    values = [0] * len(chunks)
    
    for rank, result in enumerate(resp.results):
        idx = result.index
        raw = result.relevance_score
        transformed = transform(raw)
        similarity[idx] = transformed
        values[idx] = np.exp(-rank/decay_rate) * transformed
        
    return similarity, values

```

### Extracting the Optimal Segment

The core RSE algorithm implements a brute-force search to find the contiguous segment with maximum summed value:

```python
def find_best_segment(chunk_values: List[float]):
    """Brute-force max-sum sub-array with early break heuristics."""
    best_sum, best_start, best_end = -float("inf"), 0, 0
    n = len(chunk_values)
    
    for i in range(n):
        cur = 0
        for j in range(i, n):
            cur += chunk_values[j]
            if cur > best_sum:
                best_sum, best_start, best_end = cur, i, j
                
    return best_start, best_end, best_sum

# Usage example

document = """Your long document text here..."""
query = "What is the main contribution of the paper?"

chunks = split_into_chunks(document, chunk_size=800)
_, chunk_vals = rerank_chunks(query, chunks)

start, end, score = find_best_segment(chunk_vals)
relevant_segment = " ".join(chunks[start : end + 1])

print(f"Segment [{start}-{end}] (score {score:.2f}):\n{relevant_segment[:1000]}...")

```

## Key Functions and Source Files

The RSE implementation spans multiple files in the repository:

- **`all_rag_techniques/relevant_segment_extraction.ipynb`**: Contains the full pipeline including `plot_relevance_scores` for visualization, scoring logic, and evaluation metrics
- **[`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py)**: Provides shared utilities for data loading and plotting functions used across the RSE notebook
- **Scoring functions**: `transform()` applies the Beta-CDF with `a=b=0.4`, while the decay calculation uses `decay_rate=30` to weight rank position

The notebook also includes visualization tools to debug relevance distributions across chunk sequences before segment extraction.

## Summary

- **Relevant Segment Extraction** converts isolated retrieved chunks into ordered, contiguous text segments that preserve document flow for LLM prompting
- The algorithm uses **β-CDF transformation** (`a=b=0.4`) and **exponential decay** (`decay_rate=30`) to weight chunks by both absolute relevance and retrieval rank
- A **maximum-sum sub-array search** identifies the optimal contiguous segment in 5-10ms per document, enabling real-time RAG applications
- **Sandwich expansion** optionally includes low-relevance intervening chunks to maintain narrative coherence when context temporarily dips between high-value passages
- The complete implementation resides in `NirDiamant/RAG_Techniques`, specifically `all_rag_techniques/relevant_segment_extraction.ipynb`

## Frequently Asked Questions

### How does Relevant Segment Extraction differ from standard chunk retrieval?

Standard retrieval returns individual chunks ranked by similarity, often breaking narrative flow when relevant passages span multiple chunks. RSE reconstructs **contiguous multi-chunk segments** that preserve original document order, ensuring the LLM receives coherent context rather than scattered excerpts. This addresses the "lost in the middle" problem where isolated chunks lack surrounding context.

### What is the "sandwich expansion" feature in RSE?

Sandwich expansion occurs when the algorithm detects a high-value segment that skips over one or more low-relevance chunks. Rather than returning disjointed fragments, RSE **includes the intervening "sandwich" chunks** between high-value sections. This ensures the LLM receives complete narrative flow even when relevance scores temporarily dip, preventing loss of critical transitional context.

### How does the scoring fusion work in RSE?

Scoring fusion combines two metrics: **absolute relevance** (transformed via β-CDF with parameters `a=b=0.4` to normalize raw similarity scores) and **positional decay** (calculated as `exp(-rank/decay_rate)` where `decay_rate=30`). The final chunk value equals the product of these values, ensuring that highly ranked chunks contribute disproportionately to segment scores while maintaining score distribution across the full 0-1 range.

### What are the performance characteristics of RSE?

The maximum-sum sub-array search runs in approximately **5-10ms per document**, making it suitable for production RAG pipelines. The algorithm uses brute-force search with early-stop heuristics rather than complex dynamic programming, balancing implementation simplicity with speed. This latency is negligible compared to vector search and LLM inference times, allowing RSE to operate as a real-time post-processing layer without impacting user experience.