# Implementing Proposition Chunking for RAG: A Complete Technical Guide

> Learn proposition chunking for RAG and boost retrieval precision. This guide details how to decompose documents into factual statements for higher accuracy.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: how-to-guide
- Published: 2026-02-19

---

**Proposition chunking decomposes documents into atomic, self-contained factual statements that are individually embedded and quality-filtered, enabling retrieval systems to match specific queries with significantly higher precision than traditional passage-based chunking.**

Implementing proposition chunking for RAG requires shifting from arbitrary text segmentation to semantic fact extraction. This technique, as implemented in the `NirDiamant/RAG_Techniques` repository, uses a multi-stage pipeline involving LLM-based generation and automated quality gating to create a high-precision retrieval index. The approach minimizes noise by ensuring each indexed unit contains exactly one verifiable fact rather than overlapping narrative context.

## Architecture of the Proposition Chunking Pipeline

The implementation in `all_rag_techniques/proposition_chunking.ipynb` follows a six-stage architecture that transforms raw documents into a searchable proposition store.

### Stage 1: Environment Setup and API Configuration

The pipeline begins by loading API credentials from a `.env` file to authenticate with the Groq LLM service. According to the source code, the notebook uses `load_dotenv()` to read environment variables and explicitly sets `os.environ['GROQ_API_KEY']` from the loaded values (lines 94-102). This ensures the subsequent LLM calls for proposition generation and grading are properly authenticated.

### Stage 2: Initial Document Chunking

Before proposition generation, the system splits raw text into manageable windows using **LangChain's** `RecursiveCharacterTextSplitter`. The implementation specifically uses `from_tiktoken_encoder(chunk_size=200, chunk_overlap=50)` to create approximately 200-token segments with 50-token overlaps (lines 115-124). This intermediate chunking ensures the LLM receives contextually coherent input without exceeding token limits during the proposition extraction phase.

### Stage 3: Proposition Generation with LLMs

Each chunk is processed by a `ChatGroq` instance running the `llama-3.1-70b-versatile` model. The pipeline constructs a `ChatPromptTemplate` that instructs the model to output structured data conforming to the `GeneratePropositions` Pydantic schema, which defines propositions as a list of atomic, factual, self-contained strings (lines 129-145). This structured output approach ensures consistent parsing of the generated facts.

### Stage 4: Quality Assessment and Filtering

A second LLM invocation grades every generated proposition on four criteria: **accuracy**, **clarity**, **completeness**, and **conciseness**. The `grade_chain.invoke()` method (lines 149-155) evaluates each statement against defined thresholds, filtering out ambiguous or incomplete facts before they enter the vector store. This quality gate is critical for maintaining retrieval reliability.

### Stage 5: Embedding and Vector Storage

Accepted propositions are embedded using **Ollama's** `nomic-embed-text:v1.5` model via `OllamaEmbeddings`. The system instantiates a FAISS vector store by calling `FAISS.from_texts(propositions, embedding_model)` (lines 162-170), creating a high-performance similarity search index optimized for granular factual retrieval rather than broad semantic similarity.

### Stage 6: Dual Retrieval and Comparative Analysis

The notebook implements two parallel retrievers to demonstrate the technique's effectiveness: `prop_retriever` queries the proposition store while `chunk_retriever` searches the original larger text segments (lines 176-181). This dual setup allows practitioners to compare relevance scores and contextual breadth between granular and traditional chunking strategies.

## Step-by-Step Implementation

The following runnable Python excerpt reproduces the core pipeline from the repository. It requires `langchain`, `langchain-groq`, `langchain-community`, `faiss-cpu`, `python-dotenv`, and `pydantic`.

```python
import os
from dotenv import load_dotenv
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import FAISS
from langchain_community.embeddings import OllamaEmbeddings
from langchain_groq import ChatGroq
from langchain_core.prompts import ChatPromptTemplate
from pydantic import BaseModel, Field
from typing import List

# 1️⃣ Load API key

load_dotenv()
os.environ["GROQ_API_KEY"] = os.getenv("GROQ_API_KEY")

# 2️⃣ Split the document

raw_text = """Paul Graham argues that founder mode is distinct from manager mode. 
In founder mode, founders must micromanage certain critical aspects of the company 
while delegating others. This is different from the standard advice given to scale 
companies, which suggests hiring good people and giving them space to work."""
splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
    chunk_size=200, chunk_overlap=50
)
chunks = splitter.split_text(raw_text)

# 3️⃣ Generate propositions

class GeneratePropositions(BaseModel):
    propositions: List[str] = Field(
        description="Atomic, factual, self‑contained statements"
    )

prompt = ChatPromptTemplate.from_template(
    """Given the following text chunk, extract a list of atomic propositions. 
    Each proposition should be a single, verifiable fact that stands alone without context.
    
    Chunk: {chunk}"""
)
llm = ChatGroq(model="llama-3.1-70b-versatile", temperature=0)
chain = prompt | llm.with_structured_output(GeneratePropositions)

all_props = []
for c in chunks:
    result = chain.invoke({"chunk": c})
    all_props.extend(result.propositions)

# 4️⃣ Quality check (simplified version)

def is_valid_proposition(p: str) -> bool:
    """Placeholder for LLM-based grading. Real implementation uses a second 
    chain with structured output for accuracy, clarity, completeness, and conciseness."""
    return len(p.split()) < 30 and "?" not in p and len(p) > 10

good_props = [p for p in all_props if is_valid_proposition(p)]

# 5️⃣ Embed and store

embedder = OllamaEmbeddings(model="nomic-embed-text:v1.5")
vectorstore = FAISS.from_texts(good_props, embedder)

# 6️⃣ Retrieve

retriever = vectorstore.as_retriever(search_kwargs={"k": 5})
query = "What does Paul Graham say about founder mode?"
answers = retriever.invoke(query)
for doc in answers:
    print(f"- {doc.page_content}")

```

## Why Proposition Chunking Improves RAG Performance

The technique delivers superior retrieval precision through three core mechanisms:

- **Granular indexing reduces noise**: Each vector represents exactly one semantic fact, eliminating the dilution effect where similarity scores average across multiple unrelated concepts within a large chunk.
- **Quality gating ensures reliability**: By filtering propositions through a secondary LLM evaluation for accuracy and conciseness, the system prevents hallucinated or ambiguous statements from polluting the knowledge base.
- **Optimal context boundaries**: Unlike fixed-size chunks that may split mid-sentence or combine unrelated facts, propositions align retrieval units with natural information boundaries.

## Key Files and Implementation Details

The proposition chunking system is distributed across the following files in the `NirDiamant/RAG_Techniques` repository:

- **`all_rag_techniques/proposition_chunking.ipynb`**: The primary implementation notebook containing the six-stage pipeline, dual-retriever comparison logic, and visualization of results. This file demonstrates the complete workflow from raw text to searchable proposition index.

- **[`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py)**: Shared utility modules providing text preprocessing, PDF loading helpers, and generic FAISS persistence functions used across multiple RAG technique implementations in the repository.

- **`images/proposition_chunking.svg`**: Architectural diagram illustrating the data flow from document ingestion through proposition generation to final retrieval.

## Summary

Implementing proposition chunking for RAG transforms document processing from coarse text segmentation into semantic fact extraction. The approach combines `RecursiveCharacterTextSplitter` for initial windowing, `ChatGroq` with structured Pydantic outputs for atomic proposition generation, and secondary LLM grading for quality assurance. Key takeaways include:

- Proposition chunking creates **atomic, self-contained factual units** that align precisely with specific query intents
- The **dual-retriever architecture** in the reference implementation allows empirical comparison between granular and traditional chunking strategies
- **Quality filtering** via secondary LLM evaluation is essential for maintaining retrieval accuracy
- **Ollama embeddings** with FAISS provide a cost-effective, high-performance vector storage solution for the generated propositions

## Frequently Asked Questions

### What is the difference between proposition chunking and standard text chunking?

Standard chunking divides documents into fixed-size segments (e.g., 200 tokens) that may split sentences or combine unrelated facts. Proposition chunking uses an LLM to extract **atomic factual statements** that are self-contained and semantically complete. This granularity ensures that similarity searches match specific facts rather than broad text segments, significantly improving precision for targeted queries.

### Which LLM models work best for generating propositions?

The reference implementation uses **Groq's `llama-3.1-70b-versatile`** for both proposition generation and quality grading. Any capable instruction-tuned model with structured output support (such as GPT-4, Claude 3, or Llama 3 variants) can perform this task, provided it follows the Pydantic schema constraints and produces concise, factual statements. Smaller models may require more restrictive prompting to maintain output quality.

### How does the quality assessment mechanism filter propositions?

The quality gate employs a second LLM invocation (distinct from the generation step) that evaluates each proposition on four axes: **accuracy** (factual correctness), **clarity** (unambiguous phrasing), **completeness** (self-contained context), and **conciseness** (brevity without loss of meaning). Propositions must meet thresholds on all criteria to be accepted into the FAISS vector store, ensuring only high-confidence facts are indexed.

### When should I use proposition chunking instead of traditional RAG?

Use proposition chunking when your use case requires **high-precision retrieval of specific facts** rather than broad context gathering. It excels for question-answering systems dealing with dense technical documentation, legal texts, or scientific papers where users ask precise questions. Traditional chunking remains preferable for tasks requiring narrative flow or when processing documents where facts cannot be easily isolated into atomic statements.