Implementing Semantic Chunking for RAG: A Complete Guide to Meaning-Based Document Splitting
Semantic chunking splits documents into meaning-based sections using embedding similarity rather than fixed character counts, significantly improving retrieval relevance in RAG pipelines.
Implementing semantic chunking for RAG allows you to replace arbitrary text splits with intelligent boundaries that preserve context and meaning. The NirDiamant/RAG_Techniques repository provides a production-ready implementation through the SemanticChunkingRAG class, which leverages LangChain's experimental semantic chunker and FAISS vector storage to create a modular pipeline for any RAG workflow.
What Is Semantic Chunking for RAG?
Traditional chunking methods split documents at fixed character or token boundaries, often severing sentences or semantic units mid-thought. Semantic chunking analyzes embedding similarities between text segments to identify natural topic boundaries, creating chunks that are self-contained and contextually coherent.
In the RAG_Techniques implementation, this process uses langchain_experimental.text_splitter.SemanticChunker with configurable breakpoint detection strategies (percentile, standard deviation, interquartile range, or gradient) to determine where semantic shifts occur.
Architecture of the Semantic Chunking Implementation
The SemanticChunkingRAG class in all_rag_techniques_runnable_scripts/semantic_chunking.py orchestrates a four-stage pipeline that transforms raw PDFs into retrievable semantic chunks.
Document Ingestion and PDF Processing
The pipeline begins with read_pdf_to_string from helper_functions.py, which uses PyMuPDF (fitz) to extract raw text from PDF files. This utility handles the initial document loading and returns a clean string representation of the entire document content.
from helper_functions import read_pdf_to_string
text = read_pdf_to_string("data/Understanding_Climate_Change.pdf")
Embedding Generation and Breakpoint Detection
The SemanticChunker initializes with an embedding model—defaulting to OpenAIEmbeddings but compatible with any LangChain embedding provider. The chunker analyzes embedding distances between consecutive sentences to identify semantic breakpoints based on the specified BreakpointThresholdType.
from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings()
chunker = SemanticChunker(
embeddings,
breakpoint_threshold_type="percentile",
breakpoint_threshold_amount=85
)
Vector Storage and Retrieval
Once chunked, documents are embedded again and stored in a FAISS in-memory vector index via FAISS.from_documents. The as_retriever method creates a retriever interface that returns the k most semantically similar chunks for any query.
from langchain_community.vectorstores import FAISS
chunks = chunker.create_documents([text])
vectorstore = FAISS.from_documents(chunks, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 3})
Running Semantic Chunking from the Command Line
The repository includes a full CLI interface in semantic_chunking.py that allows immediate experimentation with different chunking parameters without writing code.
python all_rag_techniques_runnable_scripts/semantic_chunking.py \
--path data/Understanding_Climate_Change.pdf \
--n_retrieved 3 \
--breakpoint_threshold_type percentile \
--breakpoint_threshold_amount 85 \
--query "What are the primary drivers of climate change?"
CLI Arguments Explained:
--path: Path to the input PDF document.--n_retrieved: Number of semantic chunks to retrieve for the query.--breakpoint_threshold_type: Algorithm for detecting semantic shifts (percentile,standard_deviation,interquartile, orgradient).--breakpoint_threshold_amount: Sensitivity threshold (higher values produce fewer, larger chunks).--query: The natural language question used to test retrieval.
Programmatic Implementation with Python
For integration into existing RAG pipelines, instantiate the SemanticChunkingRAG class directly with custom parameters.
from all_rag_techniques_runnable_scripts.semantic_chunking import SemanticChunkingRAG
# Initialize with custom chunking parameters
rag = SemanticChunkingRAG(
path="data/Understanding_Climate_Change.pdf",
n_retrieved=2,
breakpoint_type="percentile",
breakpoint_amount=90,
)
# Execute retrieval and capture timing metrics
times = rag.run("How does deforestation affect global warming?")
print(f"Chunking time: {times['Chunking']:.2f}s")
print(f"Retrieval time: {times['Retrieval']:.2f}s")
The run method returns a dictionary containing execution times for the chunking and retrieval phases, enabling performance benchmarking across different breakpoint strategies.
Key Configuration Parameters
The effectiveness of semantic chunking depends heavily on breakpoint configuration. The SemanticChunker supports four threshold types:
- Percentile: Splits when embedding distance exceeds the Nth percentile of all distances (recommended starting point: 85-95).
- Standard Deviation: Splits when distance exceeds the mean plus N standard deviations (useful for documents with uniform topic distribution).
- Interquartile: Uses the interquartile range to detect outliers in semantic distance.
- Gradient: Detects rapid changes in embedding similarity gradients.
According to the RAG_Techniques source code, the percentile method provides the most consistent results across diverse document types, while gradient excels with technical documents containing abrupt topic shifts.
Summary
- Semantic chunking replaces fixed-size text splits with meaning-based boundaries, preserving context and improving retrieval accuracy in RAG systems.
- The NirDiamant/RAG_Techniques repository provides a complete implementation via the
SemanticChunkingRAGclass inall_rag_techniques_runnable_scripts/semantic_chunking.py. - The pipeline uses LangChain's
SemanticChunkerwith configurable breakpoint strategies (percentile, standard deviation, interquartile, gradient) and FAISS for vector storage. - Users can run the implementation via CLI arguments or integrate the
SemanticChunkingRAGclass directly into existing Python RAG pipelines.
Frequently Asked Questions
What is the difference between semantic chunking and recursive character chunking?
Recursive character chunking splits text at fixed character or token boundaries, often breaking sentences or ideas mid-thought. Semantic chunking analyzes embedding similarities between text segments to identify natural topic boundaries, creating chunks that are self-contained and contextually coherent. According to the RAG_Techniques implementation, semantic chunking typically yields higher retrieval relevance with fewer chunks than character-based methods.
How do I choose the right breakpoint threshold type?
The choice depends on your document structure. Use percentile (85-95) for general documents with mixed topics, as implemented in the default CLI example. Choose standard_deviation for uniformly distributed content where outliers indicate topic shifts. Select gradient for technical documents with abrupt subject changes. The interquartile method works best when you need robust statistical outlier detection resistant to extreme values.
Can I use different embedding models with this implementation?
Yes. The SemanticChunkingRAG class accepts any LangChain-compatible embedding provider. While the default implementation uses OpenAIEmbeddings, you can substitute Cohere, HuggingFace, Bedrock, or custom embeddings by passing the desired embedding instance to the SemanticChunker initialization in all_rag_techniques_runnable_scripts/semantic_chunking.py.
Why use FAISS as the vector store for semantic chunking?
FAISS provides fast, in-memory similarity search optimized for dense vector embeddings, making it ideal for prototyping and production RAG pipelines with semantic chunks. The RAG_Techniques implementation uses FAISS.from_documents to index the semantically-split chunks and as_retriever to configure the number of results (k) returned for each query, enabling sub-millisecond retrieval latency even with large document collections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →