How to Customize the Chunking Strategy in Cognee: Complete Implementation Guide
You customize chunking in Cognee by configuring a ChunkingConfig dictionary with your chosen ChunkEngine (Default, LangChain, or Haystack), selecting a ChunkStrategy (EXACT, PARAGRAPH, SENTENCE, CODE, or LANGCHAIN_CHARACTER), and passing these to create_chunking_engine() from cognee/infrastructure/data/chunking/create_chunking_engine.py.
Cognee is an open-source knowledge graph framework that processes documents into searchable chunks. The chunking subsystem determines how raw text splits into retrievable segments, directly impacting retrieval accuracy and context preservation. This guide explains how to configure and extend the chunking pipeline using the actual implementation in the topoteretes/cognee repository.
Understanding Cognee's Chunking Architecture
Cognee's chunking system operates through three interconnected layers defined in the source code. Understanding these layers helps you select the right combination of engine and strategy for your use case.
The Three Core Components
-
Chunk Engine: Chooses the concrete implementation that performs the split. The factory in
cognee/infrastructure/data/chunking/create_chunking_engine.pyinstantiatesDefaultChunkEngine,LangchainChunkEngine, orHaystackChunkEnginebased on your configuration. -
Chunk Strategy: Determines how the text is split. The
ChunkStrategyenum incognee/shared/data_models.pydefines the available algorithms:EXACT,PARAGRAPH,SENTENCE,CODE, andLANGCHAIN_CHARACTER. -
Engine Implementations: Contain the actual splitting logic.
DefaultChunkEngine.pyprovides regex-based splitting for paragraphs, sentences, and exact lengths, whileLangchainChunkingEngine.pyleverageslangchain_text_splittersfor code-aware and character-based splitting.
Configuring Your Chunking Strategy
Configuration happens through a dictionary-like ChunkingConfig object that specifies the engine, strategy, size limits, and overlap. These parameters control how documents transform into chunks before vectorization.
The configuration requires four key fields:
chunk_engine: AChunkEngineenum value (DEFAULT_ENGINE,LANGCHAIN_ENGINE, orHAYSTACK_ENGINE)chunk_strategy: AChunkStrategyenum value determining the splitting methodchunk_size: Maximum characters (or tokens) per chunkchunk_overlap: Number of characters repeated between consecutive chunks to preserve context
Available Chunk Engines and Strategies
Cognee provides three distinct engines, each supporting specific strategies. Choose the engine that aligns with your document type and splitting requirements.
DefaultChunkEngine
The DefaultChunkEngine in cognee/infrastructure/data/chunking/DefaultChunkEngine.py provides pure-Python regex-based splitting. It works out-of-the-box without external dependencies and supports:
ChunkStrategy.PARAGRAPH: Splits on double newlinesChunkStrategy.SENTENCE: Splits on sentence boundariesChunkStrategy.EXACT: Splits at exact character counts
LangchainChunkEngine
The LangchainChunkEngine in cognee/infrastructure/data/chunking/LangchainChunkingEngine.py wraps LangChain's text splitters. Use this engine for:
ChunkStrategy.CODE: Language-aware splitting for Python, JavaScript, and other programming languagesChunkStrategy.LANGCHAIN_CHARACTER: Character-based recursive splitting with configurable separators
HaystackChunkEngine
The HaystackChunkEngine in cognee/infrastructure/data/chunking/HaystackChunkEngine.py currently serves as a placeholder stub for future Haystack-based implementations.
Implementation Examples
Basic Configuration with Default Engine
Configure paragraph-based chunking using the default regex engine. This example creates chunks of 1500 characters with 200-character overlaps between consecutive chunks.
from cognee.shared.data_models import ChunkEngine, ChunkStrategy
from cognee.infrastructure.data.chunking.create_chunking_engine import create_chunking_engine
config = {
"chunk_engine": ChunkEngine.DEFAULT_ENGINE,
"chunk_size": 1500,
"chunk_overlap": 200,
"chunk_strategy": ChunkStrategy.PARAGRAPH,
}
engine = create_chunking_engine(config)
chunks, numbers = engine.chunk_data(
source_data="Your long document text here...",
chunk_strategy=config["chunk_strategy"],
chunk_size=config["chunk_size"],
chunk_overlap=config["chunk_overlap"],
)
Code-Aware Chunking with LangChain
Use the LangChain engine for programming languages. This configuration automatically detects language syntax and splits at logical boundaries like function definitions and class blocks.
from cognee.shared.data_models import ChunkEngine, ChunkStrategy
from cognee.infrastructure.data.chunking.create_chunking_engine import create_chunking_engine
config = {
"chunk_engine": ChunkEngine.LANGCHAIN_ENGINE,
"chunk_size": 1000,
"chunk_overlap": 100,
"chunk_strategy": ChunkStrategy.CODE,
}
engine = create_chunking_engine(config)
code = """def foo():\n return "bar"\n\nclass Baz:\n pass"""
chunks, numbers = engine.chunk_data(
source_data=code,
chunk_strategy=config["chunk_strategy"],
chunk_size=config["chunk_size"],
chunk_overlap=config["chunk_overlap"],
)
Pipeline Integration
Integrate custom chunking into document processing pipelines using extract_chunks_from_documents from cognee/tasks/documents/extract_chunks_from_documents.py. This task handles document reading, chunking execution, and metadata tracking.
from cognee.tasks.documents.extract_chunks_from_documents import extract_chunks_from_documents
from cognee.modules.chunking.TextChunker import TextChunker
async for chunk in extract_chunks_from_documents(
documents=documents,
max_chunk_size=1200,
chunker=TextChunker,
):
# Process chunk.text, chunk.chunk_size, chunk.chunk_id
process(chunk)
Creating a Custom Chunker Class
Extend the base Chunker class to implement domain-specific logic. This example wraps the DefaultChunkEngine with custom paragraph handling.
from uuid import UUID
from cognee.modules.chunking.Chunker import Chunker
from cognee.infrastructure.data.chunking.DefaultChunkEngine import DefaultChunkEngine
from cognee.shared.data_models import ChunkStrategy
class MyParagraphChunker(Chunker):
@classmethod
async def read(cls, max_chunk_size, chunker_cls=DefaultChunkEngine, **kwargs):
engine = chunker_cls(
chunk_strategy=ChunkStrategy.PARAGRAPH,
chunk_size=max_chunk_size,
chunk_overlap=0,
)
chunks, _ = engine.chunk_data(source_data=kwargs["source_data"])
for i, txt in enumerate(chunks, 1):
yield cls.Chunk(
text=txt,
chunk_size=len(txt),
chunk_id=str(UUID(int=i)),
)
Pass MyParagraphChunker to extract_chunks_from_documents to override default behavior across your pipeline.
Summary
- Chunking configuration in Cognee uses a dictionary with
chunk_engine,chunk_strategy,chunk_size, andchunk_overlapkeys. - Three engines are available:
DefaultChunkEnginefor regex-based splitting,LangchainChunkEnginefor code-aware splitting, andHaystackChunkEnginefor future extensions. - Five strategies cover most use cases:
EXACT,PARAGRAPH,SENTENCE,CODE, andLANGCHAIN_CHARACTER. - Factory pattern:
create_chunking_engine()incognee/infrastructure/data/chunking/create_chunking_engine.pyinstantiates the correct engine based on your config. - Pipeline integration: Use
extract_chunks_from_documentsto process documents with your custom chunker class.
Frequently Asked Questions
What is the difference between ChunkEngine and ChunkStrategy in Cognee?
ChunkEngine selects the software implementation (Default, LangChain, or Haystack) that executes the split, while ChunkStrategy defines the algorithmic approach (paragraph, sentence, code, etc.). The engine determines which code performs the split; the strategy determines how that code splits the text. You configure both in the ChunkingConfig dictionary passed to create_chunking_engine().
How do I choose between DefaultChunkEngine and LangchainChunkEngine?
Use DefaultChunkEngine for simple text documents requiring fast, dependency-free processing with paragraph, sentence, or exact-length splitting. Use LangchainChunkEngine when processing code files or when you need LangChain's recursive character text splitting with custom separators. The LangChain engine requires the langchain_text_splitters package but provides more sophisticated language-aware boundaries.
Can I implement a completely custom chunking algorithm in Cognee?
Yes. Create a subclass of Chunker from cognee/modules/chunking/Chunker.py and implement the read class method. Your implementation can use any splitting logic—whether wrapping an existing engine like DefaultChunkEngine or using external libraries. Pass your custom class to extract_chunks_from_documents or use it directly in your data pipeline.
What file should I modify to change chunking behavior globally?
Modify cognee/infrastructure/data/chunking/create_chunking_engine.py if you need to add new engine types to the factory function. For strategy definitions, update cognee/shared/data_models.py where the ChunkStrategy enum lives. However, for most use cases, you should pass custom configurations to the existing factory rather than modifying source files, ensuring your changes persist through updates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →