How to Set the Chunking Engine in Cognee: CLI and Python API Guide
Pass the --chunker flag in the CLI or provide the chunker parameter to cognee.cognify() to switch between TextChunker, LangchainChunker, or CsvChunker.
Cognee is an open-source graph RAG framework that uses a chunking engine to partition raw documents into manageable text segments before processing them through LLM-based pipelines. The framework supports multiple chunking strategies that you can configure at runtime through either command-line arguments or direct API calls. Understanding how to set the chunking engine in Cognee allows you to optimize tokenization granularity for your specific document types and downstream vector stores.
Understanding Cognee's Chunking Architecture
The chunking engine selection determines how documents are split into DocumentChunk objects during the ingestion pipeline. According to the Cognee source code, the system supports three distinct chunking implementations defined in cognee/cli/config.py:
# cognee/cli/config.py
CHUNKER_CHOICES = ["TextChunker", "LangchainChunker", "CsvChunker"]
These chunkers are instantiated dynamically based on your configuration and are passed through the pipeline to document-type-specific readers in cognee/modules/data/processing/document_types/.
Setting the Chunking Engine via CLI
The simplest way to set the chunking engine in Cognee is through the cognee-cli cognify command. The CLI handler is implemented in cognee/cli/commands/cognify_command.py, which parses the --chunker argument and resolves it to the appropriate class.
Supported CLI Options
The CLI recognizes the following chunker names:
- TextChunker: Default paragraph-based splitting (no external dependencies)
- LangchainChunker: Uses LangChain's
RecursiveCharacterTextSplitterwith overlap support - CsvChunker: Row-based splitting for tabular data ingestion
CLI Usage Example
To process documents with the LangchainChunker, run:
cognee-cli cognify \
--chunker LangchainChunker \
--chunk-size 512 \
--chunks-per-batch 100 \
path/to/documents
Behind the scenes, cognee/cli/commands/cognify_command.py dynamically imports the requested class:
# cognee/cli/commands/cognify_command.py (excerpt)
if args.chunker == "LangchainChunker":
from cognee.modules.chunking.LangchainChunker import LangchainChunker
chunker_class = LangchainChunker
elif args.chunker == "CsvChunker":
from cognee.modules.chunking.CsvChunker import CsvChunker
chunker_class = CsvChunker
else:
chunker_class = TextChunker
The selected chunker_class is then forwarded to the cognify pipeline defined in cognee/pipelines.py.
Setting the Chunking Engine via Python API
For programmatic control, you can set the chunking engine by passing a chunker class (not an instance) to the high-level cognify function in cognee/api/v1/cognify/cognify.py.
The cognify Function Signature
The chunker parameter accepts a callable class definition:
# cognee/api/v1/cognify/cognify.py (excerpt)
async def cognify(
user: User = Depends(get_authenticated_user),
documents: List[Document],
chunker: Callable = TextChunker,
chunk_size: Optional[int] = None,
chunks_per_batch: Optional[int] = None,
...
):
await run_cognify_pipeline(
user=user,
documents=documents,
chunker=chunker,
chunk_size=chunk_size,
chunks_per_batch=chunks_per_batch,
)
Python Usage Example
To use a custom chunking engine in your Python code:
import asyncio
from cognee.api.v1.cognify.cognify import cognify
from cognee.modules.chunking.LangchainChunker import LangchainChunker
from cognee.modules.data.processing.document_types import TextDocument
async def main():
docs = [TextDocument("example.txt")]
await cognify(
user=None,
documents=docs,
chunker=LangchainChunker, # Set the chunking engine
chunk_size=1024,
chunks_per_batch=50,
)
if __name__ == "__main__":
asyncio.run(main())
How Chunkers Work Internally
Once you set the chunking engine, the system instantiates your chosen class within document-specific readers. For example, TextDocument.py implements the read method that accepts the chunker class:
# cognee/modules/data/processing/document_types/TextDocument.py (excerpt)
async def read(self, chunker_cls: Chunker, max_chunk_size: int):
chunker = chunker_cls(self, max_chunk_size=max_chunk_size, get_text=get_text)
async for chunk in chunker.read():
yield chunk
This pattern is consistent across all document types including UnstructuredDocument.py and PdfDocument.py. The chunker's read() coroutine yields DocumentChunk objects that flow into graph-building and vector-indexing stages.
Available Chunking Engines
Cognee provides three built-in chunking engines located in cognee/modules/chunking/:
TextChunker (TextChunker.py)
The default engine that performs simple paragraph-based splitting. This is the most reliable option with no external dependencies, ideal for general text documents.
LangchainChunker (LangchainChunker.py)
Leverages LangChain's RecursiveCharacterTextSplitter with configurable overlap. Better suited for long, unstructured text where you need semantic boundaries preserved across chunks.
CsvChunker (CsvChunker.py)
Specialized for tabular data, splitting CSV rows into separate chunks. Essential when ingesting structured datasets that require row-level granularity in your knowledge graph.
Summary
- Set the chunking engine in Cognee by passing
--chunkerto the CLI or thechunkerparameter tocognify(). - Available options are
TextChunker(default),LangchainChunker, andCsvChunker, defined incognee/cli/config.py. - The CLI handler in
cognee/cli/commands/cognify_command.pydynamically imports chunker classes based on the string name provided. - In the Python API, pass the chunker class (not an instance) to the
chunkerargument incognee/api/v1/cognify/cognify.py. - Document readers in
cognee/modules/data/processing/document_types/instantiate the chunker class and yield chunks through theread()coroutine.
Frequently Asked Questions
How do I change the default chunker for all operations?
Pass the desired chunker class explicitly in every API call or CLI command. Cognee does not currently persist a global configuration file for chunking defaults, but you can wrap the cognify function in your own module that presets chunker=LangchainChunker.
What is the difference between TextChunker and LangchainChunker?
TextChunker uses simple paragraph boundaries and requires no external dependencies, making it faster for basic text. LangchainChunker uses RecursiveCharacterTextSplitter with overlap support, providing better semantic coherence for complex documents but requiring the LangChain library.
Can I use a custom chunking engine not listed in CHUNKER_CHOICES?
Yes. When using the Python API, you can pass any callable class that implements the chunker interface to the chunker parameter. The CLI restricts choices to CHUNKER_CHOICES for safety, but programmatic usage accepts custom implementations as long as they match the expected signature in cognee/modules/data/processing/document_types/.
Does the chunk size parameter affect all chunking engines?
Yes. The chunk_size and chunks_per_batch parameters are forwarded to the pipeline regardless of which engine you set. However, specific chunkers may interpret chunk_size differently—LangchainChunker uses it as a token limit, while CsvChunker may use it as a row limit depending on implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →