How to Customize the Chunking Strategy in Cognee: Complete Implementation Guide

You customize chunking in Cognee by configuring a ChunkingConfig dictionary with your chosen ChunkEngine (Default, LangChain, or Haystack), selecting a ChunkStrategy (EXACT, PARAGRAPH, SENTENCE, CODE, or LANGCHAIN_CHARACTER), and passing these to create_chunking_engine() from cognee/infrastructure/data/chunking/create_chunking_engine.py.

Cognee is an open-source knowledge graph framework that processes documents into searchable chunks. The chunking subsystem determines how raw text splits into retrievable segments, directly impacting retrieval accuracy and context preservation. This guide explains how to configure and extend the chunking pipeline using the actual implementation in the topoteretes/cognee repository.

Understanding Cognee's Chunking Architecture

Cognee's chunking system operates through three interconnected layers defined in the source code. Understanding these layers helps you select the right combination of engine and strategy for your use case.

The Three Core Components

  • Chunk Engine: Chooses the concrete implementation that performs the split. The factory in cognee/infrastructure/data/chunking/create_chunking_engine.py instantiates DefaultChunkEngine, LangchainChunkEngine, or HaystackChunkEngine based on your configuration.

  • Chunk Strategy: Determines how the text is split. The ChunkStrategy enum in cognee/shared/data_models.py defines the available algorithms: EXACT, PARAGRAPH, SENTENCE, CODE, and LANGCHAIN_CHARACTER.

  • Engine Implementations: Contain the actual splitting logic. DefaultChunkEngine.py provides regex-based splitting for paragraphs, sentences, and exact lengths, while LangchainChunkingEngine.py leverages langchain_text_splitters for code-aware and character-based splitting.

Configuring Your Chunking Strategy

Configuration happens through a dictionary-like ChunkingConfig object that specifies the engine, strategy, size limits, and overlap. These parameters control how documents transform into chunks before vectorization.

The configuration requires four key fields:

  • chunk_engine: A ChunkEngine enum value (DEFAULT_ENGINE, LANGCHAIN_ENGINE, or HAYSTACK_ENGINE)
  • chunk_strategy: A ChunkStrategy enum value determining the splitting method
  • chunk_size: Maximum characters (or tokens) per chunk
  • chunk_overlap: Number of characters repeated between consecutive chunks to preserve context

Available Chunk Engines and Strategies

Cognee provides three distinct engines, each supporting specific strategies. Choose the engine that aligns with your document type and splitting requirements.

DefaultChunkEngine

The DefaultChunkEngine in cognee/infrastructure/data/chunking/DefaultChunkEngine.py provides pure-Python regex-based splitting. It works out-of-the-box without external dependencies and supports:

  • ChunkStrategy.PARAGRAPH: Splits on double newlines
  • ChunkStrategy.SENTENCE: Splits on sentence boundaries
  • ChunkStrategy.EXACT: Splits at exact character counts

LangchainChunkEngine

The LangchainChunkEngine in cognee/infrastructure/data/chunking/LangchainChunkingEngine.py wraps LangChain's text splitters. Use this engine for:

  • ChunkStrategy.CODE: Language-aware splitting for Python, JavaScript, and other programming languages
  • ChunkStrategy.LANGCHAIN_CHARACTER: Character-based recursive splitting with configurable separators

HaystackChunkEngine

The HaystackChunkEngine in cognee/infrastructure/data/chunking/HaystackChunkEngine.py currently serves as a placeholder stub for future Haystack-based implementations.

Implementation Examples

Basic Configuration with Default Engine

Configure paragraph-based chunking using the default regex engine. This example creates chunks of 1500 characters with 200-character overlaps between consecutive chunks.

from cognee.shared.data_models import ChunkEngine, ChunkStrategy
from cognee.infrastructure.data.chunking.create_chunking_engine import create_chunking_engine

config = {
    "chunk_engine": ChunkEngine.DEFAULT_ENGINE,
    "chunk_size": 1500,
    "chunk_overlap": 200,
    "chunk_strategy": ChunkStrategy.PARAGRAPH,
}

engine = create_chunking_engine(config)
chunks, numbers = engine.chunk_data(
    source_data="Your long document text here...",
    chunk_strategy=config["chunk_strategy"],
    chunk_size=config["chunk_size"],
    chunk_overlap=config["chunk_overlap"],
)

Code-Aware Chunking with LangChain

Use the LangChain engine for programming languages. This configuration automatically detects language syntax and splits at logical boundaries like function definitions and class blocks.

from cognee.shared.data_models import ChunkEngine, ChunkStrategy
from cognee.infrastructure.data.chunking.create_chunking_engine import create_chunking_engine

config = {
    "chunk_engine": ChunkEngine.LANGCHAIN_ENGINE,
    "chunk_size": 1000,
    "chunk_overlap": 100,
    "chunk_strategy": ChunkStrategy.CODE,
}

engine = create_chunking_engine(config)
code = """def foo():\n    return "bar"\n\nclass Baz:\n    pass"""
chunks, numbers = engine.chunk_data(
    source_data=code,
    chunk_strategy=config["chunk_strategy"],
    chunk_size=config["chunk_size"],
    chunk_overlap=config["chunk_overlap"],
)

Pipeline Integration

Integrate custom chunking into document processing pipelines using extract_chunks_from_documents from cognee/tasks/documents/extract_chunks_from_documents.py. This task handles document reading, chunking execution, and metadata tracking.

from cognee.tasks.documents.extract_chunks_from_documents import extract_chunks_from_documents
from cognee.modules.chunking.TextChunker import TextChunker

async for chunk in extract_chunks_from_documents(
    documents=documents,
    max_chunk_size=1200,
    chunker=TextChunker,
):
    # Process chunk.text, chunk.chunk_size, chunk.chunk_id

    process(chunk)

Creating a Custom Chunker Class

Extend the base Chunker class to implement domain-specific logic. This example wraps the DefaultChunkEngine with custom paragraph handling.

from uuid import UUID
from cognee.modules.chunking.Chunker import Chunker
from cognee.infrastructure.data.chunking.DefaultChunkEngine import DefaultChunkEngine
from cognee.shared.data_models import ChunkStrategy

class MyParagraphChunker(Chunker):
    @classmethod
    async def read(cls, max_chunk_size, chunker_cls=DefaultChunkEngine, **kwargs):
        engine = chunker_cls(
            chunk_strategy=ChunkStrategy.PARAGRAPH,
            chunk_size=max_chunk_size,
            chunk_overlap=0,
        )
        chunks, _ = engine.chunk_data(source_data=kwargs["source_data"])
        for i, txt in enumerate(chunks, 1):
            yield cls.Chunk(
                text=txt,
                chunk_size=len(txt),
                chunk_id=str(UUID(int=i)),
            )

Pass MyParagraphChunker to extract_chunks_from_documents to override default behavior across your pipeline.

Summary

  • Chunking configuration in Cognee uses a dictionary with chunk_engine, chunk_strategy, chunk_size, and chunk_overlap keys.
  • Three engines are available: DefaultChunkEngine for regex-based splitting, LangchainChunkEngine for code-aware splitting, and HaystackChunkEngine for future extensions.
  • Five strategies cover most use cases: EXACT, PARAGRAPH, SENTENCE, CODE, and LANGCHAIN_CHARACTER.
  • Factory pattern: create_chunking_engine() in cognee/infrastructure/data/chunking/create_chunking_engine.py instantiates the correct engine based on your config.
  • Pipeline integration: Use extract_chunks_from_documents to process documents with your custom chunker class.

Frequently Asked Questions

What is the difference between ChunkEngine and ChunkStrategy in Cognee?

ChunkEngine selects the software implementation (Default, LangChain, or Haystack) that executes the split, while ChunkStrategy defines the algorithmic approach (paragraph, sentence, code, etc.). The engine determines which code performs the split; the strategy determines how that code splits the text. You configure both in the ChunkingConfig dictionary passed to create_chunking_engine().

How do I choose between DefaultChunkEngine and LangchainChunkEngine?

Use DefaultChunkEngine for simple text documents requiring fast, dependency-free processing with paragraph, sentence, or exact-length splitting. Use LangchainChunkEngine when processing code files or when you need LangChain's recursive character text splitting with custom separators. The LangChain engine requires the langchain_text_splitters package but provides more sophisticated language-aware boundaries.

Can I implement a completely custom chunking algorithm in Cognee?

Yes. Create a subclass of Chunker from cognee/modules/chunking/Chunker.py and implement the read class method. Your implementation can use any splitting logic—whether wrapping an existing engine like DefaultChunkEngine or using external libraries. Pass your custom class to extract_chunks_from_documents or use it directly in your data pipeline.

What file should I modify to change chunking behavior globally?

Modify cognee/infrastructure/data/chunking/create_chunking_engine.py if you need to add new engine types to the factory function. For strategy definitions, update cognee/shared/data_models.py where the ChunkStrategy enum lives. However, for most use cases, you should pass custom configurations to the existing factory rather than modifying source files, ensuring your changes persist through updates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →