How to Use Text Chunking with Chonkie vs NaiveChunker in Sieves

Sieves provides two built-in chunking strategies—Chonkie for token-aware splitting and NaiveChunker for character-interval slicing—accessible through the unified Chunking task or individual task aliases.

The mantisai/sieves repository offers a flexible preprocessing module that handles text segmentation for downstream NLP pipelines. Whether you need linguistically coherent chunks that respect token boundaries or fast character-based slicing, Sieves abstracts both approaches behind a consistent API defined in sieves/tasks/preprocessing/chunking/core.py.

Understanding the Two Chunking Strategies

Sieves implements two distinct chunking philosophies:

  • Chonkie: A wrapper around the third-party chonkie library that provides token-aware chunking with configurable overlap, ideal for LLM context windows.
  • NaiveChunker: A lightweight, interval-based splitter that divides text every n characters without considering token boundaries.

Both are exposed through the high-level Chunking task, which accepts either a chonkie.BaseChunker instance or an integer to determine which implementation to instantiate.

Using the Chonkie Chunker for Token-Aware Splitting

Implementation Details

The Chonkie integration lives in sieves/tasks/preprocessing/chunking/chonkie_.py and is invoked when you pass a chonkie.BaseChunker to the task constructor. According to the source in core.py (lines 30-68), the logic branches based on type:


# From sieves/tasks/preprocessing/chunking/core.py

if isinstance(chunker, chonkie.BaseChunker):
    chunker_task = chonkie_.Chonkie(chunker=chunker)   # Chonkie path

elif isinstance(chunker, int):
    chunker_task = naive.NaiveChunker(interval=chunker)   # Naive path

Practical Example with TokenChunker

To create semantically coherent chunks using GPT-2 tokenization:

from sieves import Pipeline, tasks
import chonkie
from chonkie import tokenizers

tokenizer = tokenizers.Tokenizer.from_pretrained("gpt2")
chonkie_chunker = chonkie.TokenChunker(
    tokenizer, 
    chunk_size=512, 
    chunk_overlap=50
)

pipeline = Pipeline([tasks.Chunking(chonkie_chunker)])
doc = pipeline([sieves.Doc(text="Your very long document text here...")])[0]
print(doc.meta["Chunker"])

This configuration respects token boundaries, maintains a 50-token overlap between chunks, and stores metadata about the chunking operation.

Using the NaiveChunker for Character-Interval Splitting

Implementation Details

The NaiveChunker class in sieves/tasks/preprocessing/chunking/naive.py provides deterministic, high-speed text segmentation. When the Chunking task receives an integer argument, it automatically instantiates this class using the integer as the character interval.

Practical Example

For quick prototyping or when token boundaries are irrelevant:

from sieves import Pipeline, tasks

pipeline = Pipeline([tasks.Chunking(200)])  # 200-character slices

# Or explicitly:

pipeline = Pipeline([tasks.NaiveChunker(interval=200)])

doc = pipeline([sieves.Doc(text="Another long document...")])[0]
print(doc.meta["Chunker"])

This approach cuts the text every 200 characters regardless of word or token boundaries, offering maximum speed with minimal overhead.

How Sieves Resolves Chunker Types Internally

The unified interface is implemented in sieves/tasks/preprocessing/chunking/core.py through the _ChunkerArgType union type:

_ChunkerArgType = chonkie.BaseChunker | int

def __init__(self, chunker: _ChunkerArgType, ...):
    if isinstance(chunker, chonkie.BaseChunker):
        self._chunker = chonkie_.Chonkie(chunker=chunker)
    elif isinstance(chunker, int):
        self._chunker = naive.NaiveChunker(interval=chunker)
    else:
        raise TypeError("chunker must be a chonkie.BaseChunker or int")

This type-checking mechanism allows the same Chunking task to serve both sophisticated token-aware workflows and simple character-splitting needs without requiring separate import paths.

Inspecting Chunking Metadata

Both chunkers record execution metadata in the document's meta dictionary. As shown in the test suite (sieves/tests/docs/test_preprocessing.py, lines 149-155), you can inspect the results after pipeline execution:

doc = pipe([sieves.Doc(text="Your long document")])[0]
print(doc.meta["Chunker"])

# Output: {"type": "Chonkie", "num_chunks": 7, ...}

This metadata includes the chunker type and the number of generated chunks, enabling downstream tasks to adjust their behavior based on how the text was segmented.

Choosing Between Chonkie and NaiveChunker

Select your chunking strategy based on pipeline requirements:

  • Use Chonkie when you need token-level precision, overlapping windows for context preservation, or language-specific tokenization. This yields more semantically coherent chunks for embedding generation or LLM inference.
  • Use NaiveChunker for speed-critical applications, quick prototyping, or when your downstream model can handle arbitrary character boundaries without semantic degradation.

Summary

Frequently Asked Questions

How do I configure chunk overlap with the Chonkie chunker?

Pass the chunk_overlap parameter to your chonkie.TokenChunker instance before wrapping it in the Sieves task. For example, chonkie.TokenChunker(tokenizer, chunk_size=512, chunk_overlap=50) creates 512-token chunks with 50 tokens of overlap between consecutive segments.

Can I switch between chunkers without rewriting my entire pipeline?

Yes. Since tasks.Chunking() accepts either a chonkie.BaseChunker object or an integer, you can swap tasks.Chunking(my_chonkie_obj) with tasks.Chunking(200) to switch from token-aware to character-based chunking without changing any other pipeline components.

What tokenizer should I use with Chonkie in Sieves?

Use any tokenizer compatible with the chonkie library, typically loaded via chonkie.tokenizers.Tokenizer.from_pretrained(). The examples in the Sieves test suite use "gpt2" for demonstration, but you should match your tokenizer to the specific model that will process the chunks downstream.

Does NaiveChunker support overlapping windows like Chonkie?

No. The NaiveChunker implementation in sieves/tasks/preprocessing/chunking/naive.py performs simple interval-based slicing without overlap. If you need overlapping segments, you must use the Chonkie chunker with its chunk_overlap parameter configured.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →