How to Use Text Chunking with Chonkie vs NaiveChunker in Sieves
Sieves provides two built-in chunking strategies—Chonkie for token-aware splitting and NaiveChunker for character-interval slicing—accessible through the unified Chunking task or individual task aliases.
The mantisai/sieves repository offers a flexible preprocessing module that handles text segmentation for downstream NLP pipelines. Whether you need linguistically coherent chunks that respect token boundaries or fast character-based slicing, Sieves abstracts both approaches behind a consistent API defined in sieves/tasks/preprocessing/chunking/core.py.
Understanding the Two Chunking Strategies
Sieves implements two distinct chunking philosophies:
- Chonkie: A wrapper around the third-party
chonkielibrary that provides token-aware chunking with configurable overlap, ideal for LLM context windows. - NaiveChunker: A lightweight, interval-based splitter that divides text every n characters without considering token boundaries.
Both are exposed through the high-level Chunking task, which accepts either a chonkie.BaseChunker instance or an integer to determine which implementation to instantiate.
Using the Chonkie Chunker for Token-Aware Splitting
Implementation Details
The Chonkie integration lives in sieves/tasks/preprocessing/chunking/chonkie_.py and is invoked when you pass a chonkie.BaseChunker to the task constructor. According to the source in core.py (lines 30-68), the logic branches based on type:
# From sieves/tasks/preprocessing/chunking/core.py
if isinstance(chunker, chonkie.BaseChunker):
chunker_task = chonkie_.Chonkie(chunker=chunker) # Chonkie path
elif isinstance(chunker, int):
chunker_task = naive.NaiveChunker(interval=chunker) # Naive path
Practical Example with TokenChunker
To create semantically coherent chunks using GPT-2 tokenization:
from sieves import Pipeline, tasks
import chonkie
from chonkie import tokenizers
tokenizer = tokenizers.Tokenizer.from_pretrained("gpt2")
chonkie_chunker = chonkie.TokenChunker(
tokenizer,
chunk_size=512,
chunk_overlap=50
)
pipeline = Pipeline([tasks.Chunking(chonkie_chunker)])
doc = pipeline([sieves.Doc(text="Your very long document text here...")])[0]
print(doc.meta["Chunker"])
This configuration respects token boundaries, maintains a 50-token overlap between chunks, and stores metadata about the chunking operation.
Using the NaiveChunker for Character-Interval Splitting
Implementation Details
The NaiveChunker class in sieves/tasks/preprocessing/chunking/naive.py provides deterministic, high-speed text segmentation. When the Chunking task receives an integer argument, it automatically instantiates this class using the integer as the character interval.
Practical Example
For quick prototyping or when token boundaries are irrelevant:
from sieves import Pipeline, tasks
pipeline = Pipeline([tasks.Chunking(200)]) # 200-character slices
# Or explicitly:
pipeline = Pipeline([tasks.NaiveChunker(interval=200)])
doc = pipeline([sieves.Doc(text="Another long document...")])[0]
print(doc.meta["Chunker"])
This approach cuts the text every 200 characters regardless of word or token boundaries, offering maximum speed with minimal overhead.
How Sieves Resolves Chunker Types Internally
The unified interface is implemented in sieves/tasks/preprocessing/chunking/core.py through the _ChunkerArgType union type:
_ChunkerArgType = chonkie.BaseChunker | int
def __init__(self, chunker: _ChunkerArgType, ...):
if isinstance(chunker, chonkie.BaseChunker):
self._chunker = chonkie_.Chonkie(chunker=chunker)
elif isinstance(chunker, int):
self._chunker = naive.NaiveChunker(interval=chunker)
else:
raise TypeError("chunker must be a chonkie.BaseChunker or int")
This type-checking mechanism allows the same Chunking task to serve both sophisticated token-aware workflows and simple character-splitting needs without requiring separate import paths.
Inspecting Chunking Metadata
Both chunkers record execution metadata in the document's meta dictionary. As shown in the test suite (sieves/tests/docs/test_preprocessing.py, lines 149-155), you can inspect the results after pipeline execution:
doc = pipe([sieves.Doc(text="Your long document")])[0]
print(doc.meta["Chunker"])
# Output: {"type": "Chonkie", "num_chunks": 7, ...}
This metadata includes the chunker type and the number of generated chunks, enabling downstream tasks to adjust their behavior based on how the text was segmented.
Choosing Between Chonkie and NaiveChunker
Select your chunking strategy based on pipeline requirements:
- Use Chonkie when you need token-level precision, overlapping windows for context preservation, or language-specific tokenization. This yields more semantically coherent chunks for embedding generation or LLM inference.
- Use NaiveChunker for speed-critical applications, quick prototyping, or when your downstream model can handle arbitrary character boundaries without semantic degradation.
Summary
- Chonkie wraps the external
chonkielibrary and requires a tokenizer and chunk size configuration for token-aware splitting. - NaiveChunker provides fast character-interval splitting when you pass an integer to the
Chunkingtask constructor. - The unified
Chunkingtask insieves/tasks/preprocessing/chunking/core.pyautomatically routes to the appropriate implementation based on argument type. - Both strategies populate
doc.meta["Chunker"]with execution metadata including chunk count and chunker type. - Import paths:
sieves/tasks/preprocessing/chunking/chonkie_.pyfor Chonkie logic,sieves/tasks/preprocessing/chunking/naive.pyfor the naive implementation.
Frequently Asked Questions
How do I configure chunk overlap with the Chonkie chunker?
Pass the chunk_overlap parameter to your chonkie.TokenChunker instance before wrapping it in the Sieves task. For example, chonkie.TokenChunker(tokenizer, chunk_size=512, chunk_overlap=50) creates 512-token chunks with 50 tokens of overlap between consecutive segments.
Can I switch between chunkers without rewriting my entire pipeline?
Yes. Since tasks.Chunking() accepts either a chonkie.BaseChunker object or an integer, you can swap tasks.Chunking(my_chonkie_obj) with tasks.Chunking(200) to switch from token-aware to character-based chunking without changing any other pipeline components.
What tokenizer should I use with Chonkie in Sieves?
Use any tokenizer compatible with the chonkie library, typically loaded via chonkie.tokenizers.Tokenizer.from_pretrained(). The examples in the Sieves test suite use "gpt2" for demonstration, but you should match your tokenizer to the specific model that will process the chunks downstream.
Does NaiveChunker support overlapping windows like Chonkie?
No. The NaiveChunker implementation in sieves/tasks/preprocessing/chunking/naive.py performs simple interval-based slicing without overlap. If you need overlapping segments, you must use the Chonkie chunker with its chunk_overlap parameter configured.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →