# How to Use Text Chunking with Chonkie vs NaiveChunker in Sieves

> Master text chunking in Sieves with Chonkie vs NaiveChunker. Learn token-aware splitting and character-interval slicing using the Chunking task for efficient data processing.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Sieves provides two built-in chunking strategies—Chonkie for token-aware splitting and NaiveChunker for character-interval slicing—accessible through the unified `Chunking` task or individual task aliases.**

The `mantisai/sieves` repository offers a flexible preprocessing module that handles text segmentation for downstream NLP pipelines. Whether you need linguistically coherent chunks that respect token boundaries or fast character-based slicing, Sieves abstracts both approaches behind a consistent API defined in [`sieves/tasks/preprocessing/chunking/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/core.py).

## Understanding the Two Chunking Strategies

Sieves implements two distinct chunking philosophies:

- **Chonkie**: A wrapper around the third-party `chonkie` library that provides token-aware chunking with configurable overlap, ideal for LLM context windows.
- **NaiveChunker**: A lightweight, interval-based splitter that divides text every *n* characters without considering token boundaries.

Both are exposed through the high-level `Chunking` task, which accepts either a `chonkie.BaseChunker` instance or an integer to determine which implementation to instantiate.

## Using the Chonkie Chunker for Token-Aware Splitting

### Implementation Details

The Chonkie integration lives in [`sieves/tasks/preprocessing/chunking/chonkie_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/chonkie_.py) and is invoked when you pass a `chonkie.BaseChunker` to the task constructor. According to the source in [`core.py`](https://github.com/mantisai/sieves/blob/main/core.py) (lines 30-68), the logic branches based on type:

```python

# From sieves/tasks/preprocessing/chunking/core.py

if isinstance(chunker, chonkie.BaseChunker):
    chunker_task = chonkie_.Chonkie(chunker=chunker)   # Chonkie path

elif isinstance(chunker, int):
    chunker_task = naive.NaiveChunker(interval=chunker)   # Naive path

```

### Practical Example with TokenChunker

To create semantically coherent chunks using GPT-2 tokenization:

```python
from sieves import Pipeline, tasks
import chonkie
from chonkie import tokenizers

tokenizer = tokenizers.Tokenizer.from_pretrained("gpt2")
chonkie_chunker = chonkie.TokenChunker(
    tokenizer, 
    chunk_size=512, 
    chunk_overlap=50
)

pipeline = Pipeline([tasks.Chunking(chonkie_chunker)])
doc = pipeline([sieves.Doc(text="Your very long document text here...")])[0]
print(doc.meta["Chunker"])

```

This configuration respects token boundaries, maintains a 50-token overlap between chunks, and stores metadata about the chunking operation.

## Using the NaiveChunker for Character-Interval Splitting

### Implementation Details

The `NaiveChunker` class in [`sieves/tasks/preprocessing/chunking/naive.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/naive.py) provides deterministic, high-speed text segmentation. When the `Chunking` task receives an integer argument, it automatically instantiates this class using the integer as the character interval.

### Practical Example

For quick prototyping or when token boundaries are irrelevant:

```python
from sieves import Pipeline, tasks

pipeline = Pipeline([tasks.Chunking(200)])  # 200-character slices

# Or explicitly:

pipeline = Pipeline([tasks.NaiveChunker(interval=200)])

doc = pipeline([sieves.Doc(text="Another long document...")])[0]
print(doc.meta["Chunker"])

```

This approach cuts the text every 200 characters regardless of word or token boundaries, offering maximum speed with minimal overhead.

## How Sieves Resolves Chunker Types Internally

The unified interface is implemented in [`sieves/tasks/preprocessing/chunking/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/core.py) through the `_ChunkerArgType` union type:

```python
_ChunkerArgType = chonkie.BaseChunker | int

def __init__(self, chunker: _ChunkerArgType, ...):
    if isinstance(chunker, chonkie.BaseChunker):
        self._chunker = chonkie_.Chonkie(chunker=chunker)
    elif isinstance(chunker, int):
        self._chunker = naive.NaiveChunker(interval=chunker)
    else:
        raise TypeError("chunker must be a chonkie.BaseChunker or int")

```

This type-checking mechanism allows the same `Chunking` task to serve both sophisticated token-aware workflows and simple character-splitting needs without requiring separate import paths.

## Inspecting Chunking Metadata

Both chunkers record execution metadata in the document's meta dictionary. As shown in the test suite ([`sieves/tests/docs/test_preprocessing.py`](https://github.com/mantisai/sieves/blob/main/sieves/tests/docs/test_preprocessing.py), lines 149-155), you can inspect the results after pipeline execution:

```python
doc = pipe([sieves.Doc(text="Your long document")])[0]
print(doc.meta["Chunker"])

# Output: {"type": "Chonkie", "num_chunks": 7, ...}

```

This metadata includes the chunker type and the number of generated chunks, enabling downstream tasks to adjust their behavior based on how the text was segmented.

## Choosing Between Chonkie and NaiveChunker

Select your chunking strategy based on pipeline requirements:

- **Use Chonkie** when you need token-level precision, overlapping windows for context preservation, or language-specific tokenization. This yields more semantically coherent chunks for embedding generation or LLM inference.
- **Use NaiveChunker** for speed-critical applications, quick prototyping, or when your downstream model can handle arbitrary character boundaries without semantic degradation.

## Summary

- **Chonkie** wraps the external `chonkie` library and requires a tokenizer and chunk size configuration for token-aware splitting.
- **NaiveChunker** provides fast character-interval splitting when you pass an integer to the `Chunking` task constructor.
- The unified `Chunking` task in [`sieves/tasks/preprocessing/chunking/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/core.py) automatically routes to the appropriate implementation based on argument type.
- Both strategies populate `doc.meta["Chunker"]` with execution metadata including chunk count and chunker type.
- Import paths: [`sieves/tasks/preprocessing/chunking/chonkie_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/chonkie_.py) for Chonkie logic, [`sieves/tasks/preprocessing/chunking/naive.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/naive.py) for the naive implementation.

## Frequently Asked Questions

### How do I configure chunk overlap with the Chonkie chunker?

Pass the `chunk_overlap` parameter to your `chonkie.TokenChunker` instance before wrapping it in the Sieves task. For example, `chonkie.TokenChunker(tokenizer, chunk_size=512, chunk_overlap=50)` creates 512-token chunks with 50 tokens of overlap between consecutive segments.

### Can I switch between chunkers without rewriting my entire pipeline?

Yes. Since `tasks.Chunking()` accepts either a `chonkie.BaseChunker` object or an integer, you can swap `tasks.Chunking(my_chonkie_obj)` with `tasks.Chunking(200)` to switch from token-aware to character-based chunking without changing any other pipeline components.

### What tokenizer should I use with Chonkie in Sieves?

Use any tokenizer compatible with the `chonkie` library, typically loaded via `chonkie.tokenizers.Tokenizer.from_pretrained()`. The examples in the Sieves test suite use `"gpt2"` for demonstration, but you should match your tokenizer to the specific model that will process the chunks downstream.

### Does NaiveChunker support overlapping windows like Chonkie?

No. The `NaiveChunker` implementation in [`sieves/tasks/preprocessing/chunking/naive.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/chunking/naive.py) performs simple interval-based slicing without overlap. If you need overlapping segments, you must use the Chonkie chunker with its `chunk_overlap` parameter configured.