# Entity-Aware Chunking for Preserving Named Entities in Semantica

> Semantica's entity-aware chunking keeps named entities intact for better GraphRAG and knowledge graph pipelines. Preserve semantic units for accurate relationship extraction.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-11

---

**Semantica's entity-aware chunking prevents named entity fragmentation by dynamically extending chunk boundaries to keep complete entities intact, ensuring GraphRAG and knowledge graph pipelines receive uncorrupted semantic units for relationship extraction.**

Entity-aware chunking is a specialized text-splitting strategy in the Semantica framework designed to maintain named entity integrity when processing large documents. This approach is essential for downstream GraphRAG implementations where fragmented entities—such as splitting "Barack Obama" into "Bar-" and "ack Obama"—would severely degrade entity resolution accuracy and knowledge graph construction quality.

## Architecture and Implementation

The implementation separates configuration concerns from algorithmic logic across two core modules. The public API resides in [`semantica/split/kg_chunkers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/kg_chunkers.py) within the `EntityAwareChunker` class (lines 62-99), while the core splitting algorithm lives in [`semantica/split/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/methods.py) as the `split_entity_aware` function (lines 951-981).

### The EntityAwareChunker Class

This class serves as the primary entry point for users configuring entity-aware splitting behavior. The constructor accepts four critical parameters:

- **`chunk_size`** – Target character length for each chunk
- **`chunk_overlap`** – Number of characters to overlap between consecutive chunks
- **`ner_method`** – The named entity recognition backend to employ (e.g., `"ml"`, `"huggingface"`, `"spacy"`)
- **`preserve_entities`** – Boolean flag enforcing strict entity boundary preservation

The class automatically retrieves a progress tracker, providing observability for long-running document splits across large corpora.

### The split_entity_aware Algorithm

The `split_entity_aware` function implements the boundary-preservation logic through a multi-stage pipeline:

1. **Entity Extraction** – An `NERExtractor` processes the entire input text upfront, generating entities with precise start/end character offsets
2. **Sentence Traversal** – The text iterates sentence-by-sentence to identify natural break points
3. **Boundary Extension** – When a sentence would exceed the target `chunk_size` and contains an entity that would be bisected, the algorithm extends the current chunk beyond the size limit rather than splitting the entity
4. **Metadata Enrichment** – Each output chunk includes metadata fields: `method: "entity_aware"`, `entity_count`, and the complete list of entity objects contained within

## Usage Examples

### Basic Python Implementation

```python
from semantica.split.kg_chunkers import EntityAwareChunker

# Sample text containing named entities

text = """
Apple announced its new products in California. Tim Cook said the iPhone 15 will be released next month.
"""

# Initialize chunker with 200-character target and ML-based NER

chunker = EntityAwareChunker(chunk_size=200, ner_method="ml")

# Execute the split

chunks = chunker.chunk(text)

# Review entity preservation

for i, chunk in enumerate(chunks, 1):
    print(f"Chunk {i}:")
    print(chunk.text)
    print("Entities:", [e.text for e in chunk.metadata["entities"]])

```

### Command-Line Interface

The CLI exposes entity-aware chunking through the `--method` flag in [`semantica/cli.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/cli.py):

```bash
semantica split --method entity-aware --chunk-size 1000 input_document.txt

```

### Programmatic Method Dispatch

For dynamic method selection, use the factory function in [`semantica/split/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/methods.py):

```python
from semantica.split.methods import get_split_method

split_fn = get_split_method("entity_aware")
chunks = split_fn(
    text,
    chunk_size=1000,
    ner_method="huggingface",
    preserve_entities=True,
)

```

## Integration Points and Fallback Behavior

Entity-aware chunking integrates throughout the Semantica ecosystem. The documentation in [`docs/reference/split.md`](https://github.com/semantica-agi/semantica/blob/main/docs/reference/split.md) (lines 9-15) explicitly lists this method among six supported strategies, noting its specific utility for "preserving entity boundaries" in knowledge graph workflows. The interactive cookbook at `cookbook/introduction/11_Chunking_and_Splitting.ipynb` provides practical scenarios demonstrating optimal selection criteria for KG pipelines.

The implementation includes graceful degradation: if the `semantic-extract` package is unavailable, the system automatically falls back to a standard recursive splitter, ensuring pipeline continuity without manual intervention.

## Summary

- **Entity-aware chunking** prevents named entity fragmentation by extending chunk boundaries when necessary, implemented in [`semantica/split/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/methods.py) as `split_entity_aware`
- The `EntityAwareChunker` class in [`semantica/split/kg_chunkers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/kg_chunkers.py) provides a configurable interface with `chunk_size`, `ner_method`, and `preserve_entities` parameters
- Each chunk retains rich metadata including entity counts and complete entity objects for downstream GraphRAG consumption
- The strategy is accessible via Python API, CLI (`--method entity-aware`), and programmatic dispatch functions
- Built-in fallback to recursive splitting ensures robustness when NER dependencies are unavailable

## Frequently Asked Questions

### What is entity-aware chunking and why does it matter for GraphRAG?

Entity-aware chunking is a document-splitting strategy that recognizes named entity boundaries and prevents splitting text mid-entity. It matters for GraphRAG because standard chunking algorithms often bisect entities (e.g., splitting "Barack Obama" across chunks), which corrupts entity resolution and prevents accurate relationship extraction in knowledge graphs.

### How does the algorithm handle entities larger than the specified chunk size?

When an entity exceeds the configured `chunk_size`, the `split_entity_aware` function extends the current chunk beyond the target limit to accommodate the complete entity. This boundary extension ensures semantic coherence, though users should monitor chunk size distributions if their corpus contains unusually long named entities.

### Which NER methods does Semantica support for entity-aware chunking?

According to the source code in [`semantica/split/kg_chunkers.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/kg_chunkers.py), the `ner_method` parameter accepts multiple backends including `"ml"` (default machine learning), `"huggingface"`, `"spacy"`, and pattern-based extractors. This modular design allows swapping NER implementations without modifying the chunking logic.

### What happens if the NER dependency is not installed?

The system implements graceful degradation in [`semantica/split/methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/methods.py). If the `semantic-extract` package is unavailable, the `split_entity_aware` function catches the import exception and falls back to a standard recursive text splitter, ensuring pipelines continue operating while logging the dependency limitation.