Entity-Aware Chunking for Preserving Named Entities in Semantica

Semantica's entity-aware chunking prevents named entity fragmentation by dynamically extending chunk boundaries to keep complete entities intact, ensuring GraphRAG and knowledge graph pipelines receive uncorrupted semantic units for relationship extraction.

Entity-aware chunking is a specialized text-splitting strategy in the Semantica framework designed to maintain named entity integrity when processing large documents. This approach is essential for downstream GraphRAG implementations where fragmented entities—such as splitting "Barack Obama" into "Bar-" and "ack Obama"—would severely degrade entity resolution accuracy and knowledge graph construction quality.

Architecture and Implementation

The implementation separates configuration concerns from algorithmic logic across two core modules. The public API resides in semantica/split/kg_chunkers.py within the EntityAwareChunker class (lines 62-99), while the core splitting algorithm lives in semantica/split/methods.py as the split_entity_aware function (lines 951-981).

The EntityAwareChunker Class

This class serves as the primary entry point for users configuring entity-aware splitting behavior. The constructor accepts four critical parameters:

  • chunk_size – Target character length for each chunk
  • chunk_overlap – Number of characters to overlap between consecutive chunks
  • ner_method – The named entity recognition backend to employ (e.g., "ml", "huggingface", "spacy")
  • preserve_entities – Boolean flag enforcing strict entity boundary preservation

The class automatically retrieves a progress tracker, providing observability for long-running document splits across large corpora.

The split_entity_aware Algorithm

The split_entity_aware function implements the boundary-preservation logic through a multi-stage pipeline:

  1. Entity Extraction – An NERExtractor processes the entire input text upfront, generating entities with precise start/end character offsets
  2. Sentence Traversal – The text iterates sentence-by-sentence to identify natural break points
  3. Boundary Extension – When a sentence would exceed the target chunk_size and contains an entity that would be bisected, the algorithm extends the current chunk beyond the size limit rather than splitting the entity
  4. Metadata Enrichment – Each output chunk includes metadata fields: method: "entity_aware", entity_count, and the complete list of entity objects contained within

Usage Examples

Basic Python Implementation

from semantica.split.kg_chunkers import EntityAwareChunker

# Sample text containing named entities

text = """
Apple announced its new products in California. Tim Cook said the iPhone 15 will be released next month.
"""

# Initialize chunker with 200-character target and ML-based NER

chunker = EntityAwareChunker(chunk_size=200, ner_method="ml")

# Execute the split

chunks = chunker.chunk(text)

# Review entity preservation

for i, chunk in enumerate(chunks, 1):
    print(f"Chunk {i}:")
    print(chunk.text)
    print("Entities:", [e.text for e in chunk.metadata["entities"]])

Command-Line Interface

The CLI exposes entity-aware chunking through the --method flag in semantica/cli.py:

semantica split --method entity-aware --chunk-size 1000 input_document.txt

Programmatic Method Dispatch

For dynamic method selection, use the factory function in semantica/split/methods.py:

from semantica.split.methods import get_split_method

split_fn = get_split_method("entity_aware")
chunks = split_fn(
    text,
    chunk_size=1000,
    ner_method="huggingface",
    preserve_entities=True,
)

Integration Points and Fallback Behavior

Entity-aware chunking integrates throughout the Semantica ecosystem. The documentation in docs/reference/split.md (lines 9-15) explicitly lists this method among six supported strategies, noting its specific utility for "preserving entity boundaries" in knowledge graph workflows. The interactive cookbook at cookbook/introduction/11_Chunking_and_Splitting.ipynb provides practical scenarios demonstrating optimal selection criteria for KG pipelines.

The implementation includes graceful degradation: if the semantic-extract package is unavailable, the system automatically falls back to a standard recursive splitter, ensuring pipeline continuity without manual intervention.

Summary

  • Entity-aware chunking prevents named entity fragmentation by extending chunk boundaries when necessary, implemented in semantica/split/methods.py as split_entity_aware
  • The EntityAwareChunker class in semantica/split/kg_chunkers.py provides a configurable interface with chunk_size, ner_method, and preserve_entities parameters
  • Each chunk retains rich metadata including entity counts and complete entity objects for downstream GraphRAG consumption
  • The strategy is accessible via Python API, CLI (--method entity-aware), and programmatic dispatch functions
  • Built-in fallback to recursive splitting ensures robustness when NER dependencies are unavailable

Frequently Asked Questions

What is entity-aware chunking and why does it matter for GraphRAG?

Entity-aware chunking is a document-splitting strategy that recognizes named entity boundaries and prevents splitting text mid-entity. It matters for GraphRAG because standard chunking algorithms often bisect entities (e.g., splitting "Barack Obama" across chunks), which corrupts entity resolution and prevents accurate relationship extraction in knowledge graphs.

How does the algorithm handle entities larger than the specified chunk size?

When an entity exceeds the configured chunk_size, the split_entity_aware function extends the current chunk beyond the target limit to accommodate the complete entity. This boundary extension ensures semantic coherence, though users should monitor chunk size distributions if their corpus contains unusually long named entities.

Which NER methods does Semantica support for entity-aware chunking?

According to the source code in semantica/split/kg_chunkers.py, the ner_method parameter accepts multiple backends including "ml" (default machine learning), "huggingface", "spacy", and pattern-based extractors. This modular design allows swapping NER implementations without modifying the chunking logic.

What happens if the NER dependency is not installed?

The system implements graceful degradation in semantica/split/methods.py. If the semantic-extract package is unavailable, the split_entity_aware function catches the import exception and falls back to a standard recursive text splitter, ensuring pipelines continue operating while logging the dependency limitation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →