Entity-Aware Chunking for Preserving Named Entities in Semantica
Semantica's entity-aware chunking prevents named entity fragmentation by dynamically extending chunk boundaries to keep complete entities intact, ensuring GraphRAG and knowledge graph pipelines receive uncorrupted semantic units for relationship extraction.
Entity-aware chunking is a specialized text-splitting strategy in the Semantica framework designed to maintain named entity integrity when processing large documents. This approach is essential for downstream GraphRAG implementations where fragmented entities—such as splitting "Barack Obama" into "Bar-" and "ack Obama"—would severely degrade entity resolution accuracy and knowledge graph construction quality.
Architecture and Implementation
The implementation separates configuration concerns from algorithmic logic across two core modules. The public API resides in semantica/split/kg_chunkers.py within the EntityAwareChunker class (lines 62-99), while the core splitting algorithm lives in semantica/split/methods.py as the split_entity_aware function (lines 951-981).
The EntityAwareChunker Class
This class serves as the primary entry point for users configuring entity-aware splitting behavior. The constructor accepts four critical parameters:
chunk_size– Target character length for each chunkchunk_overlap– Number of characters to overlap between consecutive chunksner_method– The named entity recognition backend to employ (e.g.,"ml","huggingface","spacy")preserve_entities– Boolean flag enforcing strict entity boundary preservation
The class automatically retrieves a progress tracker, providing observability for long-running document splits across large corpora.
The split_entity_aware Algorithm
The split_entity_aware function implements the boundary-preservation logic through a multi-stage pipeline:
- Entity Extraction – An
NERExtractorprocesses the entire input text upfront, generating entities with precise start/end character offsets - Sentence Traversal – The text iterates sentence-by-sentence to identify natural break points
- Boundary Extension – When a sentence would exceed the target
chunk_sizeand contains an entity that would be bisected, the algorithm extends the current chunk beyond the size limit rather than splitting the entity - Metadata Enrichment – Each output chunk includes metadata fields:
method: "entity_aware",entity_count, and the complete list of entity objects contained within
Usage Examples
Basic Python Implementation
from semantica.split.kg_chunkers import EntityAwareChunker
# Sample text containing named entities
text = """
Apple announced its new products in California. Tim Cook said the iPhone 15 will be released next month.
"""
# Initialize chunker with 200-character target and ML-based NER
chunker = EntityAwareChunker(chunk_size=200, ner_method="ml")
# Execute the split
chunks = chunker.chunk(text)
# Review entity preservation
for i, chunk in enumerate(chunks, 1):
print(f"Chunk {i}:")
print(chunk.text)
print("Entities:", [e.text for e in chunk.metadata["entities"]])
Command-Line Interface
The CLI exposes entity-aware chunking through the --method flag in semantica/cli.py:
semantica split --method entity-aware --chunk-size 1000 input_document.txt
Programmatic Method Dispatch
For dynamic method selection, use the factory function in semantica/split/methods.py:
from semantica.split.methods import get_split_method
split_fn = get_split_method("entity_aware")
chunks = split_fn(
text,
chunk_size=1000,
ner_method="huggingface",
preserve_entities=True,
)
Integration Points and Fallback Behavior
Entity-aware chunking integrates throughout the Semantica ecosystem. The documentation in docs/reference/split.md (lines 9-15) explicitly lists this method among six supported strategies, noting its specific utility for "preserving entity boundaries" in knowledge graph workflows. The interactive cookbook at cookbook/introduction/11_Chunking_and_Splitting.ipynb provides practical scenarios demonstrating optimal selection criteria for KG pipelines.
The implementation includes graceful degradation: if the semantic-extract package is unavailable, the system automatically falls back to a standard recursive splitter, ensuring pipeline continuity without manual intervention.
Summary
- Entity-aware chunking prevents named entity fragmentation by extending chunk boundaries when necessary, implemented in
semantica/split/methods.pyassplit_entity_aware - The
EntityAwareChunkerclass insemantica/split/kg_chunkers.pyprovides a configurable interface withchunk_size,ner_method, andpreserve_entitiesparameters - Each chunk retains rich metadata including entity counts and complete entity objects for downstream GraphRAG consumption
- The strategy is accessible via Python API, CLI (
--method entity-aware), and programmatic dispatch functions - Built-in fallback to recursive splitting ensures robustness when NER dependencies are unavailable
Frequently Asked Questions
What is entity-aware chunking and why does it matter for GraphRAG?
Entity-aware chunking is a document-splitting strategy that recognizes named entity boundaries and prevents splitting text mid-entity. It matters for GraphRAG because standard chunking algorithms often bisect entities (e.g., splitting "Barack Obama" across chunks), which corrupts entity resolution and prevents accurate relationship extraction in knowledge graphs.
How does the algorithm handle entities larger than the specified chunk size?
When an entity exceeds the configured chunk_size, the split_entity_aware function extends the current chunk beyond the target limit to accommodate the complete entity. This boundary extension ensures semantic coherence, though users should monitor chunk size distributions if their corpus contains unusually long named entities.
Which NER methods does Semantica support for entity-aware chunking?
According to the source code in semantica/split/kg_chunkers.py, the ner_method parameter accepts multiple backends including "ml" (default machine learning), "huggingface", "spacy", and pattern-based extractors. This modular design allows swapping NER implementations without modifying the chunking logic.
What happens if the NER dependency is not installed?
The system implements graceful degradation in semantica/split/methods.py. If the semantic-extract package is unavailable, the split_entity_aware function catches the import exception and falls back to a standard recursive text splitter, ensuring pipelines continue operating while logging the dependency limitation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →