How LangExtract Processes Long Documents with Chunking and Parallel Workers
LangExtract processes long documents by splitting text into token-aware chunks, grouping them into batches, and executing parallel inference via ThreadPoolExecutor to maximize throughput while respecting LLM context limits.
LangExtract is an open-source Python library developed by Google that extracts structured data from unstructured text using large language models. When processing long documents such as books or lengthy reports, the library implements a sophisticated three-stage pipeline that combines chunking and parallel workers to efficiently handle texts exceeding typical LLM token limits.
The Three-Stage Processing Pipeline
Stage 1: Token-Aware Chunking
The pipeline begins in langextract/chunking.py, where the ChunkIterator class greedily constructs text chunks that respect both sentence boundaries and the max_char_buffer limit. The algorithm iterates through tokenized sentences, accumulating text until adding another sentence would exceed the character buffer, then yields a TextChunk containing the accumulated content plus optional context from the previous chunk.
Key implementation details from langextract/chunking.py lines 43-84:
- Uses
RegexTokenizerby default to identify sentence boundaries - Handles edge cases where a single token exceeds
max_char_bufferby creating a single-token chunk - Supports
context_window_charsto prepend characters from the previous chunk for cross-chunk coherence
Stage 2: Batch Formation
Once chunks are generated, the Annotator class in langextract/annotation.py groups consecutive chunks into batches. The _annotate_documents_single_pass method (lines 58-84) collects chunks until reaching batch_length, then converts each chunk into a structured prompt using PromptTemplateStructured.
Batch formation serves two purposes:
- API Efficiency: Grouping prompts allows providers to use batch APIs when available
- Memory Management: Processing batches sequentially prevents memory exhaustion on extremely large documents
Stage 3: Parallel Inference
The final stage executes prompts concurrently using worker pools. When a batch contains multiple prompts and max_workers > 1, providers spawn a ThreadPoolExecutor to parallelize inference.
Implementation in langextract/providers/gemini.py (lines 7-14):
# Simplified representation of the parallel execution logic
with ThreadPoolExecutor(max_workers=min(self.max_workers, len(batch_prompts))) as executor:
futures = [executor.submit(self._process_single_prompt, prompt) for prompt in batch_prompts]
results = [f.result() for f in futures]
The OpenAI provider follows an identical pattern in langextract/providers/openai.py (lines 7-14), ensuring consistent parallelism across different LLM backends.
Configuration Parameters for Chunking and Parallelism
LangExtract exposes several parameters in the extract() function (defined in langextract/extraction.py) to tune the pipeline:
| Parameter | Default | Description |
|---|---|---|
max_char_buffer |
1000 | Maximum characters per chunk before splitting |
batch_length |
10 | Number of chunks to group into a single batch |
max_workers |
10 | Maximum parallel threads for inference |
context_window_chars |
0 | Characters from previous chunk to prepend for context |
extraction_passes |
1 | Number of times to run the entire pipeline over the text |
Note: When batch_length < max_workers, langextract/extraction.py emits a warning (lines 3-8) because the configuration underutilizes available worker capacity.
Practical Code Examples
Basic Extraction with Explicit Chunking and Parallelism
import langextract as lx
import textwrap
prompt = textwrap.dedent(
"""\
Extract all character names and their emotions from the story.
Return each extraction as a JSON object with fields "character" and "emotion".
"""
)
examples = [
lx.data.ExampleData(
text="ROMEO: But soft! What light through yonder window breaks?",
extractions=[
lx.data.Extraction(
extraction_class="character",
extraction_text="ROMEO",
attributes={"emotion": "wonder"},
)
],
)
]
result = lx.extract(
text_or_documents="https://www.gutenberg.org/files/1513/1513-0.txt", # full Romeo & Juliet
prompt_description=prompt,
examples=examples,
max_char_buffer=1200, # each LLM call sees ≤ 1200 chars
batch_length=20, # group 20 chunks per batch
max_workers=8, # run up to 8 prompts in parallel
extraction_passes=1, # single pass (increase for higher recall)
show_progress=False,
)
Under the hood, this configuration processes the entire play by downloading the text via io.download_text_from_url, creating approximately 10,000 chunks using ChunkIterator with 1200-character buffers, grouping chunks into batches of 20, and executing each batch with 8 parallel workers via ThreadPoolExecutor.
Multi-Pass Extraction for Higher Recall
result = lx.extract(
text_or_documents=long_text,
prompt_description=prompt,
examples=examples,
max_char_buffer=800,
batch_length=15,
max_workers=5,
extraction_passes=3, # run the whole pipeline three times
)
Each pass reprocesses the same chunks independently, allowing the model to discover entities missed in earlier iterations. The Resolver merges results across passes, preserving the first-found annotation for overlapping entities.
Adding Cross-Chunk Context
result = lx.extract(
text_or_documents=long_text,
prompt_description=prompt,
examples=examples,
max_char_buffer=1000,
batch_length=10,
max_workers=4,
context_window_chars=200, # prepend last 200 chars from previous chunk
)
The Annotator copies the final 200 characters from the preceding chunk into the current prompt, helping the model resolve anaphoric references like "she" or "it" that depend on earlier context.
Key Source Files and Implementation Details
Understanding the internal architecture helps debug performance bottlenecks:
| File | Primary Responsibility | Notable Sections |
|---|---|---|
langextract/chunking.py |
Token-aware text segmentation | ChunkIterator class (lines 43-84) implements the greedy chunking algorithm with sentence boundary detection |
langextract/annotation.py |
Batch orchestration and prompt assembly | _annotate_documents_single_pass (lines 58-84) groups chunks and manages the inference loop |
langextract/providers/gemini.py |
Gemini-specific parallel execution | Lines 7-14 contain the ThreadPoolExecutor logic that caps workers at min(max_workers, len(batch_prompts)) |
langextract/providers/openai.py |
OpenAI-specific parallel execution | Lines 7-14 mirror the Gemini pattern for consistent cross-provider parallelism |
langextract/extraction.py |
Public API and configuration validation | Lines 3-8 emit warnings when batch_length < max_workers; lines 34-47 handle model instantiation |
These files implement the "chunk → batch → parallel inference" workflow that enables LangExtract to efficiently process megabytes of text while staying within LLM token limits.
Summary
LangExtract efficiently processes long documents through a three-stage pipeline that balances memory constraints with parallel throughput:
- Token-aware chunking in
langextract/chunking.pysplits documents intomax_char_buffer-sized segments while preserving sentence boundaries, with optionalcontext_window_charsfor cross-chunk coherence. - Batch formation in
langextract/annotation.pygroups consecutive chunks intobatch_lengthsized collections to optimize API usage and memory consumption. - Parallel inference via
ThreadPoolExecutorin bothlangextract/providers/gemini.pyandlangextract/providers/openai.pyexecutes up tomax_workersprompts concurrently, automatically scaling down when batches are smaller than the worker limit.
Configure the pipeline using max_char_buffer, batch_length, max_workers, and extraction_passes to tune the trade-off between latency, cost, and recall when processing large-scale text corpora.
Frequently Asked Questions
What is the optimal batch_length to max_workers ratio for LangExtract?
For maximum throughput, set batch_length equal to or greater than max_workers. When batch_length is smaller than max_workers, the system emits a warning in langextract/extraction.py because you have idle threads that cannot be utilized within a single batch. However, if memory constraints force small batches, the pipeline still functions correctly but underutilizes available parallelism.
How does LangExtract handle chunks that exceed max_char_buffer?
The ChunkIterator in langextract/chunking.py implements a greedy algorithm that checks token sizes before adding sentences to the current chunk. If a single token (sentence) exceeds max_char_buffer, the system creates a chunk containing only that token rather than splitting it mid-sentence, ensuring semantic coherence is preserved even with aggressive size constraints.
Can I use different chunking strategies with LangExtract?
Currently, LangExtract uses the built-in RegexTokenizer and ChunkIterator defined in langextract/chunking.py. While the library does not expose pluggable chunking strategies in the public API, you can influence behavior through max_char_buffer and context_window_chars parameters. The context_window_chars option prepends previous context to each chunk, effectively creating overlapping sliding windows without changing the underlying segmentation logic.
Does parallel execution work with both Gemini and OpenAI providers?
Yes, both providers implement identical parallel execution patterns. The Gemini class in langextract/providers/gemini.py and the OpenAI class in langextract/providers/openai.py both use ThreadPoolExecutor to run _process_single_prompt across multiple threads when max_workers > 1 and the batch contains multiple prompts. The executor size is automatically capped at min(self.max_workers, len(batch_prompts)) to prevent spawning unnecessary threads.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →