Best Practices for Large-Scale Dataset Processing with max_parallel_insert in LightRAG
Set max_parallel_insert between 2 and 8 based on your CPU, GPU, and storage backend limits, starting with the default value of 2 and scaling up only after profiling system resource usage.
LightRAG is an open-source retrieval-augmented generation framework designed for efficient knowledge graph construction from massive text corpora. When ingesting large-scale datasets, the max_parallel_insert parameter serves as a critical concurrency throttle that prevents resource exhaustion while maintaining optimal throughput. Proper tuning of this semaphore-based control ensures you maximize ingestion speed without overwhelming embedding models, vector stores, or memory constraints.
Understanding max_parallel_insert in LightRAG
Source Code Definition and Defaults
The max_parallel_insert parameter appears in three critical locations within the LightRAG codebase. In lightrag/lightrag.py, it is defined as a dataclass field that reads from environment variables: int(os.getenv("MAX_PARALLEL_INSERT", DEFAULT_MAX_PARALLEL_INSERT)) [link]. The fallback constant DEFAULT_MAX_PARALLEL_INSERT = 2 resides in lightrag/constants.py [link], while the CLI and environment handling in lightrag/api/config.py maps the variable using args.max_parallel_insert = get_env_value("MAX_PARALLEL_INSERT", 2, int) [link].
Concurrency Control Mechanism
During the document insertion phase, LightRAG converts this integer into an asyncio.Semaphore that caps concurrent document processing. The implementation in lightrag/lightrag.py (lines 1814-1820) creates the semaphore and acquires it for each document:
semaphore = asyncio.Semaphore(self.max_parallel_insert)
...
async with semaphore:
# extract / embed / store a single document
Because the semaphore applies once per document, it directly controls how many files undergo chunking, embedding, and storage operations simultaneously.
Resource Optimization Strategies
Tuning max_parallel_insert addresses five specific bottlenecks that emerge during large-scale ingestion:
-
CPU and GPU saturation – Embedding many chunks concurrently can overwhelm the model worker; lowering the value ensures only a few embedding calls remain pending.
-
Storage backend limits – PostgreSQL, Qdrant, and OpenSearch impose connection limits; keeping
max_parallel_insertbelow your database'smax_connectionsprevents "too many connections" errors. -
I/O thrashing – Reading large PDFs or remote files can flood disk and network queues; modest concurrency (2-4) maintains steady throughput without thrashing.
-
Memory pressure – Each coroutine retains chunk objects until flush; capping concurrency reduces peak RAM usage, especially with large
chunk_sizevalues. -
Graceful cancellation – LightRAG checks
pipeline_status.get("cancellation_requested")within bounded active tasks, ensuring responsive shutdown even during million-document ingests.
Practical Configuration Guidelines
Follow these evidence-based recommendations when scaling your LightRAG ingestion pipeline:
-
Start with the default (
2) and profile – Run a pilot batch and monitor CPU, GPU, and storage latency through log entries likeExtracting stage …before increasing concurrency. -
Scale incrementally based on utilization – Increase to
4or8only when CPU/GPU usage remains below 50% and storage latency stays low. -
Respect storage connection limits – For Qdrant with default
max_connections = 10, maintainmax_parallel_insert ≤ 6to preserve headroom for other operations. -
Coordinate with embedding batch size – The default
DEFAULT_EMBEDDING_BATCH_NUM = 10inconstants.pydefines chunk grouping; increasemax_parallel_insertonly after confirming your embedding endpoint handles the combined payload. -
Prefer environment variables for deployment – Set
MAX_PARALLEL_INSERT=4in your.envfile or Docker configuration to maintain reproducible deployments without code changes. -
Maintain hierarchy with max_async – Keep
max_parallel_insert≤max_asyncto prevent deadlock scenarios, asmax_asyncgoverns other async-heavy operations like reranking. -
Monitor semaphore saturation – Log
semaphore._value(available slots) during production runs to detect when the pipeline constantly hits its concurrency ceiling.
Implementation Examples
Basic Instantiation
Configure parallelism directly through the constructor to override environment defaults:
from lightrag.lightrag import LightRAG
rag = LightRAG(
working_dir="my_workdir",
max_parallel_insert=6,
embedding_func=my_embedding,
)
rag.insert(docs=my_large_dataset)
The constructor stores the value in self.max_parallel_insert, which the insert pipeline transforms into an asyncio.Semaphore during document processing.
Environment Variable Configuration
For CI/CD pipelines and containerized deployments, use environment variables:
# .env file
MAX_PARALLEL_INSERT=4
# No code modification required
rag = LightRAG(working_dir="my_workdir")
Performance Monitoring
Measure the concrete impact of concurrency adjustments:
import time
import psutil
from lightrag.lightrag import LightRAG
rag = LightRAG(max_parallel_insert=8)
start = time.time()
rag.insert(docs=big_corpus)
duration = time.time() - start
print(f"Insert completed in {duration:.1f}s")
print("Peak RAM:", psutil.Process().memory_info().rss / 1e9, "GB")
Typical observations show that max_parallel_insert=2 yields longer wall-time but minimal RAM usage, while increasing to 8 improves speed by approximately 30% at the cost of proportional memory spikes.
Handling Cancellation
The bounded concurrency enables responsive cancellation during long-running ingests:
import asyncio
from lightrag.lightrag import LightRAG
rag = LightRAG(max_parallel_insert=4)
async def run_with_timeout():
insert_task = asyncio.create_task(rag.insert(docs=huge_set))
await asyncio.sleep(10)
rag.cancel() # Flips pipeline_status["cancellation_requested"]
await insert_task
asyncio.run(run_with_timeout())
The cancellation flag is checked inside the process_document coroutine against pipeline_status.get("cancellation_requested"), ensuring the semaphore releases correctly during shutdown.
Key Source Files
Understanding the implementation requires familiarity with these specific files:
-
lightrag/lightrag.py– Contains the dataclass definition, semaphore creation (lines 1814-1820), and the full insert pipeline where the limit is applied. -
lightrag/constants.py– DefinesDEFAULT_MAX_PARALLEL_INSERT = 2and related async defaults likeDEFAULT_EMBEDDING_BATCH_NUM. -
lightrag/api/config.py– Handles CLI and environment variable parsing that mapsMAX_PARALLEL_INSERTto constructor arguments. -
tests/test_doc_status_chunk_preservation.py– Unit test explicitly settingmax_parallel_insert=1to verify deterministic behavior on small datasets. -
examples/insert_custom_kg.py– Example script demonstrating typicalLightRAGinstantiation with custom parameters includingmax_parallel_insert.
Summary
-
max_parallel_insertcontrols concurrent document processing through anasyncio.Semaphorecreated inlightrag/lightrag.py. -
Default value is 2, suitable for safe operation on modest hardware, but should be tuned based on profiling.
-
Resource constraints including GPU memory, storage connections, and I/O bandwidth dictate your optimal setting—never exceed backend connection limits.
-
Environment variable configuration via
MAX_PARALLEL_INSERTensures reproducible deployments across environments. -
Cancellation safety relies on bounded concurrency to ensure the
pipeline_statuscheck remains responsive during shutdown.
Frequently Asked Questions
What happens if I set max_parallel_insert too high?
Setting the value too high causes CPU or GPU overload on embedding models, "too many connections" errors from vector stores like PostgreSQL or Qdrant, and potential out-of-memory crashes as each coroutine retains chunk data until flush. Start at 2 and increase only after confirming your infrastructure handles the load.
How does max_parallel_insert differ from max_async?
While max_parallel_insert specifically governs the document ingestion semaphore in lightrag.py, max_async caps other asynchronous operations throughout the system including reranking and retrieval. Always maintain max_parallel_insert ≤ max_async to prevent resource contention and potential deadlocks during mixed workloads.
Can I change max_parallel_insert after initializing LightRAG?
No, the value is read during initialization and converted immediately into an asyncio.Semaphore stored in the instance. Changing the concurrency limit requires creating a new LightRAG instance with the updated parameter, as the semaphore cannot be resized dynamically once documents begin processing.
Why is the default value only 2?
The conservative default of 2 ensures safe operation across diverse environments including limited hardware and connection-restricted backends, as defined in lightrag/constants.py. This value prevents resource exhaustion for new users while providing a baseline that works with standard PostgreSQL or local vector store configurations without requiring immediate tuning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →