OpenDataLoader PDF Hybrid Server Limits: Maximum File Size and Complexity Explained

The OpenDataLoader PDF hybrid server enforces a hard 100 MiB upload limit, restricts concurrent backend requests to 4, and caps LLM enrichment tokens at 300 to prevent memory exhaustion and maintain predictable performance.

The hybrid server in the opendataloader-project/opendataloader-pdf repository serves as a FastAPI wrapper that orchestrates PDF conversion through the Docling engine. Understanding its hard limits for maximum file size and complexity is essential for production deployments. These constraints are defined in hybrid_server.py and HybridConfig.java and work together to protect server resources while balancing throughput.

Hard Limits Defined in the Source Code

The system implements three specific tiers of protection: file upload size, backend concurrency, and generation token ceilings.

Maximum Uploaded File Size (100 MiB)

The maximum file size limit is defined by the MAX_FILE_SIZE constant in python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py. Set to 100 * 1024 * 1024 (100 MiB), this guard prevents out-of-memory crashes and protects the FastAPI server from excessively large multipart uploads.

When a client POSTs a PDF to /v1/convert/file, the server reads the entire file into memory. If len(content) > MAX_FILE_SIZE, the request is rejected immediately with a JSON error response. This check occurs at lines 353-357 of hybrid_server.py, using the constant defined at line 64.

Concurrent Backend Request Cap (4 Requests)

The Java side of the hybrid pipeline controls parallelism through HybridConfig.java. The maxConcurrentRequests parameter defaults to 4 concurrent requests (DEFAULT_MAX_CONCURRENT_REQUESTS = 4), limiting parallel calls to remote hybrid backends such as Docling or Hancom.

This limit, defined at lines 30-44 of java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java, prevents overwhelming the remote service or the host machine during heavy parallel workloads. Additional requests beyond this cap queue at the FastAPI level rather than flooding the backend.

LLM Enrichment Token Ceiling (300 Tokens)

When using --enrich-formula or --enrich-picture-description, the server calls a language model to generate text. The generation_config in hybrid_server.py sets max_new_tokens to 300 at line 231, capping the length of generated formula LaTeX strings or image captions.

This constraint ensures that documents with many enriched elements stay within reasonable latency and memory bounds, keeping response times predictable regardless of document complexity.

How the Limits Work in Practice

These three constraints operate sequentially during the document processing lifecycle.

First, the file upload guard validates the raw byte size of the incoming PDF. If the file exceeds 100 MiB, the server returns an immediate error without attempting conversion, saving computational resources.

Second, for files under the limit, the concurrency throttle manages how many simultaneous conversions can invoke the backend Docling engine. With a default of four parallel requests, the system balances throughput against memory pressure on the conversion workers.

Third, during enrichment phases, the token limit restricts LLM output to 300 tokens per call. This prevents runaway generation on complex mathematical formulas or detailed images that could otherwise delay the entire conversion pipeline.

Practical Code Examples

Successful Conversion Under the Limit

The following command converts a 20 MiB PDF with enrichment features enabled, operating comfortably within the default constraints:


# Convert a 20 MiB PDF, requesting pages 1-10

opendataloader-pdf-hybrid \
  --host 0.0.0.0 --port 5002 \
  --ocr-lang "en" \
  --enrich-formula \
  --enrich-picture-description \
  POST http://localhost:5002/v1/convert/file \
  -F "file=@small_document.pdf" \
  -F "page_ranges=1-10"

The server returns a JSON document with status "success" and populated enrichment fields.

Handling File Size Errors

When attempting to upload a 150 MiB PDF, the server triggers the size guard at lines 353-357 of hybrid_server.py:


# Attempt to upload a 150 MiB PDF

curl -X POST http://localhost:5002/v1/convert/file \
     -F "file=@huge_document.pdf"

Response:

{
  "status": "error",
  "errors": ["File size exceeds maximum allowed (100MB)"]
}

Modifying Limits for Advanced Use Cases

To process larger PDFs, edit the constant and rebuild the package:


# In hybrid_server.py

MAX_FILE_SIZE = 200 * 1024 * 1024   # increase to 200 MiB

After changing the source, reinstall the package:

pip install -e .[hybrid]   # editable install with hybrid extras

Caution: Raising the limit requires ensuring the host machine has sufficient RAM and that the web server (uvicorn) is configured with appropriate request-body size limits.

Summary

  • Maximum file size: Hard-coded to 100 MiB in hybrid_server.py (lines 64 and 353-357) to prevent memory exhaustion from large uploads.
  • Concurrent requests: Capped at 4 parallel backend calls via HybridConfig.java (lines 30-44) to protect backend resources.
  • LLM tokens: Limited to 300 max_new_tokens in generation_config (line 231) for predictable enrichment latency.
  • Error handling: Oversized files receive an immediate JSON error response without partial processing.
  • Customization: Limits can be increased by modifying source constants and reinstalling, though this requires sufficient host resources.

Frequently Asked Questions

What happens if I upload a PDF larger than 100 MiB?

The server rejects the request immediately with a JSON error stating "File size exceeds maximum allowed (100MB)". This check occurs at lines 353-357 of hybrid_server.py before any conversion processing begins, preventing memory allocation failures.

Can I process multiple large PDFs simultaneously?

The hybrid server allows multiple uploads but throttles concurrent backend processing to 4 requests by default. Additional requests queue at the FastAPI level. This limit is controlled by maxConcurrentRequests in HybridConfig.java (lines 30-44).

Why is the LLM token limit set to 300?

The 300 token ceiling for max_new_tokens in generation_config (line 231 of hybrid_server.py) ensures that formula LaTeX strings and image captions remain concise. This prevents runaway generation times and memory usage when processing documents with numerous enrichable elements.

Where are these limits configured in the codebase?

File size limits reside in python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py (the MAX_FILE_SIZE constant). Concurrency limits are defined in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java (the maxConcurrentRequests parameter). Token limits appear in the generation_config dictionary within hybrid_server.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →