# OpenDataLoader PDF Hybrid Server Limits: Maximum File Size and Complexity Explained

> Discover OpenDataLoader PDF hybrid server limits including the 100 MiB file size cap, 4 concurrent backend requests, and 300 LLM tokens to ensure optimal performance and prevent memory issues.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: performance
- Published: 2026-03-20

---

**The OpenDataLoader PDF hybrid server enforces a hard 100 MiB upload limit, restricts concurrent backend requests to 4, and caps LLM enrichment tokens at 300 to prevent memory exhaustion and maintain predictable performance.**

The hybrid server in the `opendataloader-project/opendataloader-pdf` repository serves as a FastAPI wrapper that orchestrates PDF conversion through the Docling engine. Understanding its hard limits for maximum file size and complexity is essential for production deployments. These constraints are defined in [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py) and [`HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridConfig.java) and work together to protect server resources while balancing throughput.

## Hard Limits Defined in the Source Code

The system implements three specific tiers of protection: file upload size, backend concurrency, and generation token ceilings.

### Maximum Uploaded File Size (100 MiB)

The **maximum file size** limit is defined by the `MAX_FILE_SIZE` constant in [`python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py). Set to `100 * 1024 * 1024` (100 MiB), this guard prevents out-of-memory crashes and protects the FastAPI server from excessively large multipart uploads.

When a client POSTs a PDF to `/v1/convert/file`, the server reads the entire file into memory. If `len(content) > MAX_FILE_SIZE`, the request is rejected immediately with a JSON error response. This check occurs at lines 353-357 of [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py), using the constant defined at line 64.

### Concurrent Backend Request Cap (4 Requests)

The Java side of the hybrid pipeline controls parallelism through [`HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridConfig.java). The `maxConcurrentRequests` parameter defaults to **4** concurrent requests (`DEFAULT_MAX_CONCURRENT_REQUESTS = 4`), limiting parallel calls to remote hybrid backends such as Docling or Hancom.

This limit, defined at lines 30-44 of [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java), prevents overwhelming the remote service or the host machine during heavy parallel workloads. Additional requests beyond this cap queue at the FastAPI level rather than flooding the backend.

### LLM Enrichment Token Ceiling (300 Tokens)

When using `--enrich-formula` or `--enrich-picture-description`, the server calls a language model to generate text. The `generation_config` in [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py) sets `max_new_tokens` to **300** at line 231, capping the length of generated formula LaTeX strings or image captions.

This constraint ensures that documents with many enriched elements stay within reasonable latency and memory bounds, keeping response times predictable regardless of document complexity.

## How the Limits Work in Practice

These three constraints operate sequentially during the document processing lifecycle.

First, the **file upload guard** validates the raw byte size of the incoming PDF. If the file exceeds 100 MiB, the server returns an immediate error without attempting conversion, saving computational resources.

Second, for files under the limit, the **concurrency throttle** manages how many simultaneous conversions can invoke the backend Docling engine. With a default of four parallel requests, the system balances throughput against memory pressure on the conversion workers.

Third, during enrichment phases, the **token limit** restricts LLM output to 300 tokens per call. This prevents runaway generation on complex mathematical formulas or detailed images that could otherwise delay the entire conversion pipeline.

## Practical Code Examples

### Successful Conversion Under the Limit

The following command converts a 20 MiB PDF with enrichment features enabled, operating comfortably within the default constraints:

```bash

# Convert a 20 MiB PDF, requesting pages 1-10

opendataloader-pdf-hybrid \
  --host 0.0.0.0 --port 5002 \
  --ocr-lang "en" \
  --enrich-formula \
  --enrich-picture-description \
  POST http://localhost:5002/v1/convert/file \
  -F "file=@small_document.pdf" \
  -F "page_ranges=1-10"

```

The server returns a JSON document with status `"success"` and populated enrichment fields.

### Handling File Size Errors

When attempting to upload a 150 MiB PDF, the server triggers the size guard at lines 353-357 of [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py):

```bash

# Attempt to upload a 150 MiB PDF

curl -X POST http://localhost:5002/v1/convert/file \
     -F "file=@huge_document.pdf"

```

**Response:**

```json
{
  "status": "error",
  "errors": ["File size exceeds maximum allowed (100MB)"]
}

```

### Modifying Limits for Advanced Use Cases

To process larger PDFs, edit the constant and rebuild the package:

```python

# In hybrid_server.py

MAX_FILE_SIZE = 200 * 1024 * 1024   # increase to 200 MiB

```

After changing the source, reinstall the package:

```bash
pip install -e .[hybrid]   # editable install with hybrid extras

```

**Caution:** Raising the limit requires ensuring the host machine has sufficient RAM and that the web server (`uvicorn`) is configured with appropriate request-body size limits.

## Summary

- **Maximum file size**: Hard-coded to 100 MiB in [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py) (lines 64 and 353-357) to prevent memory exhaustion from large uploads.
- **Concurrent requests**: Capped at 4 parallel backend calls via [`HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridConfig.java) (lines 30-44) to protect backend resources.
- **LLM tokens**: Limited to 300 `max_new_tokens` in `generation_config` (line 231) for predictable enrichment latency.
- **Error handling**: Oversized files receive an immediate JSON error response without partial processing.
- **Customization**: Limits can be increased by modifying source constants and reinstalling, though this requires sufficient host resources.

## Frequently Asked Questions

### What happens if I upload a PDF larger than 100 MiB?

The server rejects the request immediately with a JSON error stating `"File size exceeds maximum allowed (100MB)"`. This check occurs at lines 353-357 of [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py) before any conversion processing begins, preventing memory allocation failures.

### Can I process multiple large PDFs simultaneously?

The hybrid server allows multiple uploads but throttles concurrent backend processing to 4 requests by default. Additional requests queue at the FastAPI level. This limit is controlled by `maxConcurrentRequests` in [`HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridConfig.java) (lines 30-44).

### Why is the LLM token limit set to 300?

The 300 token ceiling for `max_new_tokens` in `generation_config` (line 231 of [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py)) ensures that formula LaTeX strings and image captions remain concise. This prevents runaway generation times and memory usage when processing documents with numerous enrichable elements.

### Where are these limits configured in the codebase?

File size limits reside in [`python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py) (the `MAX_FILE_SIZE` constant). Concurrency limits are defined in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/hybrid/HybridConfig.java) (the `maxConcurrentRequests` parameter). Token limits appear in the `generation_config` dictionary within [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py).