# How to Extract LaTeX Formulas with the `--enrich-formula` Option in OpenDataLoader-PDF

> Extract LaTeX formulas with the --enrich-formula option in OpenDataLoader-PDF. Learn how to get LaTeX source and bounding boxes in JSON format.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: how-to-guide
- Published: 2026-03-20

---

**Enable LaTeX formula extraction by starting the hybrid server with `--enrich-formula`, then run client conversions with `--hybrid-mode full` to receive JSON objects containing the LaTeX source and bounding box coordinates for each detected formula.**

The `opendataloader-pdf` repository provides specialized PDF conversion capabilities that identify mathematical expressions and return them as structured LaTeX code. Activating the `--enrich-formula` option loads a dedicated enrichment model within the hybrid server architecture, enabling machine-readable formula extraction alongside standard document conversion.

## Starting the Hybrid Server with Formula Enrichment

To activate LaTeX formula extraction, you must start the backend server with the enrichment flag enabled. In [`python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py), the argument parser defines `--enrich-formula` at lines 60-64, which sets the internal `enrich_formula` flag to `True`.

When the server initializes, the `create_app` function passes this flag to the `DocumentConverter` via the `create_converter` method at lines 91-96:

```bash
opendataloader-pdf-hybrid --enrich-formula

```

This command loads the **formula enrichment model** into memory and configures the Java-based Docling engine to process mathematical regions during conversion. The server remains active, waiting for client requests that specify full hybrid mode processing.

## Running Client Conversions in Full Hybrid Mode

After starting the enriched server, client conversions must explicitly request full processing to trigger formula extraction. The `--hybrid-mode full` flag is mandatory; without it, the enrichment step is skipped on the client side according to the documentation in [`README.md`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/README.md) (lines 24-33).

Execute the conversion using:

```bash
opendataloader-pdf --hybrid docling-fast \
    --hybrid-mode full file1.pdf file2.pdf folder/

```

This sends the PDF data to the hybrid server, which processes each page through the formula detection pipeline. The server returns a JSON array containing both standard document elements and specialized formula objects.

## Understanding the Formula Output Format

For each detected formula, the engine emits a JSON object with four specific fields. The `content` field contains the extracted LaTeX string, while `bounding box` provides the spatial coordinates in points:

```json
{
  "type": "formula",
  "page number": 1,
  "bounding box": [226.2, 144.7, 377.1, 168.7],
  "content": "\\frac{f(x+h) - f(x)}{h}"
}

```

- **type**: Always set to `"formula"` to identify the element category
- **page number**: Integer indicating the PDF page where the formula appears
- **bounding box**: Array of four coordinates defining the region `[x1, y1, x2, y2]`
- **content**: The extracted LaTeX source code as an escaped string

These entries integrate seamlessly with the standard document structure, allowing you to filter and process formulas independently from text content.

## Practical Implementation Examples

### Command Line Interface

Start the server in one terminal with formula enrichment enabled:

```bash

# Terminal 1 - Launch backend

opendataloader-pdf-hybrid --enrich-formula

```

Run the client conversion in a second terminal, filtering results to show only formulas:

```bash

# Terminal 2 - Convert and extract formulas

opendataloader-pdf --hybrid docling-fast \
    --hybrid-mode full my_scientific_paper.pdf | \
    jq '.[] | select(.type == "formula")'

```

### Python API

Access formula data programmatically while the enriched hybrid server runs in the background:

```python
import opendataloader_pdf

results = opendataloader_pdf.convert(
    input_path="my_scientific_paper.pdf",
    output_dir="out/",
    format="json",
    hybrid="docling-fast",
    hybrid_mode="full"
)

formulas = [e for e in results if e.get("type") == "formula"]
for f in formulas:
    print(f["content"])  # Output: \frac{f(x+h) - f(x)}{h}

```

### Node.js API

Extract formulas using the JavaScript client library:

```javascript
import { convert } from '@opendataloader/pdf';

const docs = await convert(
  ['my_scientific_paper.pdf'],
  { outputDir: 'out/', format: 'json', hybrid: 'docling-fast', hybridMode: 'full' }
);

const formulas = docs.filter(d => d.type === 'formula');
formulas.forEach(f => console.log(f.content));

```

## Summary

- **Server Requirement**: Start `opendataloader-pdf-hybrid` with `--enrich-formula` to load the formula detection model, as implemented in [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py) lines 60-64.
- **Client Requirement**: Always specify `--hybrid-mode full` when running conversions; otherwise the enrichment pipeline is bypassed.
- **Implementation Flow**: The flag propagates from CLI arguments through `create_converter` (lines 91-96) to the Java Docling engine.
- **Output Format**: Formulas return as JSON objects with `type: "formula"`, page numbers, bounding boxes in points, and LaTeX content strings.

## Frequently Asked Questions

### What is the difference between `--enrich-formula` and `--hybrid-mode full`?

The `--enrich-formula` flag is a **server-side** configuration that loads the machine learning model and enables the extraction capability within the hybrid server. The `--hybrid-mode full` flag is a **client-side** instruction that tells the converter to actually execute the enrichment pipeline during a specific conversion job. Both are required to successfully extract formulas.

### Where does the actual formula recognition happen?

The heavy computation occurs in the Java-based Docling engine running inside the hybrid server. When `enrich_formula=True` is passed to the `DocumentConverter` (as seen in [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py) lines 91-96), the server invokes specialized computer vision models to detect mathematical regions and parse them into LaTeX syntax before returning the JSON response.

### What coordinate system does the bounding box use?

The `bounding box` array uses **PDF points** (1/72 of an inch) relative to the bottom-left corner of the page, following standard PDF coordinate conventions. The four values represent `[x1, y1, x2, y2]` where `(x1, y1)` is the lower-left corner and `(x2, y2)` is the upper-right corner of the formula region.

### Can I extract formulas without using the hybrid server architecture?

No. According to the source code in `opendataloader-pdf`, the `--enrich-formula` functionality is tightly integrated into the hybrid server architecture defined in [`hybrid_server.py`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/hybrid_server.py). The formula enrichment model requires the persistent backend process to manage the Java Docling engine and ML model lifecycle, making standalone CLI extraction without the hybrid server unsupported.