How to Extract LaTeX Formulas with the `--enrich-formula` Option in OpenDataLoader-PDF

Enable LaTeX formula extraction by starting the hybrid server with --enrich-formula, then run client conversions with --hybrid-mode full to receive JSON objects containing the LaTeX source and bounding box coordinates for each detected formula.

The opendataloader-pdf repository provides specialized PDF conversion capabilities that identify mathematical expressions and return them as structured LaTeX code. Activating the --enrich-formula option loads a dedicated enrichment model within the hybrid server architecture, enabling machine-readable formula extraction alongside standard document conversion.

Starting the Hybrid Server with Formula Enrichment

To activate LaTeX formula extraction, you must start the backend server with the enrichment flag enabled. In python/opendataloader-pdf/src/opendataloader_pdf/hybrid_server.py, the argument parser defines --enrich-formula at lines 60-64, which sets the internal enrich_formula flag to True.

When the server initializes, the create_app function passes this flag to the DocumentConverter via the create_converter method at lines 91-96:

opendataloader-pdf-hybrid --enrich-formula

This command loads the formula enrichment model into memory and configures the Java-based Docling engine to process mathematical regions during conversion. The server remains active, waiting for client requests that specify full hybrid mode processing.

Running Client Conversions in Full Hybrid Mode

After starting the enriched server, client conversions must explicitly request full processing to trigger formula extraction. The --hybrid-mode full flag is mandatory; without it, the enrichment step is skipped on the client side according to the documentation in README.md (lines 24-33).

Execute the conversion using:

opendataloader-pdf --hybrid docling-fast \
    --hybrid-mode full file1.pdf file2.pdf folder/

This sends the PDF data to the hybrid server, which processes each page through the formula detection pipeline. The server returns a JSON array containing both standard document elements and specialized formula objects.

Understanding the Formula Output Format

For each detected formula, the engine emits a JSON object with four specific fields. The content field contains the extracted LaTeX string, while bounding box provides the spatial coordinates in points:

{
  "type": "formula",
  "page number": 1,
  "bounding box": [226.2, 144.7, 377.1, 168.7],
  "content": "\\frac{f(x+h) - f(x)}{h}"
}
  • type: Always set to "formula" to identify the element category
  • page number: Integer indicating the PDF page where the formula appears
  • bounding box: Array of four coordinates defining the region [x1, y1, x2, y2]
  • content: The extracted LaTeX source code as an escaped string

These entries integrate seamlessly with the standard document structure, allowing you to filter and process formulas independently from text content.

Practical Implementation Examples

Command Line Interface

Start the server in one terminal with formula enrichment enabled:


# Terminal 1 - Launch backend

opendataloader-pdf-hybrid --enrich-formula

Run the client conversion in a second terminal, filtering results to show only formulas:


# Terminal 2 - Convert and extract formulas

opendataloader-pdf --hybrid docling-fast \
    --hybrid-mode full my_scientific_paper.pdf | \
    jq '.[] | select(.type == "formula")'

Python API

Access formula data programmatically while the enriched hybrid server runs in the background:

import opendataloader_pdf

results = opendataloader_pdf.convert(
    input_path="my_scientific_paper.pdf",
    output_dir="out/",
    format="json",
    hybrid="docling-fast",
    hybrid_mode="full"
)

formulas = [e for e in results if e.get("type") == "formula"]
for f in formulas:
    print(f["content"])  # Output: \frac{f(x+h) - f(x)}{h}

Node.js API

Extract formulas using the JavaScript client library:

import { convert } from '@opendataloader/pdf';

const docs = await convert(
  ['my_scientific_paper.pdf'],
  { outputDir: 'out/', format: 'json', hybrid: 'docling-fast', hybridMode: 'full' }
);

const formulas = docs.filter(d => d.type === 'formula');
formulas.forEach(f => console.log(f.content));

Summary

  • Server Requirement: Start opendataloader-pdf-hybrid with --enrich-formula to load the formula detection model, as implemented in hybrid_server.py lines 60-64.
  • Client Requirement: Always specify --hybrid-mode full when running conversions; otherwise the enrichment pipeline is bypassed.
  • Implementation Flow: The flag propagates from CLI arguments through create_converter (lines 91-96) to the Java Docling engine.
  • Output Format: Formulas return as JSON objects with type: "formula", page numbers, bounding boxes in points, and LaTeX content strings.

Frequently Asked Questions

What is the difference between --enrich-formula and --hybrid-mode full?

The --enrich-formula flag is a server-side configuration that loads the machine learning model and enables the extraction capability within the hybrid server. The --hybrid-mode full flag is a client-side instruction that tells the converter to actually execute the enrichment pipeline during a specific conversion job. Both are required to successfully extract formulas.

Where does the actual formula recognition happen?

The heavy computation occurs in the Java-based Docling engine running inside the hybrid server. When enrich_formula=True is passed to the DocumentConverter (as seen in hybrid_server.py lines 91-96), the server invokes specialized computer vision models to detect mathematical regions and parse them into LaTeX syntax before returning the JSON response.

What coordinate system does the bounding box use?

The bounding box array uses PDF points (1/72 of an inch) relative to the bottom-left corner of the page, following standard PDF coordinate conventions. The four values represent [x1, y1, x2, y2] where (x1, y1) is the lower-left corner and (x2, y2) is the upper-right corner of the formula region.

Can I extract formulas without using the hybrid server architecture?

No. According to the source code in opendataloader-pdf, the --enrich-formula functionality is tightly integrated into the hybrid server architecture defined in hybrid_server.py. The formula enrichment model requires the persistent backend process to manage the Java Docling engine and ML model lifecycle, making standalone CLI extraction without the hybrid server unsupported.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →