How MinerU Handles Mathematical Formulas in PDFs: Detection, Recognition, and LaTeX Conversion

MinerU detects mathematical formulas in PDFs by rendering pages to images, running a specialized layout model to identify formula bounding boxes (categories 13 and 14), converting each formula image to LaTeX using a ViT-based recognition model, and finally injecting the cleaned LaTeX into Markdown output with configurable delimiters.

Understanding how MinerU handles mathematical formulas in PDFs reveals why this open-source tool produces publication-ready Markdown from scientific documents. The pipeline treats formulas as a distinct content type, separating them from plain text during layout analysis and applying specialized computer vision models to recover the underlying LaTeX markup.

The Formula Processing Pipeline

MinerU’s formula handling operates as a multi-stage pipeline controlled by the formula_enable configuration flag. When enabled, the system executes detection, recognition, and post-processing steps before merging results into the final document.

Step 1: PDF to Image Conversion

The process begins in mineru/utils/pdf_reader.py, where the pdf_to_images function rasterizes each PDF page into high-resolution PIL Image objects. This conversion is essential because subsequent layout and formula recognition models operate on pixel data rather than vector PDF commands.

Step 2: Layout Analysis and Formula Detection

The pipeline_doc_analyze function in mineru/backend/pipeline/pipeline_analyze.py orchestrates layout detection using the MFD (Mathematical Formula Detection) model. This YOLO-based detector emits bounding boxes with specific category IDs:

  • Category 13: Inline formulas embedded within text lines
  • Category 14: Display formulas appearing as separate blocks

In mineru/model/mfr/unimernet/Unimernet.py, the _filter_boxes_by_iou method filters overlapping detections before passing valid formula crops to the recognition stage.

Step 3: Formula Recognition (MFR)

Each detected formula image undergoes recognition via the PP-Formulanet-Plus-M model, a Vision Transformer (ViT) based encoder-decoder architecture. The implementation in mineru/model/mfr/pp_formulanet_plus_m/predict_formula.py processes the image tensor and generates the raw LaTeX string.

The UnimernetModel.predict method in mineru/model/mfr/unimernet/Unimernet.py collects these results and stores them in the latex field of each formula dictionary:

new_item = {
    "category_id": 13 + int(cla.item()),
    "poly": [...],
    "score": round(float(conf.item()), 2),
    "latex": "",  # Populated after MFR

}

Step 4: LaTeX Post-Processing

Raw LaTeX output from vision models often contains structural errors such as mismatched \left/\right delimiters or unbalanced braces. The mineru/model/mfr/utils.py module provides correction utilities:

  • fix_latex_left_right: Balances delimiter pairs
  • fix_unbalanced_braces: Corrects orphaned opening or closing braces

These functions ensure that the final LaTeX strings compile correctly in standard Markdown renderers.

Step 5: Markdown Integration

The cleaned LaTeX merges into the document structure through merge_para_with_text and mk_blocks_to_markdown functions. These appear in both the standard pipeline (mineru/backend/pipeline/pipeline_middle_json_mkcontent.py) and the VLM backend (mineru/backend/vlm/vlm_middle_json_mkcontent.py).

The system respects delimiter configuration from mineru/utils/config_reader.get_latex_delimiter_config, defaulting to:

  • Inline math: $...$
  • Display math: $$...$$

Configuration and Control Flags

MinerU exposes formula handling through the formula_enable boolean flag, accessible via CLI arguments, environment variables, or API payloads.

Enabling Formula Extraction via CLI

Activate formula processing when parsing PDFs from the command line:

minerU parse input.pdf --formula-enable true --backend pipeline

The --formula-enable argument propagates through mineru/cli/common.py to pipeline_analyze, controlling whether the MFR branch executes or whether formulas are treated as static images.

Programmatic API Usage

For Python integrations, pass formula_enable=True to pipeline_doc_analyze:

from mineru.cli.common import read_fn, prepare_env, _process_pipeline
from mineru.utils.pdf_reader import pdf_to_images
from mineru.backend.pipeline.pipeline_analyze import pipeline_doc_analyze

pdf_bytes = read_fn("sample.pdf")
images = pdf_to_images(pdf_bytes)                 # step 1

# run layout+OCR; ask for formulas

infer_results, _, _, _, _ = pipeline_doc_analyze(
    [pdf_bytes],
    ["ch"],                # language

    parse_method="pipeline",
    formula_enable=True,   # turn on formula detection

    table_enable=False,
)

# `infer_results` now contains formula dicts with `latex` fields

for page in infer_results[0]:
    if page["category_id"] in (13, 14):
        print(page["latex"])

Key Implementation Files

Role File Path
PDF to image conversion mineru/utils/pdf_reader.py
Layout & OCR pipeline entry mineru/backend/pipeline/pipeline_analyze.py
Formula detection & cropping mineru/model/mfr/unimernet/Unimernet.py
Formula recognition model (PP-Formulanet-Plus-M) mineru/model/mfr/pp_formulanet_plus_m/predict_formula.py
LaTeX post-processing utilities mineru/model/mfr/utils.py
Config flag reader (formula_enable) mineru/utils/config_reader.py
Markdown merging (pipeline) mineru/backend/pipeline/pipeline_middle_json_mkcontent.py
Markdown merging (VLM backend) mineru/backend/vlm/vlm_middle_json_mkcontent.py
CLI entry point mineru/cli/gradio_app.py

Summary

  • MinerU handles mathematical formulas in PDFs through a dedicated pipeline that detects formulas as distinct layout elements (categories 13 and 14), recognizes them using the PP-Formulanet-Plus-M vision model, and converts them to LaTeX.
  • Detection occurs on rasterized images produced by pdf_to_images in mineru/utils/pdf_reader.py, with bounding boxes filtered by IOU in Unimernet.py.
  • Recognition generates raw LaTeX that undergoes mandatory post-processing in mineru/model/mfr/utils.py to fix delimiter mismatches and unbalanced braces.
  • Output integration embeds formulas into Markdown using configurable delimiters ($...$ for inline, $$...$$ for display) via pipeline_middle_json_mkcontent.py or the VLM backend equivalent.
  • Control is centralized through the formula_enable flag, accessible via CLI (--formula-enable), Python API (formula_enable=True), or environment variables.

Frequently Asked Questions

What model does MinerU use for formula recognition?

MinerU uses the PP-Formulanet-Plus-M model, a Vision Transformer (ViT)-based encoder-decoder architecture implemented in mineru/model/mfr/pp_formulanet_plus_m/predict_formula.py. This model converts cropped formula images into raw LaTeX strings, which are then post-processed to correct structural errors before insertion into the final Markdown.

Can I disable formula detection if I only need plain text?

Yes. Formula extraction is controlled by the formula_enable boolean flag. Set --formula-enable false in the CLI, pass formula_enable=False to the Python API (pipeline_doc_analyze), or configure it via environment variables. When disabled, MinerU skips the MFR (Mathematical Formula Recognition) branch and treats formula regions as ordinary image blocks or ignores them depending on the layout configuration.

How does MinerU distinguish between inline and display formulas?

During layout analysis in mineru/model/mfr/unimernet/Unimernet.py, the MFD (Mathematical Formula Detection) model assigns category ID 13 to inline formulas (embedded within text lines) and category ID 14 to display formulas (standalone blocks). The predict method uses these IDs to determine whether to wrap the resulting LaTeX in single dollar signs $...$ (inline) or double dollar signs $$...$$ (display) during the Markdown generation phase.

What happens if the formula recognition fails?

If the PP-Formulanet-Plus-M model returns an empty or malformed LaTeX string, the post-processing utilities in mineru/model/mfr/utils.py attempt to repair common structural issues such as mismatched \left/\right pairs and unbalanced braces. If the LaTeX remains invalid after cleaning, the pipeline typically retains the formula image as a fallback or inserts the raw (potentially broken) LaTeX string into the Markdown, depending on the error handling configuration in the specific backend (pipeline vs. VLM).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →