# How MinerU Detects and Handles Hallucinations in PDF Parsing

> MinerU prevents PDF parsing hallucinations by classifying documents. Learn how it avoids false text generation for corrupted or image-heavy files. Read more.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: how-to-guide
- Published: 2026-02-23

---

**MinerU prevents hallucinations by classifying PDFs into text-based or OCR-based extraction pipelines before processing, ensuring corrupted or image-heavy documents never trigger false text generation.**

MinerU, the open-source PDF parsing toolkit from `opendatalab/MinerU`, implements a sophisticated **hallucination detection and handling** system that analyzes document structure before extraction. Rather than attempting to parse every PDF with a single method, the system proactively identifies documents likely to produce garbled or fabricated text—such as those with corrupted fonts or scan-based pages—and routes them to specialized OCR pipelines.

## The Classification-Based Defense Against Hallucinations

The core defense mechanism resides in [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py), where the `classify()` function implements a four-stage validation process. This function samples up to ten random pages from the input PDF and runs heuristic checks to determine whether the document should use native text extraction or OCR-based processing.

### Sampling Strategy for Large Documents

To maintain performance on multi-hundred-page documents, the system uses `extract_pages` to pull a random sample of up to ten pages. This statistical approach ensures the classification decision reflects the overall document quality without requiring full PDF traversal, keeping the hallucination detection overhead minimal even for large files.

### Text Density Validation

The `get_avg_cleaned_chars_per_page` function calculates the average character count per sampled page after stripping whitespace. If the resulting average falls **below 50 characters per page**, the document is classified as text-poor and immediately flagged for OCR processing. This threshold prevents the system from attempting to extract meaningful content from nearly blank or heavily image-based pages where native extraction would likely hallucinate structure.

### Corruption Detection via CID Tokens

MinerU detects font corruption through the `detect_invalid_chars` function, which scans the sampled pages for `"(cid:xxx)"` patterns—placeholder tokens that appear when PDF fonts are missing or improperly mapped. When these CID tokens constitute **more than 5%** of the total character count, the document is flagged as corrupted and routed to OCR. This specific check targets a common source of extraction hallucinations where garbled CID strings appear as nonsensical text.

### Image Coverage Analysis

The `get_high_image_coverage_ratio` function computes the proportion of page area covered by images across the sampled pages. If **80% or more** of the pages contain high image coverage, the document is classified as scan-based or presentation-heavy and sent to OCR. This prevents the text extraction pipeline from attempting to read text from complex diagrams or photographs where it might invent text content.

## Hallucination-Free Pipeline Architecture

Once classified, MinerU enforces a strict separation between extraction methods. As documented in [`mineru/cli/gradio_app.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/gradio_app.py) (line 288), the system advertises a **"hallucination-free"** pipeline backend because it **never mixes OCR-generated text with native PDF text**. The classification step guarantees a single, coherent extraction path—either pure text extraction or pure OCR—eliminating the inconsistencies that arise when switching methods mid-document.

```python
"backend_info_pipeline": "Traditional Multi-model pipeline parsing, supports multiple languages, hallucination-free."

```

## Advanced Safeguards for Hybrid Deployments

For deployments using hybrid backends, MinerU provides an environment variable override that forces the text-extraction stage to use the smaller "pipeline" model rather than OCR, even when the classifier suggests otherwise. Setting `MINERU_HYBRID_FORCE_PIPELINE_ENABLE=true` (documented in [`docs/en/usage/cli_tools.md`](https://github.com/opendatalab/MinerU/blob/main/docs/en/usage/cli_tools.md), lines 32-35) reduces hallucinations in edge cases where OCR might introduce artifacts.

```bash
export MINERU_HYBRID_FORCE_PIPELINE_ENABLE=true
python -m mineru.cli.gradio_app

```

When enabled, this flag ensures that even borderline documents receive the hallucination-free pipeline treatment rather than risking OCR-induced text generation errors.

## Practical Implementation

To implement hallucination detection in your own MinerU workflow, use the `classify` function from [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py) to determine the appropriate extraction mode before processing:

```python
from mineru.utils.pdf_classify import classify

# Load PDF as bytes

with open("document.pdf", "rb") as f:
    pdf_bytes = f.read()

# Detect potential hallucination sources

extraction_mode = classify(pdf_bytes)  # Returns 'txt' or 'ocr'

print(f"Selected extraction mode: {extraction_mode}")

```

Based on the classification result, route the document to the appropriate backend:

```python
if extraction_mode == "txt":
    # Use hallucination-free text extraction

    from mineru.backend.pipeline import pipeline_process
    result = pipeline_process(pdf_bytes)
else:
    # Use OCR for corrupted or image-heavy documents

    from mineru.backend.ocr import ocr_process
    result = ocr_process(pdf_bytes)

```

## Summary

- **Proactive Classification**: MinerU analyzes PDFs before extraction using [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py) to detect low text density, CID token corruption, and high image coverage.
- **Automatic Routing**: Documents triggering any hallucination risk factor (under 50 characters per page, over 5% CID tokens, or 80% image coverage) are automatically routed to OCR rather than native text extraction.
- **Pipeline Isolation**: The system maintains a strict separation between text-extraction and OCR pipelines, ensuring hallucination-free processing by never mixing extraction methods within a single document.
- **Override Controls**: The `MINERU_HYBRID_FORCE_PIPELINE_ENABLE` environment variable provides an additional safeguard for hybrid deployments, forcing the use of the hallucination-free pipeline model when needed.

## Frequently Asked Questions

### How does MinerU detect corrupted fonts that could cause hallucinations?

MinerU detects corrupted fonts through the `detect_invalid_chars` function in [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py). This function scans sampled pages for `"(cid:xxx)"` tokens—placeholders that appear when PDF fonts are missing or improperly mapped. When these CID tokens exceed 5% of the total character count, the document is flagged as corrupted and routed to OCR processing to prevent the extraction of garbled text.

### What is the MINERU_HYBRID_FORCE_PIPELINE_ENABLE environment variable used for?

The `MINERU_HYBRID_FORCE_PIPELINE_ENABLE` environment variable, documented in [`docs/en/usage/cli_tools.md`](https://github.com/opendatalab/MinerU/blob/main/docs/en/usage/cli_tools.md), allows users to override the automatic classification system in hybrid deployments. When set to `true`, it forces the text-extraction stage to use the smaller "pipeline" model rather than OCR, even if the classifier suggests OCR is needed. This provides an additional safeguard against OCR-induced hallucinations in edge cases.

### Why does MinerU use random page sampling instead of analyzing the entire PDF?

MinerU samples up to ten random pages using `extract_pages` in [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py) to maintain performance on large documents. This statistical approach allows the hallucination detection system to make accurate classification decisions without the overhead of parsing multi-hundred-page PDFs in full. The sampling strategy ensures that the text density, corruption, and image coverage checks remain fast and scalable while still representative of the document's overall characteristics.

### How does MinerU prevent mixing OCR and native text extraction within the same document?

MinerU prevents mixing extraction methods by enforcing a strict classification-based routing system. The `classify` function in [`mineru/utils/pdf_classify.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/utils/pdf_classify.py) returns either `'txt'` or `'ocr'` before any extraction begins, and the system commits to that single pipeline for the entire document. As noted in [`mineru/cli/gradio_app.py`](https://github.com/opendatalab/MinerU/blob/main/mineru/cli/gradio_app.py), this architecture is explicitly labeled "hallucination-free" because it eliminates the inconsistencies that arise when switching between OCR-generated text and native PDF text within the same parsing operation.