How MinerU Detects and Handles Hallucinations in PDF Parsing
MinerU prevents hallucinations by classifying PDFs into text-based or OCR-based extraction pipelines before processing, ensuring corrupted or image-heavy documents never trigger false text generation.
MinerU, the open-source PDF parsing toolkit from opendatalab/MinerU, implements a sophisticated hallucination detection and handling system that analyzes document structure before extraction. Rather than attempting to parse every PDF with a single method, the system proactively identifies documents likely to produce garbled or fabricated text—such as those with corrupted fonts or scan-based pages—and routes them to specialized OCR pipelines.
The Classification-Based Defense Against Hallucinations
The core defense mechanism resides in mineru/utils/pdf_classify.py, where the classify() function implements a four-stage validation process. This function samples up to ten random pages from the input PDF and runs heuristic checks to determine whether the document should use native text extraction or OCR-based processing.
Sampling Strategy for Large Documents
To maintain performance on multi-hundred-page documents, the system uses extract_pages to pull a random sample of up to ten pages. This statistical approach ensures the classification decision reflects the overall document quality without requiring full PDF traversal, keeping the hallucination detection overhead minimal even for large files.
Text Density Validation
The get_avg_cleaned_chars_per_page function calculates the average character count per sampled page after stripping whitespace. If the resulting average falls below 50 characters per page, the document is classified as text-poor and immediately flagged for OCR processing. This threshold prevents the system from attempting to extract meaningful content from nearly blank or heavily image-based pages where native extraction would likely hallucinate structure.
Corruption Detection via CID Tokens
MinerU detects font corruption through the detect_invalid_chars function, which scans the sampled pages for "(cid:xxx)" patterns—placeholder tokens that appear when PDF fonts are missing or improperly mapped. When these CID tokens constitute more than 5% of the total character count, the document is flagged as corrupted and routed to OCR. This specific check targets a common source of extraction hallucinations where garbled CID strings appear as nonsensical text.
Image Coverage Analysis
The get_high_image_coverage_ratio function computes the proportion of page area covered by images across the sampled pages. If 80% or more of the pages contain high image coverage, the document is classified as scan-based or presentation-heavy and sent to OCR. This prevents the text extraction pipeline from attempting to read text from complex diagrams or photographs where it might invent text content.
Hallucination-Free Pipeline Architecture
Once classified, MinerU enforces a strict separation between extraction methods. As documented in mineru/cli/gradio_app.py (line 288), the system advertises a "hallucination-free" pipeline backend because it never mixes OCR-generated text with native PDF text. The classification step guarantees a single, coherent extraction path—either pure text extraction or pure OCR—eliminating the inconsistencies that arise when switching methods mid-document.
"backend_info_pipeline": "Traditional Multi-model pipeline parsing, supports multiple languages, hallucination-free."
Advanced Safeguards for Hybrid Deployments
For deployments using hybrid backends, MinerU provides an environment variable override that forces the text-extraction stage to use the smaller "pipeline" model rather than OCR, even when the classifier suggests otherwise. Setting MINERU_HYBRID_FORCE_PIPELINE_ENABLE=true (documented in docs/en/usage/cli_tools.md, lines 32-35) reduces hallucinations in edge cases where OCR might introduce artifacts.
export MINERU_HYBRID_FORCE_PIPELINE_ENABLE=true
python -m mineru.cli.gradio_app
When enabled, this flag ensures that even borderline documents receive the hallucination-free pipeline treatment rather than risking OCR-induced text generation errors.
Practical Implementation
To implement hallucination detection in your own MinerU workflow, use the classify function from mineru/utils/pdf_classify.py to determine the appropriate extraction mode before processing:
from mineru.utils.pdf_classify import classify
# Load PDF as bytes
with open("document.pdf", "rb") as f:
pdf_bytes = f.read()
# Detect potential hallucination sources
extraction_mode = classify(pdf_bytes) # Returns 'txt' or 'ocr'
print(f"Selected extraction mode: {extraction_mode}")
Based on the classification result, route the document to the appropriate backend:
if extraction_mode == "txt":
# Use hallucination-free text extraction
from mineru.backend.pipeline import pipeline_process
result = pipeline_process(pdf_bytes)
else:
# Use OCR for corrupted or image-heavy documents
from mineru.backend.ocr import ocr_process
result = ocr_process(pdf_bytes)
Summary
- Proactive Classification: MinerU analyzes PDFs before extraction using
mineru/utils/pdf_classify.pyto detect low text density, CID token corruption, and high image coverage. - Automatic Routing: Documents triggering any hallucination risk factor (under 50 characters per page, over 5% CID tokens, or 80% image coverage) are automatically routed to OCR rather than native text extraction.
- Pipeline Isolation: The system maintains a strict separation between text-extraction and OCR pipelines, ensuring hallucination-free processing by never mixing extraction methods within a single document.
- Override Controls: The
MINERU_HYBRID_FORCE_PIPELINE_ENABLEenvironment variable provides an additional safeguard for hybrid deployments, forcing the use of the hallucination-free pipeline model when needed.
Frequently Asked Questions
How does MinerU detect corrupted fonts that could cause hallucinations?
MinerU detects corrupted fonts through the detect_invalid_chars function in mineru/utils/pdf_classify.py. This function scans sampled pages for "(cid:xxx)" tokens—placeholders that appear when PDF fonts are missing or improperly mapped. When these CID tokens exceed 5% of the total character count, the document is flagged as corrupted and routed to OCR processing to prevent the extraction of garbled text.
What is the MINERU_HYBRID_FORCE_PIPELINE_ENABLE environment variable used for?
The MINERU_HYBRID_FORCE_PIPELINE_ENABLE environment variable, documented in docs/en/usage/cli_tools.md, allows users to override the automatic classification system in hybrid deployments. When set to true, it forces the text-extraction stage to use the smaller "pipeline" model rather than OCR, even if the classifier suggests OCR is needed. This provides an additional safeguard against OCR-induced hallucinations in edge cases.
Why does MinerU use random page sampling instead of analyzing the entire PDF?
MinerU samples up to ten random pages using extract_pages in mineru/utils/pdf_classify.py to maintain performance on large documents. This statistical approach allows the hallucination detection system to make accurate classification decisions without the overhead of parsing multi-hundred-page PDFs in full. The sampling strategy ensures that the text density, corruption, and image coverage checks remain fast and scalable while still representative of the document's overall characteristics.
How does MinerU prevent mixing OCR and native text extraction within the same document?
MinerU prevents mixing extraction methods by enforcing a strict classification-based routing system. The classify function in mineru/utils/pdf_classify.py returns either 'txt' or 'ocr' before any extraction begins, and the system commits to that single pipeline for the entire document. As noted in mineru/cli/gradio_app.py, this architecture is explicitly labeled "hallucination-free" because it eliminates the inconsistencies that arise when switching between OCR-generated text and native PDF text within the same parsing operation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →