MinerU OmniDocBench Accuracy: Benchmark Results for Pipeline and VLM Backends
MinerU 2.5 achieves state-of-the-art End-to-End Evaluation Overall scores on OmniDocBench v1.5, surpassing larger multimodal models like GPT-4o and Gemini 2.5 Pro with only 1.2B parameters.
MinerU is an open-source document parsing toolkit that extracts structured content from PDFs and other formats. Understanding MinerU OmniDocBench accuracy is crucial for evaluating its performance against other document understanding solutions, as this benchmark represents the current standard for end-to-end document parsing evaluation.
Understanding MinerU OmniDocBench Accuracy Metrics
End-to-End Evaluation Overall Score
According to the repository's changelog at docs/en/reference/changelog.md, MinerU 2.5 is evaluated using the End-to-End Evaluation Overall metric from OmniDocBench v1.5. This score measures the complete document parsing pipeline—from raw PDF input to structured JSON output—rather than evaluating individual components in isolation.
The evaluation compares the generated middle.json and content_list.json files against ground truth annotations. These output files represent the final structured representation of the document, containing extracted text, tables, figures, and layout information.
Model Architecture and Parameter Efficiency
MinerU 2.5 operates with only 1.2 billion parameters, making it significantly more efficient than competing solutions. Despite this compact architecture, the model achieves superior OmniDocBench accuracy compared to:
- Gemini 2.5 Pro
- GPT-4o
- Qwen 2.5-VL-72B
- dots.ocr
- MonkeyOCR
- PP-StructureV3
This efficiency makes MinerU particularly suitable for deployment environments with limited computational resources while maintaining state-of-the-art parsing accuracy.
Pipeline vs VLM Backend Performance on OmniDocBench
Shared Benchmark Results
Both the pipeline and vlm backends share identical benchmark results on OmniDocBench v1.5. The documentation does not publish separate accuracy numbers for each backend because the End-to-End Evaluation Overall score is computed on the final structured output, regardless of which backend generated it.
The pipeline backend, implemented in minerU/backend/pipeline/*.py, uses traditional computer vision and rule-based approaches combined with deep learning models. The VLM backend, implemented in minerU/cli/vlm_server.py and minerU/cli/client.py, leverages vision-language models for document understanding.
Output Format and Evaluation Consistency
Both backends produce output conforming to the schema defined in minerU/template.json, generating middle.json files that OmniDocBench evaluates. While the internal processing differs:
- The pipeline backend processes documents through staged extraction pipelines
- The VLM backend uses
vlm-auto-enginefor end-to-end generation
The final accuracy metric remains consistent because evaluation occurs on the structured JSON output rather than intermediate representations.
Competitive Comparison on OmniDocBench
MinerU 2.5's OmniDocBench accuracy positions it as the leading open-source document parsing solution. The changelog explicitly states that MinerU 2.5 surpasses both general-purpose multimodal models and specialized OCR systems:
General Multimodal Models:
- Outperforms Gemini 2.5 Pro despite having 100x fewer parameters
- Exceeds GPT-4o accuracy on document structure recognition
- Surpasses Qwen 2.5-VL-72B, a 72-billion parameter vision-language model
Dedicated Document Parsing Models:
- Outperforms dots.ocr on complex layout preservation
- Exceeds MonkeyOCR on table structure recognition
- Surpasses PP-StructureV3 (PaddlePaddle's document analysis system)
This performance is achieved while maintaining a lightweight 1.2B parameter footprint, making MinerU deployable on consumer hardware and edge devices where larger models would be impractical.
Reproducing MinerU OmniDocBench Accuracy Results
To verify the benchmark results locally, install MinerU with support for both backends and process documents through the evaluation pipeline.
Installation
Install the complete package including core dependencies, VLM support, and pipeline components:
pip install "mineru[core,vlm,pipeline]"
Running Both Backends
Process the same document using both backends to generate comparable outputs:
Pipeline Backend:
mineru -p docs/example.pdf \
-o out/pipeline \
-b pipeline \
--method auto
VLM Backend:
mineru -p docs/example.pdf \
-o out/vlm \
-b vlm-auto-engine \
--method auto
Both commands generate middle.json files in their respective output directories. These JSON files contain the structured document representation that OmniDocBench evaluates.
Evaluation
Feed the generated middle.json files to the OmniDocBench evaluator (available in the separate OmniDocBench repository) to compute the End-to-End Evaluation Overall score. The evaluation compares your output against ground truth annotations to calculate precision, recall, and F1 scores for document structure recognition.
Implementation Details and Source Files
The accuracy results depend on specific implementations across the MinerU codebase:
| File | Relevance to Benchmark Accuracy |
|---|---|
docs/en/reference/changelog.md |
Documents MinerU 2.5's OmniDocBench v1.5 results and competitive comparisons |
README.md / README_zh-CN.md |
Provides benchmark overview and links to OmniDocBench methodology |
minerU/cli/vlm_server.py |
Implements VLM backend server for document processing |
minerU/cli/client.py |
Client interface for VLM backend operations |
minerU/backend/pipeline/*.py |
Pipeline backend implementation with staged extraction logic |
minerU/template.json |
Defines output schema including _backend field for evaluation |
The exact numeric accuracy scores are maintained in the external OmniDocBench repository rather than the MinerU codebase. The MinerU documentation references these results qualitatively, emphasizing the model's superior performance relative to competing solutions.
Summary
- MinerU 2.5 achieves state-of-the-art accuracy on OmniDocBench v1.5 using the End-to-End Evaluation Overall metric.
- With only 1.2B parameters, it surpasses larger multimodal models (Gemini 2.5 Pro, GPT-4o, Qwen 2.5-VL-72B) and dedicated OCR systems.
- Both pipeline and vlm backends share the same benchmark results, as evaluation occurs on the final
middle.jsonoutput. - Exact numeric scores reside in the external OmniDocBench repository; MinerU documentation provides qualitative superiority claims.
- Reproduce results by installing
mineru[core,vlm,pipeline]and processing documents through either backend for OmniDocBench evaluation.
Frequently Asked Questions
What is the exact accuracy score MinerU achieves on OmniDocBench?
The exact numeric value of the End-to-End Evaluation Overall score is not stored in the MinerU repository. According to the documentation in docs/en/reference/changelog.md, the specific quantitative results are maintained in the external OmniDocBench repository's JSON result files. The MinerU documentation emphasizes qualitative performance claims—specifically that version 2.5 surpasses competing models—rather than publishing specific precision or F1 percentages.
Does the VLM backend produce different accuracy results than the pipeline backend?
No, both backends share the same benchmark results on OmniDocBench v1.5. The evaluation measures accuracy on the final structured output (middle.json or content_list.json), not the intermediate processing steps. Whether you use the pipeline backend (minerU/backend/pipeline/*.py) or the VLM backend (minerU/cli/vlm_server.py), the End-to-End Evaluation Overall score remains identical because both use the same underlying MinerU 2.5 model and post-processing pipeline.
How does MinerU's 1.2B parameter model compare to larger multimodal models?
Despite having only 1.2 billion parameters, MinerU 2.5 surpasses significantly larger models on OmniDocBench. The changelog explicitly states that MinerU outperforms Gemini 2.5 Pro, GPT-4o, and Qwen 2.5-VL-72B (a 72-billion parameter model). This efficiency advantage makes MinerU deployable on consumer hardware and edge devices where resource-intensive models would be impractical, while maintaining superior document parsing accuracy.
Where can I find the official OmniDocBench evaluation results for MinerU?
The official quantitative results are maintained in the OmniDocBench repository, not the MinerU codebase. While MinerU's docs/en/reference/changelog.md and README.md reference these results and describe the model's superior performance, the specific JSON result files containing numeric scores reside in the external benchmark repository. To verify claims or access detailed precision/recall metrics, consult the OmniDocBench repository's results directory and compare the middle.json outputs generated by MinerU against ground truth annotations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →