# MinerU OmniDocBench Accuracy: Benchmark Results for Pipeline and VLM Backends

> Discover MinerU OmniDocBench accuracy. See how MinerU's pipeline and VLM backends outperform GPT-4o and Gemini 2.5 Pro with just 1.2B parameters. Get benchmark results now.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: performance
- Published: 2026-02-23

---

**MinerU 2.5 achieves state-of-the-art End-to-End Evaluation Overall scores on OmniDocBench v1.5, surpassing larger multimodal models like GPT-4o and Gemini 2.5 Pro with only 1.2B parameters.**

MinerU is an open-source document parsing toolkit that extracts structured content from PDFs and other formats. Understanding MinerU OmniDocBench accuracy is crucial for evaluating its performance against other document understanding solutions, as this benchmark represents the current standard for end-to-end document parsing evaluation.

## Understanding MinerU OmniDocBench Accuracy Metrics

### End-to-End Evaluation Overall Score

According to the repository's changelog at [`docs/en/reference/changelog.md`](https://github.com/opendatalab/MinerU/blob/main/docs/en/reference/changelog.md), MinerU 2.5 is evaluated using the **End-to-End Evaluation Overall** metric from OmniDocBench v1.5. This score measures the complete document parsing pipeline—from raw PDF input to structured JSON output—rather than evaluating individual components in isolation.

The evaluation compares the generated [`middle.json`](https://github.com/opendatalab/MinerU/blob/main/middle.json) and [`content_list.json`](https://github.com/opendatalab/MinerU/blob/main/content_list.json) files against ground truth annotations. These output files represent the final structured representation of the document, containing extracted text, tables, figures, and layout information.

### Model Architecture and Parameter Efficiency

MinerU 2.5 operates with **only 1.2 billion parameters**, making it significantly more efficient than competing solutions. Despite this compact architecture, the model achieves superior OmniDocBench accuracy compared to:

- **Gemini 2.5 Pro**
- **GPT-4o**
- **Qwen 2.5-VL-72B**
- **dots.ocr**
- **MonkeyOCR**
- **PP-StructureV3**

This efficiency makes MinerU particularly suitable for deployment environments with limited computational resources while maintaining state-of-the-art parsing accuracy.

## Pipeline vs VLM Backend Performance on OmniDocBench

### Shared Benchmark Results

Both the **pipeline** and **vlm** backends share identical benchmark results on OmniDocBench v1.5. The documentation does not publish separate accuracy numbers for each backend because the End-to-End Evaluation Overall score is computed on the final structured output, regardless of which backend generated it.

The **pipeline backend**, implemented in `minerU/backend/pipeline/*.py`, uses traditional computer vision and rule-based approaches combined with deep learning models. The **VLM backend**, implemented in [`minerU/cli/vlm_server.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/vlm_server.py) and [`minerU/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/client.py), leverages vision-language models for document understanding.

### Output Format and Evaluation Consistency

Both backends produce output conforming to the schema defined in [`minerU/template.json`](https://github.com/opendatalab/MinerU/blob/main/minerU/template.json), generating [`middle.json`](https://github.com/opendatalab/MinerU/blob/main/middle.json) files that OmniDocBench evaluates. While the internal processing differs:

- The **pipeline backend** processes documents through staged extraction pipelines
- The **VLM backend** uses `vlm-auto-engine` for end-to-end generation

The final accuracy metric remains consistent because evaluation occurs on the structured JSON output rather than intermediate representations.

## Competitive Comparison on OmniDocBench

MinerU 2.5's OmniDocBench accuracy positions it as the leading open-source document parsing solution. The changelog explicitly states that MinerU 2.5 **surpasses** both general-purpose multimodal models and specialized OCR systems:

**General Multimodal Models:**
- Outperforms Gemini 2.5 Pro despite having 100x fewer parameters
- Exceeds GPT-4o accuracy on document structure recognition
- Surpasses Qwen 2.5-VL-72B, a 72-billion parameter vision-language model

**Dedicated Document Parsing Models:**
- Outperforms dots.ocr on complex layout preservation
- Exceeds MonkeyOCR on table structure recognition
- Surpasses PP-StructureV3 (PaddlePaddle's document analysis system)

This performance is achieved while maintaining a lightweight 1.2B parameter footprint, making MinerU deployable on consumer hardware and edge devices where larger models would be impractical.

## Reproducing MinerU OmniDocBench Accuracy Results

To verify the benchmark results locally, install MinerU with support for both backends and process documents through the evaluation pipeline.

### Installation

Install the complete package including core dependencies, VLM support, and pipeline components:

```bash
pip install "mineru[core,vlm,pipeline]"

```

### Running Both Backends

Process the same document using both backends to generate comparable outputs:

**Pipeline Backend:**

```bash
mineru -p docs/example.pdf \
       -o out/pipeline \
       -b pipeline \
       --method auto

```

**VLM Backend:**

```bash
mineru -p docs/example.pdf \
       -o out/vlm \
       -b vlm-auto-engine \
       --method auto

```

Both commands generate [`middle.json`](https://github.com/opendatalab/MinerU/blob/main/middle.json) files in their respective output directories. These JSON files contain the structured document representation that OmniDocBench evaluates.

### Evaluation

Feed the generated [`middle.json`](https://github.com/opendatalab/MinerU/blob/main/middle.json) files to the OmniDocBench evaluator (available in the separate OmniDocBench repository) to compute the End-to-End Evaluation Overall score. The evaluation compares your output against ground truth annotations to calculate precision, recall, and F1 scores for document structure recognition.

## Implementation Details and Source Files

The accuracy results depend on specific implementations across the MinerU codebase:

| File | Relevance to Benchmark Accuracy |
|------|--------------------------------|
| [`docs/en/reference/changelog.md`](https://github.com/opendatalab/MinerU/blob/main/docs/en/reference/changelog.md) | Documents MinerU 2.5's OmniDocBench v1.5 results and competitive comparisons |
| [`README.md`](https://github.com/opendatalab/MinerU/blob/main/README.md) / [`README_zh-CN.md`](https://github.com/opendatalab/MinerU/blob/main/README_zh-CN.md) | Provides benchmark overview and links to OmniDocBench methodology |
| [`minerU/cli/vlm_server.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/vlm_server.py) | Implements VLM backend server for document processing |
| [`minerU/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/client.py) | Client interface for VLM backend operations |
| `minerU/backend/pipeline/*.py` | Pipeline backend implementation with staged extraction logic |
| [`minerU/template.json`](https://github.com/opendatalab/MinerU/blob/main/minerU/template.json) | Defines output schema including `_backend` field for evaluation |

The exact numeric accuracy scores are maintained in the external OmniDocBench repository rather than the MinerU codebase. The MinerU documentation references these results qualitatively, emphasizing the model's superior performance relative to competing solutions.

## Summary

- **MinerU 2.5** achieves state-of-the-art accuracy on **OmniDocBench v1.5** using the End-to-End Evaluation Overall metric.
- With only **1.2B parameters**, it surpasses larger multimodal models (Gemini 2.5 Pro, GPT-4o, Qwen 2.5-VL-72B) and dedicated OCR systems.
- Both **pipeline** and **vlm** backends share the same benchmark results, as evaluation occurs on the final [`middle.json`](https://github.com/opendatalab/MinerU/blob/main/middle.json) output.
- Exact numeric scores reside in the external OmniDocBench repository; MinerU documentation provides qualitative superiority claims.
- Reproduce results by installing `mineru[core,vlm,pipeline]` and processing documents through either backend for OmniDocBench evaluation.

## Frequently Asked Questions

### What is the exact accuracy score MinerU achieves on OmniDocBench?

The exact numeric value of the End-to-End Evaluation Overall score is not stored in the MinerU repository. According to the documentation in [`docs/en/reference/changelog.md`](https://github.com/opendatalab/MinerU/blob/main/docs/en/reference/changelog.md), the specific quantitative results are maintained in the external OmniDocBench repository's JSON result files. The MinerU documentation emphasizes qualitative performance claims—specifically that version 2.5 surpasses competing models—rather than publishing specific precision or F1 percentages.

### Does the VLM backend produce different accuracy results than the pipeline backend?

No, both backends share the same benchmark results on OmniDocBench v1.5. The evaluation measures accuracy on the final structured output ([`middle.json`](https://github.com/opendatalab/MinerU/blob/main/middle.json) or [`content_list.json`](https://github.com/opendatalab/MinerU/blob/main/content_list.json)), not the intermediate processing steps. Whether you use the **pipeline** backend (`minerU/backend/pipeline/*.py`) or the **VLM** backend ([`minerU/cli/vlm_server.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/vlm_server.py)), the End-to-End Evaluation Overall score remains identical because both use the same underlying MinerU 2.5 model and post-processing pipeline.

### How does MinerU's 1.2B parameter model compare to larger multimodal models?

Despite having only **1.2 billion parameters**, MinerU 2.5 surpasses significantly larger models on OmniDocBench. The changelog explicitly states that MinerU outperforms **Gemini 2.5 Pro**, **GPT-4o**, and **Qwen 2.5-VL-72B** (a 72-billion parameter model). This efficiency advantage makes MinerU deployable on consumer hardware and edge devices where resource-intensive models would be impractical, while maintaining superior document parsing accuracy.

### Where can I find the official OmniDocBench evaluation results for MinerU?

The official quantitative results are maintained in the **OmniDocBench** repository, not the MinerU codebase. While MinerU's [`docs/en/reference/changelog.md`](https://github.com/opendatalab/MinerU/blob/main/docs/en/reference/changelog.md) and [`README.md`](https://github.com/opendatalab/MinerU/blob/main/README.md) reference these results and describe the model's superior performance, the specific JSON result files containing numeric scores reside in the external benchmark repository. To verify claims or access detailed precision/recall metrics, consult the OmniDocBench repository's results directory and compare the [`middle.json`](https://github.com/opendatalab/MinerU/blob/main/middle.json) outputs generated by MinerU against ground truth annotations.