How to Install MinerU Using Docker: Complete Setup Guide
To install MinerU using Docker, build the GPU-optimized image from the official docker/global/Dockerfile and deploy one of three service profiles—OpenAI-compatible server, REST API, or Gradio UI—using either docker run or Docker Compose.
The opendatalab/MinerU repository provides a production-ready containerization strategy that bundles the vLLM inference engine, Python dependencies, and pre-downloaded models into a single deployable unit. This guide walks through building the image from source, configuring GPU access, and launching the specific service tier that matches your integration requirements.
Docker Architecture and Base Image
The official MinerU Dockerfile at docker/global/Dockerfile extends the vllm/vllm-openai base image, which provides CUDA support for compute capabilities ≥ 8.0 (Ampere and newer architectures). The build process installs system dependencies including OpenCV’s libgl and Noto fonts for Chinese character rendering, then installs MinerU via pip install -U 'minerU[core]>=2.7.0'.
During image construction, the command minerU-models-download -s huggingface -m all pre-populates the container with every supported model. The Dockerfile entrypoint automatically sets MINERU_MODEL_SOURCE=local, ensuring the runtime loads these bundled models without requiring external network calls or volume mounts at startup.
Prerequisites
Before installing MinerU using Docker, verify your environment meets the following requirements:
- NVIDIA GPU with compute capability 8.0 or higher
- Docker Engine 20.10+ with the NVIDIA Container Toolkit installed
- CUDA drivers compatible with the vLLM base image
- 8GB+ VRAM recommended for standard models (adjustable via GPU memory utilization settings)
Building the MinerU Docker Image
Clone the repository and build the image locally to ensure you have the latest version with all model dependencies baked in:
git clone https://github.com/opendatalab/MinerU.git
cd MinerU
docker build -t mineru:latest -f docker/global/Dockerfile .
The build process downloads several gigabytes of model weights from Hugging Face. Once complete, the mineru:latest tag references a self-contained image ready for GPU-accelerated PDF processing.
Running Individual Services
The MinerU Docker image supports three distinct runtime modes. Each exposes a different interface but shares the same underlying image and model cache.
OpenAI-Compatible Server
Deploy the mineru-openai-server service on port 30000 to expose a local endpoint matching the OpenAI API schema:
docker run -d --gpus all \
-p 30000:30000 \
-e MINERU_MODEL_SOURCE=local \
--name mineru-openai-server \
mineru:latest \
mineru-openai-server --host 0.0.0.0 --port 30000
This mode is ideal for integrating MinerU into existing LLM pipelines that expect OpenAI-style /v1/chat/completions or structured extraction endpoints.
REST API Server
For a dedicated HTTP interface without OpenAI compatibility layers, launch the mineru-api service on port 8000:
docker run -d --gpus all \
-p 8000:8000 \
-e MINERU_MODEL_SOURCE=local \
--name mineru-api \
mineru:latest \
mineru-api --host 0.0.0.0 --port 8000
Navigate to http://localhost:8000/docs to access the interactive Swagger UI documenting all available endpoints.
Gradio Web UI
For interactive testing and manual PDF processing, run the mineru-gradio service on port 7860:
docker run -d --gpus all \
-p 7860:7860 \
-e MINERU_MODEL_SOURCE=local \
--name mineru-gradio \
mineru:latest \
mineru-gradio --server-name 0.0.0.0 --server-port 7860
Access the interface at http://localhost:7860 to upload documents and configure extraction parameters through a browser-based GUI.
Docker Compose Deployment
For production environments or multi-service orchestration, use the docker/compose.yaml file provided in the repository. The compose configuration defines three profiles that share the same mineru:latest image but expose different entry commands:
openai-server– Exposes port 30000 for OpenAI-compatible inferenceapi– Exposes port 8000 for the native REST APIgradio– Exposes port 7860 for the web interface
Launch a specific profile using the --profile flag:
# Start only the Gradio UI
docker compose --profile gradio up -d
# Start both the OpenAI server and REST API simultaneously
docker compose --profile openai-server --profile api up -d
The compose file automatically handles GPU device reservation and inherits the MINERU_MODEL_SOURCE=local environment variable from the image definition.
Advanced Configuration Options
Customize resource allocation and runtime behavior by appending parameters to the service commands in docker/compose.yaml or directly in docker run invocations:
- GPU Memory Utilization – Reduce KV-cache pressure with
--gpu-memory-utilization 0.5to fit large models on limited VRAM (e.g., 8GB cards) - Data Parallelism – Enable multi-GPU processing with
--data-parallel-size 2(requires listing multiple device IDs in thedevice_idsarray) - API Disabling – Add
--enable-api falseto the Gradio service to prevent automatic generation of the/apiendpoint - Page Limits – Cap processing with
--max-convert-pages 20to prevent timeout errors on extremely large PDFs
These options appear as commented examples in docker/compose.yaml at lines 13–17, 44–48, and 70–74.
Verifying Your Installation
Confirm successful deployment using these health check commands:
- OpenAI Server:
curl http://localhost:30000/healthreturns{"status":"healthy"} - API Server: Open
http://localhost:8000/docsto verify the Swagger documentation loads - Gradio UI: Navigate to
http://localhost:7860and upload a test PDF to confirm GPU inference activates without errors
Summary
- MinerU provides an official Docker image based on
vllm/vllm-openaiwith CUDA support for modern NVIDIA GPUs - The build process at
docker/global/Dockerfileinstalls MinerU core, system dependencies, and pre-downloads all models usingminerU-models-download - Three service profiles are available: OpenAI-compatible server (port 30000), REST API (port 8000), and Gradio UI (port 7860)
- Use
docker runfor single-service deployment ordocker compose --profile <name> upfor orchestrated multi-service stacks - The container automatically sets
MINERU_MODEL_SOURCE=local, eliminating the need for external model volume mounts
Frequently Asked Questions
What GPU requirements are needed to run MinerU in Docker?
MinerU requires an NVIDIA GPU with compute capability 8.0 or higher (Ampere, Ada Lovelace, or Hopper architectures). The vLLM base image does not support older Pascal or Turing cards (compute capability < 8.0). You need at least 8GB of VRAM for standard usage, though you can reduce memory requirements by adjusting the --gpu-memory-utilization parameter.
Can I run MinerU Docker without a GPU?
No. The official docker/global/Dockerfile is built on vllm/vllm-openai, which requires NVIDIA GPU access via the --gpus flag. CPU-only inference is not supported in the containerized distribution. For CPU deployment, you must install MinerU directly via pip and configure CPU-compatible model backends.
How do I update the models inside the Docker container?
Models are baked into the image during the build process via the minerU-models-download -s huggingface -m all command. To update models, you must rebuild the Docker image from the Dockerfile, which pulls the latest model weights from Hugging Face. Alternatively, mount a host volume containing updated models to /root/.cache/mineru and adjust the MINERU_MODEL_SOURCE environment variable accordingly.
Which service profile should I choose for production API deployments?
Use the api profile (port 8000) for production REST API deployments, as it provides the cleanest HTTP interface without the overhead of OpenAI compatibility translation. For integrations requiring OpenAI SDK compatibility (e.g., LangChain or OpenAI client libraries), use the openai-server profile (port 30000). The gradio profile is intended for interactive testing and demonstrations only.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →