How to Deploy MinerU for Production Integration: Complete Setup Guide
You can deploy MinerU as a Docker container, standalone FastAPI service, OpenAI-compatible server, or Gradio UI, exposing REST endpoints at /api/v1/tasks/submit for document processing and retrieval.
MinerU is a modular document-parsing platform maintained by the opendatalab/MinerU repository. When you deploy MinerU for integration with other applications, you gain access to programmable HTTP endpoints that convert PDFs and images into structured Markdown via a scalable, GPU-accelerated backend architecture.
Understanding the MinerU Deployment Architecture
The platform separates concerns across three distinct layers defined in the source code.
CLI Entry Points
The mineru command-line interface defined in mineru/cli/client.py provides four primary entry points: mineru (pipeline), mineru-api, mineru-openai-server, and mineru-gradio. These scripts parse environment variables and configuration files before launching the appropriate backend service.
Backend Inference Engines
The system supports multiple inference backends through mineru/cli/vlm_server.py, which automatically selects between VLLM and LMDeploy engines. This layer exposes a /v1/chat/completions compatible endpoint when operating in OpenAI server mode, enabling integration with standard LLM client libraries.
FastAPI Service Layer
The production HTTP interface lives in projects/mineru_tianshu/api_server.py. This FastAPI application handles task submission, status polling via a SQLite database (TaskDB in task_db.py), and optional MinIO image storage for extracted assets.
Four Methods to Deploy MinerU
Docker Compose Deployment (Recommended)
For production environments requiring resource isolation, use the pre-configured docker/compose.yaml. The configuration mounts GPU devices via the NVIDIA Container Toolkit and exposes port 8000 for the REST API.
# docker/compose.yaml
services:
mineru-api:
image: mineru:latest # Build locally or pull from a registry
container_name: mineru-api
restart: always
ports:
- 8000:8000 # FastAPI REST interface
environment:
MINERU_MODEL_SOURCE: local # Use local model files (or "modelscope")
entrypoint: mineru-api
command:
--host 0.0.0.0
--port 8000
ulimits:
memlock: -1
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["0"] # Adjust for multi‑GPU
capabilities: [gpu]
Run the deployment:
docker compose -f docker/compose.yaml --profile api up -d
The API will be reachable at http://localhost:8000/docs (FastAPI Swagger UI).
Standalone FastAPI Server
Run mineru-api directly on the host machine for development or lightweight deployments. The CLI invokes uvicorn.run(app, ...) binding to 0.0.0.0:8000 without container overhead.
OpenAI-Compatible Server Mode
Execute mineru-openai-server to launch a VLLM or LMDeploy backend that mimics the OpenAI chat completions API. This mode enables drop-in replacement for existing LLM clients.
# Install optional VLLM dependency
uv pip install "mineru[all]" vllm
# Start the OpenAI‑compatible endpoint on port 30000
mineru-openai-server --host 0.0.0.0 --port 30000
Gradio Web Interface
Launch mineru-gradio to start a web UI on port 7860. This frontend internally calls the same backend APIs, providing a visual interface for document processing.
mineru-gradio --server-name 0.0.0.0 --server-port 7860
Open http://localhost:7860 in a browser to access the interface.
Configuring Your MinerU Deployment
All deployment modes respect environment variables prefixed with MINERU_, such as MINERU_MODEL_SOURCE, MINERU_BACKEND, and MINERU_DEVICE. The system also reads from a user-level mineru.json configuration file generated by mineru-models-download or copied from mineru.template.json.
Integrating Applications with the MinerU API
The typical integration flow involves submitting documents to /api/v1/tasks/submit, polling /api/v1/tasks/{task_id} for completion, and fetching results from /api/v1/tasks/{task_id}/data.
import requests
api_url = "http://localhost:8000/api/v1/tasks/submit"
pdf_path = "sample.pdf"
files = {"file": open(pdf_path, "rb")}
data = {
"backend": "pipeline",
"lang": "en",
"method": "auto",
"formula_enable": "true",
"table_enable": "true",
"priority": 10,
}
r = requests.post(api_url, files=files, data=data)
task = r.json()
task_id = task["task_id"]
print(f"Submitted, task_id={task_id}")
# Poll until completed
import time
while True:
status = requests.get(f"http://localhost:8000/api/v1/tasks/{task_id}").json()
if status["status"] == "completed":
break
time.sleep(2)
# Fetch full result (markdown + images)
result = requests.get(
f"http://localhost:8000/api/v1/tasks/{task_id}/data",
params={"include_fields": "md,images", "upload_images": "true"},
).json()
print(result["data"]["markdown"]["content"])
Set upload_images=true to store extracted images in MinIO and receive public URLs in the response.
Deploying the OpenAI-Compatible Endpoint
To integrate with LangChain or existing OpenAI SDK clients, start the compatible server and point your client to the local endpoint.
from langchain.llms import OpenAI
# Point to the local server
llm = OpenAI(model_name="gpt-4o-mini", openai_api_base="http://localhost:30000/v1")
response = llm.invoke("请把以下 PDF 内容转成 markdown:<PDF_URL>")
print(response)
Summary
- Deploy MinerU via Docker Compose for production isolation or run
mineru-apidirectly for development environments - Configure deployments using
MINERU_environment variables and themineru.jsonfile located in the user directory - Use the FastAPI endpoints at
/api/v1/tasks/submitfor asynchronous document processing with SQLite-backed task tracking - Enable OpenAI-compatible mode for seamless integration with LangChain and existing LLM frameworks
- Store extracted images in MinIO by setting
upload_images=truein API requests to offload asset storage
Frequently Asked Questions
What are the hardware requirements for deploying MinerU?
MinerU requires GPU acceleration for optimal performance, with support for NVIDIA GPUs via the CUDA runtime. The Docker Compose configuration uses the nvidia driver with device reservations, while CPU-only modes may be available for smaller documents depending on the backend configuration in mineru.json.
How do I configure GPU access in Docker deployments?
Modify the deploy.resources.reservations.devices section in docker/compose.yaml to specify driver: nvidia and the appropriate device_ids. Set capabilities: [gpu] to enable CUDA access within the container, and adjust device_ids: ["0"] to target specific GPUs in multi-GPU systems.
Can MinerU replace OpenAI endpoints in existing applications?
Yes. Running mineru-openai-server creates a compatible endpoint at /v1/chat/completions using VLLM or LMDeploy backends as implemented in mineru/cli/vlm_server.py. You can point any OpenAI SDK or LangChain client to http://localhost:30000/v1 to process documents through the chat completions interface.
Where does MinerU store processing results and logs?
The FastAPI server maintains task state in a lightweight SQLite database managed by TaskDB in task_db.py. Results persist locally unless you configure the optional MinIO integration by setting upload_images=true during task submission, which stores assets in external object storage. The server also provides an administrative endpoint at /api/v1/admin/cleanup for automatic removal of old results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →