How to Use MinerU in HTTP Client Mode with OpenAI-Compatible Servers
Set the backend to vlm-http-client or hybrid-http-client and provide a server_url pointing to your OpenAI-compatible endpoint to offload VLM inference to a remote server while keeping the extraction pipeline local.
MinerU, the open-source document parsing toolkit from the opendatalab/MinerU repository, supports delegating heavy vision-language-model (VLM) inference to remote OpenAI-compatible services. By using MinerU in HTTP client mode, you can run the full extraction pipeline—including layout analysis, OCR, formula recognition, and table detection—on a CPU-only machine while the GPU-intensive model inference happens on a separate server.
How HTTP Client Mode Works in MinerU
When you select an HTTP-client backend, MinerU does not load any local model weights into memory. Instead, the ModelSingleton.get_model function in minerU/backend/vlm/vlm_analyze.py instantiates a MinerUClient object that stores the remote server_url and optional HTTP tuning parameters.
During document processing, the doc_analyze or aio_doc_analyze method calls predictor.batch_two_step_extract. The MinerUClient implementation (provided by the minerU-vl-utils package) translates each image batch into OpenAI-style chat/completion POST requests, sends them to the configured endpoint, and parses the JSON responses back into the format expected by MinerU’s post-processing pipeline.
Configuration and Backend Options
Selecting the Backend
MinerU provides two HTTP-client backend identifiers:
vlm-http-client: Uses a remote VLM for all visual inference tasks, including layout detection, reading order, and element classification.hybrid-http-client: Combines local lightweight processing with remote VLM calls, typically using the HTTP client only for heavy tasks like formula recognition or complex table structures while running OCR and basic layout locally.
Set the backend via the -b or --backend CLI flag, or the backend parameter in Python.
HTTP Client Parameters
When initializing the HTTP client, you can tune connection behavior. These parameters are defined in minerU/backend/vlm/vlm_analyze.py (lines 48-52) and passed directly to MinerUClient:
max_concurrency: Maximum parallel requests per batch (default varies by implementation).http_timeout: Seconds before a request aborts (default 120).max_retries: Retry count for transient HTTP errors.retry_backoff_factor: Exponential back-off multiplier between retries.
Usage Examples
Command Line Interface
The quickest way to use MinerU in HTTP client mode is via the CLI:
minerU \
-p my_document.pdf \
-o ./output \
-b vlm-http-client \
-u http://127.0.0.1:30000 \
-m auto \
-l ch
-b vlm-http-clientselects the HTTP-client backend.-u/--urlsupplies the OpenAI-compatible endpoint URL.-m autoand-l chset the parsing method and language as usual.
FastAPI Endpoint
If you are running MinerU as a service, submit jobs via the FastAPI endpoint exposed in minerU/cli/fast_api.py:
import requests
files = {"file": ("my_document.pdf", open("my_document.pdf", "rb"))}
data = {
"backend": "vlm-http-client",
"method": "auto",
"lang": "ch",
"server_url": "http://127.0.0.1:30000",
"output_dir": "./output",
}
r = requests.post("http://localhost:8000/api/v1/tasks/submit", files=files, data=data)
print(r.json())
The handler forwards server_url to the same do_parse routine used by the CLI, ensuring consistent behavior across interfaces.
Direct Python API
For embedded use, call do_parse directly from mineru.cli.common:
from mineru.cli.common import do_parse
from pathlib import Path
pdf_path = Path("my_document.pdf")
pdf_bytes = pdf_path.read_bytes()
do_parse(
output_dir="./output",
pdf_file_names=[pdf_path.stem],
pdf_bytes_list=[pdf_bytes],
p_lang_list=["ch"],
backend="vlm-http-client",
parse_method="auto",
formula_enable=True,
table_enable=True,
server_url="http://127.0.0.1:30000",
)
This bypasses the CLI layer and gives you full control over parsing parameters while still delegating VLM inference to the remote endpoint.
Tuning HTTP Client Settings
To optimize throughput or reliability, pass additional arguments to do_parse:
do_parse(
output_dir="./output",
pdf_file_names=["my_document"],
pdf_bytes_list=[pdf_bytes],
p_lang_list=["ch"],
backend="vlm-http-client",
server_url="http://127.0.0.1:30000",
max_concurrency=20, # parallel requests per batch
http_timeout=120, # seconds before a request aborts
max_retries=5, # retry count on transient errors
retry_backoff_factor=1.0, # exponential back‑off multiplier
)
These parameters are processed by ModelSingleton in minerU/backend/vlm/vlm_analyze.py and passed to the underlying MinerUClient instance.
Key Implementation Files
Understanding the source structure helps with debugging and advanced configuration:
minerU/cli/client.py: Parses CLI arguments including--backendand--url(server_url), then invokesdo_parse.minerU/backend/vlm/vlm_analyze.py: ContainsModelSingleton.get_modelwhich instantiatesMinerUClientfor HTTP-client backends (lines 202-216) and defines HTTP tuning parameters (lines 48-52).minerU/cli/fast_api.py: Exposes the parsing functionality via FastAPI, acceptingserver_urlas a form field and forwarding it to the backend.minerU/cli/common.py: Houses thedo_parsefunction used by CLI, FastAPI, and direct Python integrations.
Summary
- MinerU supports HTTP-client mode via the
vlm-http-clientandhybrid-http-clientbackends, enabling remote VLM inference. - Set the
server_urlparameter to point to any OpenAI-compatible endpoint (OpenAI, Azure, vLLM, etc.). - The
MinerUClientclass handles batch translation of images to OpenAI chat/completion requests without loading local models. - Configure throughput and reliability via
max_concurrency,http_timeout,max_retries, andretry_backoff_factor. - All interfaces—CLI, FastAPI, and direct Python—support the same HTTP-client configuration through
minerU/cli/client.py,minerU/cli/fast_api.py, andminerU/cli/common.py.
Frequently Asked Questions
What is the difference between vlm-http-client and hybrid-http-client?
The vlm-http-client backend delegates all vision-language model inference to the remote server, including layout detection, reading order, and element classification. The hybrid-http-client backend combines local processing with remote calls, typically using the HTTP client only for computationally intensive tasks like formula recognition or complex table structures while running OCR and basic layout analysis locally.
Do I need a GPU on the client machine when using HTTP client mode?
No. When using vlm-http-client or hybrid-http-client, MinerU does not load local model weights into GPU memory. The heavy inference is performed on the remote server specified by server_url, allowing the client machine to run on CPU-only resources while still achieving full document parsing capabilities.
Which OpenAI-compatible servers work with MinerU?
MinerU's HTTP client mode works with any server implementing the OpenAI chat/completions API format, including OpenAI's official API, Azure OpenAI Service, and open-source alternatives like vLLM, TGI (Text Generation Inference), and local LLM servers exposing an OpenAI-compatible REST interface. The server_url should point to the base URL of the inference endpoint.
How do I handle authentication for the remote server?
The MinerUClient passes additional keyword arguments directly to the underlying HTTP client implementation. If your OpenAI-compatible server requires API keys or custom headers, include them in the do_parse call or CLI configuration using the standard environment variables or additional parameters supported by your specific HTTP client setup. For OpenAI-compatible endpoints requiring Bearer tokens, ensure the server_url includes the necessary path or the client is configured with the appropriate authorization headers through the extended parameter set.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →