# How to Use MinerU in HTTP Client Mode with OpenAI-Compatible Servers

> Learn to use MinerU in HTTP client mode with OpenAI-compatible servers. Offload VLM inference to remote servers for efficient processing.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: how-to-guide
- Published: 2026-02-23

---

**Set the backend to `vlm-http-client` or `hybrid-http-client` and provide a `server_url` pointing to your OpenAI-compatible endpoint to offload VLM inference to a remote server while keeping the extraction pipeline local.**

MinerU, the open-source document parsing toolkit from the opendatalab/MinerU repository, supports delegating heavy vision-language-model (VLM) inference to remote OpenAI-compatible services. By using MinerU in HTTP client mode, you can run the full extraction pipeline—including layout analysis, OCR, formula recognition, and table detection—on a CPU-only machine while the GPU-intensive model inference happens on a separate server.

## How HTTP Client Mode Works in MinerU

When you select an HTTP-client backend, MinerU does not load any local model weights into memory. Instead, the `ModelSingleton.get_model` function in [`minerU/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_analyze.py) instantiates a `MinerUClient` object that stores the remote `server_url` and optional HTTP tuning parameters.

During document processing, the `doc_analyze` or `aio_doc_analyze` method calls `predictor.batch_two_step_extract`. The `MinerUClient` implementation (provided by the `minerU-vl-utils` package) translates each image batch into OpenAI-style chat/completion POST requests, sends them to the configured endpoint, and parses the JSON responses back into the format expected by MinerU’s post-processing pipeline.

## Configuration and Backend Options

### Selecting the Backend

MinerU provides two HTTP-client backend identifiers:

- **`vlm-http-client`**: Uses a remote VLM for all visual inference tasks, including layout detection, reading order, and element classification.
- **`hybrid-http-client`**: Combines local lightweight processing with remote VLM calls, typically using the HTTP client only for heavy tasks like formula recognition or complex table structures while running OCR and basic layout locally.

Set the backend via the `-b` or `--backend` CLI flag, or the `backend` parameter in Python.

### HTTP Client Parameters

When initializing the HTTP client, you can tune connection behavior. These parameters are defined in [`minerU/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_analyze.py) (lines 48-52) and passed directly to `MinerUClient`:

- **`max_concurrency`**: Maximum parallel requests per batch (default varies by implementation).
- **`http_timeout`**: Seconds before a request aborts (default 120).
- **`max_retries`**: Retry count for transient HTTP errors.
- **`retry_backoff_factor`**: Exponential back-off multiplier between retries.

## Usage Examples

### Command Line Interface

The quickest way to use MinerU in HTTP client mode is via the CLI:

```bash
minerU \
  -p my_document.pdf \
  -o ./output \
  -b vlm-http-client \
  -u http://127.0.0.1:30000 \
  -m auto \
  -l ch

```

- `-b vlm-http-client` selects the HTTP-client backend.
- `-u/--url` supplies the OpenAI-compatible endpoint URL.
- `-m auto` and `-l ch` set the parsing method and language as usual.

### FastAPI Endpoint

If you are running MinerU as a service, submit jobs via the FastAPI endpoint exposed in [`minerU/cli/fast_api.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/fast_api.py):

```python
import requests

files = {"file": ("my_document.pdf", open("my_document.pdf", "rb"))}
data = {
    "backend": "vlm-http-client",
    "method": "auto",
    "lang": "ch",
    "server_url": "http://127.0.0.1:30000",
    "output_dir": "./output",
}
r = requests.post("http://localhost:8000/api/v1/tasks/submit", files=files, data=data)
print(r.json())

```

The handler forwards `server_url` to the same `do_parse` routine used by the CLI, ensuring consistent behavior across interfaces.

### Direct Python API

For embedded use, call `do_parse` directly from `mineru.cli.common`:

```python
from mineru.cli.common import do_parse
from pathlib import Path

pdf_path = Path("my_document.pdf")
pdf_bytes = pdf_path.read_bytes()

do_parse(
    output_dir="./output",
    pdf_file_names=[pdf_path.stem],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["ch"],
    backend="vlm-http-client",
    parse_method="auto",
    formula_enable=True,
    table_enable=True,
    server_url="http://127.0.0.1:30000",
)

```

This bypasses the CLI layer and gives you full control over parsing parameters while still delegating VLM inference to the remote endpoint.

### Tuning HTTP Client Settings

To optimize throughput or reliability, pass additional arguments to `do_parse`:

```python
do_parse(
    output_dir="./output",
    pdf_file_names=["my_document"],
    pdf_bytes_list=[pdf_bytes],
    p_lang_list=["ch"],
    backend="vlm-http-client",
    server_url="http://127.0.0.1:30000",
    max_concurrency=20,          # parallel requests per batch

    http_timeout=120,            # seconds before a request aborts

    max_retries=5,               # retry count on transient errors

    retry_backoff_factor=1.0,    # exponential back‑off multiplier

)

```

These parameters are processed by `ModelSingleton` in [`minerU/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_analyze.py) and passed to the underlying `MinerUClient` instance.

## Key Implementation Files

Understanding the source structure helps with debugging and advanced configuration:

- **[`minerU/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/client.py)**: Parses CLI arguments including `--backend` and `--url` (`server_url`), then invokes `do_parse`.
- **[`minerU/backend/vlm/vlm_analyze.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/vlm/vlm_analyze.py)**: Contains `ModelSingleton.get_model` which instantiates `MinerUClient` for HTTP-client backends (lines 202-216) and defines HTTP tuning parameters (lines 48-52).
- **[`minerU/cli/fast_api.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/fast_api.py)**: Exposes the parsing functionality via FastAPI, accepting `server_url` as a form field and forwarding it to the backend.
- **[`minerU/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/common.py)**: Houses the `do_parse` function used by CLI, FastAPI, and direct Python integrations.

## Summary

- MinerU supports **HTTP-client mode** via the `vlm-http-client` and `hybrid-http-client` backends, enabling remote VLM inference.
- Set the **`server_url`** parameter to point to any OpenAI-compatible endpoint (OpenAI, Azure, vLLM, etc.).
- The **`MinerUClient`** class handles batch translation of images to OpenAI chat/completion requests without loading local models.
- Configure throughput and reliability via **`max_concurrency`**, **`http_timeout`**, **`max_retries`**, and **`retry_backoff_factor`**.
- All interfaces—CLI, FastAPI, and direct Python—support the same HTTP-client configuration through [`minerU/cli/client.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/client.py), [`minerU/cli/fast_api.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/fast_api.py), and [`minerU/cli/common.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/cli/common.py).

## Frequently Asked Questions

### What is the difference between `vlm-http-client` and `hybrid-http-client`?

The `vlm-http-client` backend delegates all vision-language model inference to the remote server, including layout detection, reading order, and element classification. The `hybrid-http-client` backend combines local processing with remote calls, typically using the HTTP client only for computationally intensive tasks like formula recognition or complex table structures while running OCR and basic layout analysis locally.

### Do I need a GPU on the client machine when using HTTP client mode?

No. When using `vlm-http-client` or `hybrid-http-client`, MinerU does not load local model weights into GPU memory. The heavy inference is performed on the remote server specified by `server_url`, allowing the client machine to run on CPU-only resources while still achieving full document parsing capabilities.

### Which OpenAI-compatible servers work with MinerU?

MinerU's HTTP client mode works with any server implementing the OpenAI chat/completions API format, including OpenAI's official API, Azure OpenAI Service, and open-source alternatives like vLLM, TGI (Text Generation Inference), and local LLM servers exposing an OpenAI-compatible REST interface. The `server_url` should point to the base URL of the inference endpoint.

### How do I handle authentication for the remote server?

The `MinerUClient` passes additional keyword arguments directly to the underlying HTTP client implementation. If your OpenAI-compatible server requires API keys or custom headers, include them in the `do_parse` call or CLI configuration using the standard environment variables or additional parameters supported by your specific HTTP client setup. For OpenAI-compatible endpoints requiring Bearer tokens, ensure the `server_url` includes the necessary path or the client is configured with the appropriate authorization headers through the extended parameter set.