System Requirements for olmOCR: Hardware, Software, and Installation Guide

olmOCR requires Python 3.11 or higher, Linux system dependencies including poppler-utils and Microsoft fonts, and an NVIDIA GPU with at least 12GB VRAM for local inference, though a 30GB Docker image provides a self-contained alternative.

olmOCR is an open-source PDF text extraction tool developed by the Allen Institute for AI. Meeting the system requirements for olmocr involves configuring OS-level packages, Python runtime constraints, and optional GPU hardware as defined in the repository's configuration files. The following sections break down specifications found in README.md and pyproject.toml to ensure a successful deployment.

Operating System Dependencies

olmOCR relies on system-level PDF rendering libraries to convert pages into images. According to README.md lines 89–96, Ubuntu and Debian systems require poppler-utils and several font packages.

Install these dependencies using apt:

sudo apt-get update
sudo apt-get install poppler-utils \
    ttf-mscorefonts-installer msttcorefonts \
    fonts-crosextra-caladea fonts-crosextra-carlito \
    gsfonts lcdf-typetools

These packages provide the rendering backend that olmocr/pipeline.py uses to process PDF documents.

Python Environment Setup

The project mandates Python 3.11 or newer. The maintainers recommend using a clean Conda environment to avoid dependency conflicts, as documented in README.md lines 99–105.

Create and activate the environment:

conda create -n olmocr python=3.11
conda activate olmocr

While the code may function on Python 3.10, the pyproject.toml explicitly targets Python 3.11+ for full compatibility.

Hardware Requirements

GPU Specifications

For local inference using the olmocr/pipeline.py entry point, you need an NVIDIA GPU with at least 12GB of VRAM. The README.md lines 120–124 confirm testing on RTX 4090, L40S, A100, and H100 GPUs.

Storage Requirements

Allocate at least 30GB of free disk space to accommodate model weights and intermediate processing files. The Docker image alleninstituteforai/olmocr:latest-with-model alone consumes approximately 30GB when pulled.

Python Package Installation

Core dependencies are defined in pyproject.toml lines 21–41, including pypdf, pypdfium2, torch, and transformers. The repository provides optional extras for different deployment scenarios, defined in pyproject.toml lines 52–92.

Install based on your use case:

Remote inference only (no GPU):

pip install olmocr

Local GPU inference:

pip install "olmocr[gpu]" --extra-index-url https://download.pytorch.org/whl/cu128

For additional inference speed, optionally install flash-infer:

pip install https://download.pytorch.org/whl/cu128/flashinfer/flashinfer_python-0.2.5%2Bcu128torch2.7-cp38-abi3-linux_x86_64.whl

Beaker cluster support:

pip install "olmocr[gpu,beaker]" --extra-index-url https://download.pytorch.org/whl/cu128

The gpu extra adds PyTorch, Transformers, and vLLM, while beaker includes the Beaker SDK for distributed processing and bench provides benchmarking utilities.

Docker Deployment Option

If you prefer not to install system dependencies manually, use the pre-built Docker image containing all requirements and model weights.

Pull and run the container:

docker pull alleninstituteforai/olmocr:latest-with-model
docker run --gpus all -v $(pwd):/workspace \
    alleninstituteforai/olmocr:latest-with-model \
    -c "olmocr /workspace/output --markdown --pdfs /workspace/sample.pdf"

This image bundles poppler-utils, fonts, Python 3.11, and the olmOCR model into a approximately 30GB package, eliminating the need for local environment setup.

Summary

  • Operating System: Ubuntu/Debian with poppler-utils, Microsoft fonts, and Crosextra fonts installed
  • Python: Version 3.11 or newer in a clean virtual environment
  • Hardware: NVIDIA GPU with 12GB+ VRAM (RTX 4090, L40S, A100, H100 tested) and 30GB+ disk space
  • Packages: Install via pip with optional [gpu], [beaker], or [bench] extras defined in pyproject.toml
  • Alternative: Use the 30GB Docker image alleninstituteforai/olmocr:latest-with-model for containerized deployment

Frequently Asked Questions

Can I run olmOCR without a GPU?

Yes, but you must use remote inference or CPU-only mode. Install the base package with pip install olmocr without the [gpu] extra. Note that CPU inference is significantly slower than GPU-accelerated processing and is not recommended for large document batches.

What Python versions are compatible with olmOCR?

The pyproject.toml specifies Python 3.11 or higher. While the code may execute on Python 3.10, the maintainers explicitly recommend version 3.11 via Conda to ensure all dependencies resolve correctly.

Is macOS or Windows supported?

The system requirements for olmocr are documented specifically for Linux (Ubuntu/Debian). While Python packages may install on macOS or Windows, the poppler-utils and font dependencies require Linux-specific packages. Use the Docker image for cross-platform compatibility.

How much disk space do I need for the model files?

The Docker image with model weights requires approximately 30GB. For manual installations, ensure at least 30GB of free space to accommodate both the model checkpoints and temporary processing files generated during PDF conversion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →