Purpose of the olmocr Repository: Technical Architecture and PDF-to-Markdown Pipeline
The olmocr repository is an open-source toolkit that converts PDFs, PNGs, and JPEGs into clean, structured Markdown text using vision-language models, preserving logical reading order, tables, and equations while processing documents at approximately $0.10 per million input tokens.
The allenai/olmocr repository provides a production-ready pipeline for optical character recognition (OCR) that leverages modern vision-language models (VLMs) to extract structured text from complex documents. Unlike traditional OCR tools that rely on heuristic layout analysis, olmocr renders each page as an image and processes it through a VLM to generate Dolma-compatible documents with preserved formatting. Understanding the purpose of the olmocr repository reveals how it enables high-throughput, low-cost conversion of massive document collections into training-ready text corpora.
Core Architecture and Design
Vision-Language Model Integration
At the heart of olmocr lies a VLM-based processing strategy implemented in olmocr/pipeline.py. The system converts each document page into a base64-encoded PNG using render_pdf_to_base64png, then transmits this image to an OpenAI-compatible API endpoint or local vLLM server. This approach eliminates the need for complex PDF parsing libraries while accurately capturing tables, mathematical equations, and multi-column layouts that traditional OCR often fails to parse.
Dolma-Compatible Output Format
The pipeline produces structured output following the Dolma document specification. The build_dolma_document function in pipeline.py assembles page-level text extracts into a single JSON document containing metadata fields such as natural_text, primary_language, and page spans. This format ensures compatibility with AI2's Dolma training corpus and other large-scale text processing workflows.
Key System Components
pipeline.py - Orchestration and Worker Management
The olmocr/pipeline.py file serves as the primary entry point, managing the end-to-end conversion workflow. It implements worker pools that handle PDF downloads, image rendering, VLM API calls with retry logic, and markdown export. The process_page function coordinates the conversion of individual pages, while main() initializes the workspace and spawns concurrent workers to maximize throughput.
work_queue.py - Distributed Processing Infrastructure
For massive-scale document processing, olmocr/work_queue.py abstracts both local and S3-backed queue implementations through the WorkQueue class. This component enables distributed processing across multiple machines by providing LocalBackend and S3Backend classes that manage work distribution. The queue system supports glob expansion for S3 paths, allowing workers to process millions of PDFs stored in cloud storage without local download bottlenecks.
prompts/ - Structured VLM Interactions
The olmocr/prompts/prompts.py module defines the YAML-based prompts sent to the vision-language model using build_no_anchoring_v4_yaml_prompt. This ensures consistent output formatting that the parser can reliably extract. Additionally, olmocr/prompts/anchor.py provides fallback anchor-text extraction using pdftotext to assist the VLM when processing rotated or complex documents.
Supporting Infrastructure
- metrics.py: Tracks token usage, request rates, and failures through
MetricsKeeperandWorkerTrackerclasses, providing real-time visibility into processing costs and performance. - s3_utils.py: Wraps boto3 functionality for downloading, uploading, and glob-expanding S3 paths, making the system storage-agnostic.
- filter/filter.py: Implements pre-processing filters to drop non-English documents or spam PDFs before expensive VLM inference, significantly reducing compute costs.
- image_utils.py: Handles conversion of image files to PDF bytes and automatic image type detection, enabling the pipeline to accept raw PNG and JPEG inputs.
- bench/: Contains the olmOCR-Bench benchmark suite with over 7,000 test cases across 1,400 documents for tracking accuracy regressions.
Pipeline Execution Flow
The conversion process follows a six-stage pipeline:
- Workspace Initialization:
main()parses CLI arguments and instantiates aWorkQueue(local or S3-backed) to distribute tasks across workers. - Document Ingestion: Workers download PDFs via
process_pdfor extract them from tarballs usingprocess_tarball. - Page Rendering: Each page is converted to a base64 PNG via
render_pdf_to_base64pngin preparation for VLM processing. - VLM Inference: The
build_page_queryfunction constructs the API request payload, with automatic retry logic for failed or rotated pages. - Response Parsing: The
FrontMatterParserextracts structured fields from the VLM's YAML response, includingnatural_textandis_tableindicators. - Output Generation:
build_dolma_documentassembles the final document, with optional markdown export viaget_markdown_pathwhen the--markdownflag is enabled.
Practical Usage Examples
Local GPU Conversion
Convert a single PDF using a local vLLM server:
# Install the core package
pip install olmocr
# Convert PDF to markdown
olmocr ./localworkspace --markdown --pdfs mypaper.pdf
This command launches the built-in vLLM server and writes output to ./localworkspace/markdown/mypaper.md.
Remote Inference Endpoint
Process documents using an external API server:
olmocr ./localworkspace \
--server http://my-vllm-host:8000/v1 \
--model allenai/olmOCR-2-7B-1025-FP8 \
--markdown \
--pdfs tests/gnarly_pdfs/*.pdf
All VLM calls route to the specified OpenAI-compatible endpoint, eliminating local GPU requirements.
Python API Integration
Embed the pipeline into larger applications:
import asyncio
from olmocr.pipeline import main
# Configure pipeline arguments
argv = [
"s3://my-bucket/workspace",
"--pdfs", "s3://my-bucket/pdfs/*.pdf",
"--model", "allenai/olmOCR-2-7B-1025-FP8",
"--workers", "8",
"--markdown"
]
# Execute async pipeline
asyncio.run(main(argv))
Low-Level Query Construction
Access the VLM query builder directly for custom implementations:
from olmocr.pipeline import build_page_query
import asyncio
query = asyncio.run(
build_page_query(
local_pdf_path="sample.pdf",
page=1,
target_longest_image_dim=1024,
model_name="allenai/olmOCR-2-7B-1025-FP8"
)
)
print(query["messages"][0]["content"])
This returns the exact JSON payload that would be POSTed to the VLM server.
Summary
- olmocr converts PDFs and images to structured Markdown using vision-language models, preserving complex layouts including tables and equations.
- The architecture centers on
olmocr/pipeline.py, which orchestrates workers, VLM API calls, and Dolma-compatible output generation. olmocr/work_queue.pyenables scalable processing across local and S3-backed distributed systems.- Pre-processing filters in
olmocr/filter/filter.pyreduce costs by excluding non-English or low-quality documents before VLM inference. - The system achieves approximately $0.10 per million input tokens when using the 7B parameter model, making it cost-effective for large-scale document conversion.
Frequently Asked Questions
What types of documents can olmocr process?
olmocr handles PDFs, PNGs, and JPEGs, converting them into clean Markdown or Dolma-compatible JSON. The vision-language model approach excels at preserving logical reading order in complex layouts including multi-column text, mathematical equations, and tables that traditional OCR tools often misparse.
How does olmocr handle large-scale document processing?
The repository implements a distributed work queue system in olmocr/work_queue.py supporting both local and S3-backed storage. This allows multiple workers across different machines to process millions of documents concurrently, with metrics.py providing real-time tracking of token usage and processing rates.
Can I use olmocr without a local GPU?
Yes. While olmocr can run a local vLLM server on GPU hardware, you can also point the tool to any OpenAI-compatible API endpoint using the --server flag. This allows lightweight deployment where the heavy inference runs on remote infrastructure while the local pipeline handles document routing and post-processing.
What is the cost of processing documents with olmocr?
According to the repository documentation and source code benchmarks, processing documents using the allenai/olmOCR-2-7B-1025-FP8 model costs approximately $0.10 per million input tokens. The pre-filtering stage in olmocr/filter/filter.py further reduces costs by eliminating non-English and spam documents before they reach the expensive VLM inference stage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →