pdftotext vs Docling Performance Benchmark for PDF Extraction: 1,000x Speed Difference Explained
pdftotext extracts a 103-page PDF in ~0.1 seconds, while Docling takes ~164 seconds (~1.5 seconds per page), making pdftotext roughly 1,600× faster for text-heavy documents.
The virgiliojr94/book-to-skill repository implements dual extraction backends that trade raw speed against structural fidelity. Understanding this performance gap helps developers choose the right tool for their specific PDF content — whether fast plaintext or parsed technical documents with tables and code blocks.
The 1,600× Performance Gap: Measured Results
According to the repository's performance documentation in docs/performance.md, extracting a 103-page technical PDF yields starkly different timings:
| Backend | Time | Architecture |
|---|---|---|
| pdftotext | ≈ 0.1 s (instant) | Compiled C binary system call + lightweight Python post-processing |
| Docling | ≈ 164 s (~1.5 s/page) | Pure Python layout pipeline with Markdown AST generation |
These measurements appear in docs/performance.md lines 27-34, establishing the baseline comparison for production workloads.
Why pdftotext Is Instant
The pdftotext backend delegates all heavy work to the compiled C utility from poppler-utils. In book_to_skill/parsers/pdf.py lines 76-90, the extract_with_pdftotext function:
- Invokes the system binary with flags
-layout -enc UTF-8 - Captures the plain-text stream where form-feed characters delimit pages
- Runs
clean_pdftotextto strip headers, footers, page numbers, and de-hyphenate wrapped lines
Post-processing remains minimal because the binary returns already-cleaned text. No layout analysis occurs — structure is flattened entirely, meaning tables, code blocks, and visual formatting are lost.
Why Docling Is Slower but Structure-Preserving
The extract_with_docling implementation in book_to_skill/parsers/pdf.py lines 45-68 instantiates Docling's DocumentConverter with pipeline options:
from docling.datamodel.document import ConversionResult
from docling.datamodel.base_models import InputFormat
from docling.datamodel.document import PdfPipelineOptions
options = PdfPipelineOptions(
do_ocr=False, # Skip OCR for digital PDFs
do_table_structure=True # Preserve table layout
)
This Python-based pipeline parses page-by-page, analyzes spatial layout, detects tables and code blocks, and constructs a Markdown AST before emitting final text. The ~1.5 seconds per page overhead buys:
- Native Markdown tables from PDF tabular content
- Preserved code block formatting with language detection
- Document hierarchy (headings, sections) maintained
Choosing Between Extraction Modes
The high-level extract_text utility in book_to_skill/utils.py abstracts this choice through a mode parameter:
| Use Case | Recommended Mode | Backend | Rationale |
|---|---|---|---|
| Novels, prose, plain articles | mode="text" |
pdftotext | Maximum throughput, structure unnecessary |
| Programming books, technical manuals, research papers | mode="technical" |
Docling | Tables, formulas, and code blocks preserved |
Practical Code Examples
Select your backend explicitly through the mode parameter:
from book_to_skill.utils import extract_text
pdf_path = "/path/to/book.pdf"
# Fast extraction for text-heavy content
plaintext = extract_text(pdf_path, mode="text")
print(f"Characters extracted: {len(plaintext)}")
# Structure-preserving extraction for technical documents
markdown = extract_text(pdf_path, mode="technical")
print(f"Markdown preview:\n{markdown[:500]}")
For direct backend access, import from the parser module:
from book_to_skill.parsers.pdf import extract_with_pdftotext, extract_with_docling
# Direct pdftotext call
raw_text = extract_with_pdftotext("document.pdf")
# Direct Docling call with full configuration
structured_md = extract_with_docling("document.pdf")
Performance Optimization Tips
- Batch processing: Docling's per-page overhead makes it prohibitively slow for large archives — consider parallelizing across files rather than pages
- OCR disablement: The repository explicitly sets
do_ocr=FalseinPdfPipelineOptionsfor digital PDFs, avoiding unnecessary Tesseract initialization - Hybrid workflows: For mixed corpora, run
mode="text"first, then re-process failures withmode="technical"only when structure detection fails
Summary
- pdftotext achieves ~0.1 s extraction via compiled C binary delegation, flattening all layout
- Docling requires ~1.5 s per page for Python-based layout analysis preserving tables and code blocks
- The ~1,600× speed difference stems from architectural trade-off: system binary versus AST-generating pipeline
- Choose
mode="text"(pdftotext) for prose;mode="technical"(Docling) for technical books requiring structure preservation - Implementation resides in
book_to_skill/parsers/pdf.pywith dispatcher logic inbook_to_skill/utils.py
Frequently Asked Questions
How much slower is Docling than pdftotext for PDF extraction?
Docling requires approximately 164 seconds to extract a 103-page technical PDF, while pdftotext completes the same document in 0.1 seconds. This represents roughly 1,600× slower performance, or about 1.5 seconds of processing time per page in Docling versus near-instant completion with pdftotext.
Why is pdftotext so much faster than Docling?
pdftotext delegates extraction to a compiled C binary (poppler-utils) that streams raw text directly, with only lightweight Python post-processing for header/footer removal. Docling runs a pure Python pipeline that parses PDF layout page-by-page, detects structural elements like tables and code blocks, and constructs a Markdown AST — significantly more CPU-intensive work that preserves document semantics.
When should I use Docling instead of pdftotext?
Use Docling's mode="technical" when your PDF contains tables, code blocks, mathematical formulas, or hierarchical document structure that must be preserved in the output. The repository configures PdfPipelineOptions with do_table_structure=True specifically for technical books where Markdown-native table representation matters more than extraction speed.
Where is the PDF extraction logic implemented in book-to-skill?
The dual backends are implemented in book_to_skill/parsers/pdf.py: extract_with_pdftotext occupies lines 76-90, while extract_with_docling spans lines 45-68. The high-level dispatcher extract_text in book_to_skill/utils.py routes to these implementations based on the mode parameter. Benchmark data documenting the performance comparison lives in docs/performance.md lines 27-34.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →