How to Extract Text Layer from PDF for ATS Verification: A Complete Guide

The ai-job-search repository provides a dual-engine PDF text extractor in tools/verify_pdf.py that uses pypdf as the primary parser and falls back to Poppler's pdftotext to retrieve exactly the text layer Applicant Tracking Systems parse, enabling validation of page count, keyword presence, and overall parseability before submission.

Extracting the machine-readable text layer from a PDF is essential for ATS verification, as Applicant Tracking Systems read the embedded text rather than the visual rendering. The ai-job-search open-source repository provides a robust utility specifically designed to extract this text layer and validate CV compatibility according to the actual source code available in the MadsLorentzen/ai-job-search repository. This guide walks through the dual-engine extraction strategy and demonstrates how to use it for reliable pre-submission verification.

Understanding the Two-Stage PDF Text Extraction Strategy

The tools/verify_pdf.py script implements a resilient two-stage extraction process designed to maximize compatibility across different PDF generation methods. This approach ensures you can always extract the ATS-readable text layer even when one extraction method fails.

Primary Extraction with pypdf (_extract_pypdf)

The script first attempts extraction using the pure-Python pypdf library through the internal _extract_pypdf function. According to the source code at lines 54-66, this method opens the PDF with PdfReader, iterates over all pages, and calls page.extract_text() to produce a single string containing every page's text.

If pypdf is not installed, raises an exception, or returns empty text after whitespace normalization, the function returns None to trigger the fallback mechanism. This primary method is preferred because it requires no external dependencies beyond Python packages.

Fallback Extraction with Poppler pdftotext (_extract_pdftotext)

When the primary extractor fails, the script automatically invokes the external pdftotext command via _extract_pdftotext (lines 72-77). This fallback executes the command with -layout -enc UTF-8 flags to preserve column layout and ensure proper UTF-8 encoding.

Simultaneously, the script calls pdfinfo to obtain the page count, ensuring that even if textual extraction encounters issues, the page count validation remains available. This dual-validation approach prevents false negatives during the verification process.

The Core Extraction API: extract_text_layer

Both extractors return standardized results through the public extract_text_layer wrapper function (lines 80-88). This function returns a tuple containing:

  • text: The extracted ATS-readable content as a single string
  • pages: The total page count as an integer
  • extractor_name: A string identifier ("pypdf" or "pdftotext") indicating which backend succeeded

This transparency allows calling code to inspect which extraction method was used, which is valuable for debugging PDF compatibility issues across different generation tools.

Validating PDFs for ATS Compliance with verify_pdf

The high-level verify_pdf function orchestrates the complete validation workflow, combining text extraction with multiple compliance checks as implemented in the source code starting at line 90.

File Validation and Extraction

The verification process begins by validating file existence (lines 92-94), raising a VerificationError if the specified path is not a valid file. It then calls extract_text_layer to obtain both the text content and page count (lines 95-96), ensuring the PDF is parseable before proceeding with content validation.

Optional Text Dumping for Manual Inspection

Before applying any checks, the script optionally writes the extracted text to a user-specified file (lines 98-107) when the dump_text parameter is provided. This raw text dump is invaluable for manual ATS inspection, allowing you to view exactly what an automated system would see when parsing your CV.

Automated Checks for ATS Requirements

The verify_pdf function performs three critical validation steps:

  1. Page-count verification (lines 111-115): When expected_pages is supplied, the script compares it against the actual count obtained during extraction
  2. Minimum-character validation (lines 116-122): Counts non-whitespace characters after applying normalize_text to ensure the CV contains sufficient content
  3. Required-text validation (lines 123-128): Ensures each string in required_text appears in the extracted layer after normalization, verifying critical keywords like "Professional Experience" or specific skills are machine-readable

On success, the function returns (extractor, extracted_text, actual_pages) for downstream reporting (lines 128-129).

Practical Implementation Examples

Command-Line Verification

The most common usage pattern involves running the script directly from the command line to verify a CV before submission:

python tools/verify_pdf.py cv/main_example.pdf \
  --pages 2 \
  --min-chars 100 \
  --contains "Professional Experience" \
  --dump-text cv/main_example.txt

This command attempts pypdf extraction first, falls back to pdftotext if needed, validates that the PDF contains exactly 2 pages with at least 100 characters, checks for the required phrase, and saves the extracted text layer to cv/main_example.txt for manual review.

Python API Integration

For programmatic workflows, import the verification functions directly:

from tools.verify_pdf import verify_pdf, VerificationError

pdf_path = "cv/main_example.pdf"
try:
    extractor, text, pages = verify_pdf(
        pdf_path,
        expected_pages=2,
        min_chars=100,
        required_text=("Professional Experience",),
        dump_text="cv/main_example.txt"
    )
    print(f"✅ {pdf_path} verified (extractor={extractor}, pages={pages})")
except VerificationError as exc:
    print(f"❌ Verification failed: {exc}")

This pattern allows integration into automated application workflows, with VerificationError providing specific failure reasons including missing pages, insufficient text, absent keywords, or missing extraction tools.

Extracting Raw Text Only

When you need only the text layer without validation—for custom keyword scoring or analysis—use the extraction function directly:

from tools.verify_pdf import extract_text_layer

text, pages, extractor = extract_text_layer("cv/main_example.pdf")
print(f"Extractor: {extractor}, Pages: {pages}")
print("First 200 characters of the ATS text layer:")
print(text[:200])

This approach is useful for integrating with custom ATS compatibility scoring systems beyond the built-in validation rules.

Summary

  • Dual-engine extraction: The ai-job-search tool uses pypdf as the primary extractor with pdftotext (Poppler) as a fallback to ensure reliable text layer extraction across all PDF types
  • Complete validation: The verify_pdf function checks file existence, page count, minimum character requirements, and required keyword presence in the ATS-readable text
  • Debugging support: The --dump-text flag and dump_text parameter allow inspection of the exact text string that Applicant Tracking Systems will parse
  • Graceful degradation: If both extractors are unavailable, the run_tool function provides OS-specific installation hints rather than crashing the workflow

Frequently Asked Questions

What is the difference between the pypdf and pdftotext extraction methods?

pypdf is a pure-Python library that requires no external system dependencies, making it ideal for cross-platform usage and primary extraction attempts. pdftotext (from the Poppler utilities) is an external command-line tool that often handles complex layouts and embedded fonts more effectively, serving as a robust fallback when pypdf returns empty or garbled text. According to the source code in tools/verify_pdf.py, the script attempts pypdf first via _extract_pypdf before falling back to _extract_pdftotext only when necessary.

How does the tool handle cases where both extraction methods fail?

If both pypdf and pdftotext are unavailable or return empty results, the run_tool function raises a VerificationError with operating-system-specific installation instructions. This allows the calling workflow—such as the apply step in the job application assistant—to gracefully downgrade to a visual keyword review instead of failing completely, ensuring you can still proceed with manual verification even when automated parsing is impossible.

Why is extracting the text layer specifically important for ATS verification?

Applicant Tracking Systems do not "see" the visual rendering of your CV; they read the embedded text layer contained within the PDF file structure. By extracting this specific layer using the same methods ATS parsers employ, the ai-job-search tool reveals exactly what keywords, formatting, and content the automated system will actually process. This prevents situations where visually present content is invisible to the ATS due to image-based PDFs, corrupted text encoding, or font subsetting issues.

Can I verify that specific keywords exist in my CV text layer?

Yes. The verify_pdf function accepts a required_text parameter (or --contains flags via CLI) that validates the presence of specific strings after text normalization. For example, you can require that "Python," "Data Analysis," and "Professional Experience" all appear in the extracted text layer. If any required text is missing, the function raises a VerificationError, alerting you that the ATS may not recognize those critical qualifications even if they appear visually in your document.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →