# How to Extract Text Layer from PDF for ATS Verification: A Complete Guide

> Extract the PDF text layer for ATS verification with this complete guide. Learn how the ai-job-search repository validates page count and keyword presence using pypdf and pdftotext.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: how-to-guide
- Published: 2026-08-28

---

**The `ai-job-search` repository provides a dual-engine PDF text extractor in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) that uses `pypdf` as the primary parser and falls back to Poppler's `pdftotext` to retrieve exactly the text layer Applicant Tracking Systems parse, enabling validation of page count, keyword presence, and overall parseability before submission.**

Extracting the machine-readable text layer from a PDF is essential for ATS verification, as Applicant Tracking Systems read the embedded text rather than the visual rendering. The `ai-job-search` open-source repository provides a robust utility specifically designed to extract this text layer and validate CV compatibility according to the actual source code available in the MadsLorentzen/ai-job-search repository. This guide walks through the dual-engine extraction strategy and demonstrates how to use it for reliable pre-submission verification.

## Understanding the Two-Stage PDF Text Extraction Strategy

The [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) script implements a resilient two-stage extraction process designed to maximize compatibility across different PDF generation methods. This approach ensures you can always extract the ATS-readable text layer even when one extraction method fails.

### Primary Extraction with pypdf (_extract_pypdf)

The script first attempts extraction using the pure-Python `pypdf` library through the internal `_extract_pypdf` function. According to the source code at lines 54-66, this method opens the PDF with `PdfReader`, iterates over all pages, and calls `page.extract_text()` to produce a single string containing every page's text.

If `pypdf` is not installed, raises an exception, or returns empty text after whitespace normalization, the function returns `None` to trigger the fallback mechanism. This primary method is preferred because it requires no external dependencies beyond Python packages.

### Fallback Extraction with Poppler pdftotext (_extract_pdftotext)

When the primary extractor fails, the script automatically invokes the external `pdftotext` command via `_extract_pdftotext` (lines 72-77). This fallback executes the command with `-layout -enc UTF-8` flags to preserve column layout and ensure proper UTF-8 encoding.

Simultaneously, the script calls `pdfinfo` to obtain the page count, ensuring that even if textual extraction encounters issues, the page count validation remains available. This dual-validation approach prevents false negatives during the verification process.

## The Core Extraction API: extract_text_layer

Both extractors return standardized results through the public `extract_text_layer` wrapper function (lines 80-88). This function returns a tuple containing:

- **text**: The extracted ATS-readable content as a single string
- **pages**: The total page count as an integer
- **extractor_name**: A string identifier ("pypdf" or "pdftotext") indicating which backend succeeded

This transparency allows calling code to inspect which extraction method was used, which is valuable for debugging PDF compatibility issues across different generation tools.

## Validating PDFs for ATS Compliance with verify_pdf

The high-level `verify_pdf` function orchestrates the complete validation workflow, combining text extraction with multiple compliance checks as implemented in the source code starting at line 90.

### File Validation and Extraction

The verification process begins by validating file existence (lines 92-94), raising a `VerificationError` if the specified path is not a valid file. It then calls `extract_text_layer` to obtain both the text content and page count (lines 95-96), ensuring the PDF is parseable before proceeding with content validation.

### Optional Text Dumping for Manual Inspection

Before applying any checks, the script optionally writes the extracted text to a user-specified file (lines 98-107) when the `dump_text` parameter is provided. This **raw text dump** is invaluable for manual ATS inspection, allowing you to view exactly what an automated system would see when parsing your CV.

### Automated Checks for ATS Requirements

The `verify_pdf` function performs three critical validation steps:

1. **Page-count verification** (lines 111-115): When `expected_pages` is supplied, the script compares it against the actual count obtained during extraction
2. **Minimum-character validation** (lines 116-122): Counts non-whitespace characters after applying `normalize_text` to ensure the CV contains sufficient content
3. **Required-text validation** (lines 123-128): Ensures each string in `required_text` appears in the extracted layer after normalization, verifying critical keywords like "Professional Experience" or specific skills are machine-readable

On success, the function returns `(extractor, extracted_text, actual_pages)` for downstream reporting (lines 128-129).

## Practical Implementation Examples

### Command-Line Verification

The most common usage pattern involves running the script directly from the command line to verify a CV before submission:

```bash
python tools/verify_pdf.py cv/main_example.pdf \
  --pages 2 \
  --min-chars 100 \
  --contains "Professional Experience" \
  --dump-text cv/main_example.txt

```

This command attempts `pypdf` extraction first, falls back to `pdftotext` if needed, validates that the PDF contains exactly 2 pages with at least 100 characters, checks for the required phrase, and saves the extracted text layer to [`cv/main_example.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/cv/main_example.txt) for manual review.

### Python API Integration

For programmatic workflows, import the verification functions directly:

```python
from tools.verify_pdf import verify_pdf, VerificationError

pdf_path = "cv/main_example.pdf"
try:
    extractor, text, pages = verify_pdf(
        pdf_path,
        expected_pages=2,
        min_chars=100,
        required_text=("Professional Experience",),
        dump_text="cv/main_example.txt"
    )
    print(f"✅ {pdf_path} verified (extractor={extractor}, pages={pages})")
except VerificationError as exc:
    print(f"❌ Verification failed: {exc}")

```

This pattern allows integration into automated application workflows, with `VerificationError` providing specific failure reasons including missing pages, insufficient text, absent keywords, or missing extraction tools.

### Extracting Raw Text Only

When you need only the text layer without validation—for custom keyword scoring or analysis—use the extraction function directly:

```python
from tools.verify_pdf import extract_text_layer

text, pages, extractor = extract_text_layer("cv/main_example.pdf")
print(f"Extractor: {extractor}, Pages: {pages}")
print("First 200 characters of the ATS text layer:")
print(text[:200])

```

This approach is useful for integrating with custom ATS compatibility scoring systems beyond the built-in validation rules.

## Summary

- **Dual-engine extraction**: The `ai-job-search` tool uses `pypdf` as the primary extractor with `pdftotext` (Poppler) as a fallback to ensure reliable text layer extraction across all PDF types
- **Complete validation**: The `verify_pdf` function checks file existence, page count, minimum character requirements, and required keyword presence in the ATS-readable text
- **Debugging support**: The `--dump-text` flag and `dump_text` parameter allow inspection of the exact text string that Applicant Tracking Systems will parse
- **Graceful degradation**: If both extractors are unavailable, the `run_tool` function provides OS-specific installation hints rather than crashing the workflow

## Frequently Asked Questions

### What is the difference between the pypdf and pdftotext extraction methods?

**`pypdf`** is a pure-Python library that requires no external system dependencies, making it ideal for cross-platform usage and primary extraction attempts. **`pdftotext`** (from the Poppler utilities) is an external command-line tool that often handles complex layouts and embedded fonts more effectively, serving as a robust fallback when `pypdf` returns empty or garbled text. According to the source code in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), the script attempts `pypdf` first via `_extract_pypdf` before falling back to `_extract_pdftotext` only when necessary.

### How does the tool handle cases where both extraction methods fail?

If both `pypdf` and `pdftotext` are unavailable or return empty results, the `run_tool` function raises a `VerificationError` with operating-system-specific installation instructions. This allows the calling workflow—such as the `apply` step in the job application assistant—to gracefully downgrade to a visual keyword review instead of failing completely, ensuring you can still proceed with manual verification even when automated parsing is impossible.

### Why is extracting the text layer specifically important for ATS verification?

Applicant Tracking Systems do not "see" the visual rendering of your CV; they read the **embedded text layer** contained within the PDF file structure. By extracting this specific layer using the same methods ATS parsers employ, the `ai-job-search` tool reveals exactly what keywords, formatting, and content the automated system will actually process. This prevents situations where visually present content is invisible to the ATS due to image-based PDFs, corrupted text encoding, or font subsetting issues.

### Can I verify that specific keywords exist in my CV text layer?

Yes. The `verify_pdf` function accepts a `required_text` parameter (or `--contains` flags via CLI) that validates the presence of specific strings after text normalization. For example, you can require that "Python," "Data Analysis," and "Professional Experience" all appear in the extracted text layer. If any required text is missing, the function raises a `VerificationError`, alerting you that the ATS may not recognize those critical qualifications even if they appear visually in your document.