# How the PDF Verification Loop Works in the AI Job Search /apply Step

> Understand the PDF verification loop in the AI job search /apply step. Learn how it validates CVs by checking page count, character minimums, and required text before application submission.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-08-29

---

**The PDF verification loop in the `/apply` step validates CV documents through an iterative check system that enforces page count, character minimums, and required text phrases before allowing the job application to proceed.**

The `/apply` command in the MadsLorentzen/ai-job-search repository automates job applications through a Markdown-driven pipeline. Before recording any application, the system runs a mandatory PDF verification loop to ensure CV integrity. This loop, defined in the command specification and implemented in Python validation scripts, prevents corrupted or incomplete documents from entering the application tracker.

## Where the PDF Verification Loop Is Defined

The verification loop spans three critical files in the repository architecture:

- **[`.claude/commands/apply.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/commands/apply.md)** – The Markdown command specification that defines the `/apply` workflow and orchestrates the verification step
- **[`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py)** – The Python implementation containing the `verify_pdf` function that performs document extraction and validation
- **[`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py)** – The security configuration that registers the PDF verifier as a mandatory guard for the application process

## How the PDF Verification Loop Validates Documents

The loop operates as a **repeat-until-valid** control structure that extracts PDF content and validates it against configurable constraints.

### Text Extraction and Parsing

When invoked, the `verify_pdf` function in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) first extracts raw text from the CV using the `_extract_pypdf` helper:

```python
extractor, text, pages = _extract_pypdf(pdf_path)

```

This extraction uses PyPDF2 or a fallback parser to return the **extractor object**, **concatenated text content**, and **total page count** from the PDF file.

### Validation Guards and Constraints

The verification loop applies three primary validation guards against the extracted content:

1. **Page Count Validation** – If `expected_pages` is specified, the document must contain at least that many pages
2. **Character Minimum Check** – The `min_chars` parameter (defaulting to 1) ensures the PDF is not empty or trivially short
3. **Required Text Verification** – The `required_text` tuple specifies mandatory phrases (e.g., "Professional Experience") that must appear in the extracted text

Each failed guard appends a descriptive error message to an accumulator list.

### Error Handling and Retry Logic

If any validation fails, the function streams error messages to `stderr` and returns `None`:

```python
if errors:
    for msg in errors:
        print(msg, file=sys.stderr)
    return None
return extractor, text, pages

```

The calling loop in [`apply.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/apply.md) detects this failure and prompts the user to correct the PDF before retrying, creating the iterative verification cycle.

## Core Implementation: The verify_pdf Function

The validation logic resides in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) and follows this signature:

```python
def verify_pdf(pdf_path,
               expected_pages: int | None = None,
               min_chars: int = 1,
               required_text: tuple[str, ...] = (),
               dump_text: Path | None = None):
    extractor, text, pages = _extract_pypdf(pdf_path)

    errors = []
    if expected_pages and pages < expected_pages:
        errors.append(f"expected ≥{expected_pages} pages, got {pages}")

    if len(text) < min_chars:
        errors.append(f"expected ≥{min_chars} characters, got {len(text)}")

    for phrase in required_text:
        if phrase not in text:
            errors.append(f"missing required phrase: {phrase!r}")

    if errors:
        for msg in errors:
            print(msg, file=sys.stderr)
        return None
    return extractor, text, pages

```

This function acts as the gatekeeper: returning the extraction data only when all constraints are satisfied, or blocking progression by returning `None` and logging specific failure reasons.

## Security Guard Integration

The verification loop is enforced through the security guards system defined in [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py). The PDF verifier is registered as a mandatory pre-condition:

```python
SECURITY_GUARDS = [
    "Bash(python tools/verify_pdf.py:*)",
    # … other guards …

]

```

This registration ensures that [`verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/verify_pdf.py) executes automatically before the `/apply` step writes any tracker entries, guaranteeing that every recorded application references a validated CV document.

## The Loop Control Structure in apply.md

The command specification implements the iterative loop using bash control flow within the Markdown command definition:

```bash
while true; do
    result=$(python tools/verify_pdf.py "$CV_PATH" \
              --expected-pages 2 \
              --min-chars 500 \
              --required-text "Professional Experience")
    if [ $? -eq 0 ]; then
        break
    fi
    echo "PDF verification failed – see above. Please fix the file and retry."
    read -p "Press <Enter> after fixing the PDF..."
done

```

This structure creates a blocking loop that pauses the application process until the PDF passes all validation checks, providing immediate feedback to the user about document quality issues.

## Summary

- The **PDF verification loop** is a mandatory gate in the `/apply` workflow that validates CV documents before application recording
- Validation occurs in **[`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py)**, where the `verify_pdf` function checks page counts, character minimums, and required text phrases
- The loop follows a **repeat-until-valid** pattern defined in **[`.claude/commands/apply.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/commands/apply.md)**, prompting users to fix documents that fail validation
- **Security guards** in [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) enforce this verification as a prerequisite for the application process
- Failed validations return `None` and print diagnostic errors to `stderr`, while successful validations return extraction data and allow workflow progression

## Frequently Asked Questions

### What happens if my PDF fails the verification loop?

If the PDF fails validation, the `verify_pdf` function prints specific error messages to `stderr` indicating which constraints failed—such as insufficient pages, low character counts, or missing required phrases. The `/apply` command then traps these errors in a `while` loop, displays them to the user, and pauses execution until the user confirms they have fixed the document and presses Enter to retry.

### Which validation parameters can I configure in the PDF verification loop?

According to the `verify_pdf` function signature in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), you can configure three main parameters. These include `expected_pages` for minimum page count, `min_chars` for minimum character count (defaulting to 1), and `required_text` for mandatory phrases. You pass these values as command-line arguments when the verification script is invoked within the `/apply` step.

### Where does the PDF verification loop fit in the overall /apply workflow?

The loop executes during **Step 0: Parse Input** of the `/apply` command, immediately after the CV path is identified. Because [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) registers the verifier as a mandatory guard, this validation occurs before any application data is written to the tracker, ensuring data integrity at the entry point of the workflow.

### How does the system extract text from PDFs for validation?

The `verify_pdf` function relies on an internal `_extract_pypdf` helper that uses **PyPDF2** (with fallback mechanisms) to parse the PDF binary. This extractor returns the raw text content, page count, and an extractor object, which the verification loop then analyzes against the configured constraints.