# How the AI Job Search Framework Handles Unpredictable LaTeX Page Breaks During PDF Compilation

> Discover how the AI Job Search Framework tackles unpredictable LaTeX page breaks. Learn about its compile-and-inspect loop for automatic detection and fixing of pagination failures.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-08-31

---

**The framework treats LaTeX page breaks as a runtime-validation problem rather than a static-template issue, using a compile-and-inspect loop that automatically detects, reports, and fixes pagination failures before finalizing the document.**

Managing LaTeX page breaks in automated document generation poses unique challenges when strict page limits must be enforced. The AI Job Search Framework, hosted at `MadsLorentzen/ai-job-search`, solves this through a deterministic compile-and-inspect workflow that converts unpredictable LaTeX pagination into a test-driven process. By hard-coding expected page counts and programmatically inserting spacing commands, the system guarantees that every generated CV and cover letter meets precise formatting requirements.

## The Compile-and-Inspect Loop Architecture

The framework's core strategy relies on a **compile-and-inspect loop** defined in [`.claude/skills/job-application-assistant/05-cv-templates.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-application-assistant/05-cv-templates.md) and [`06-cover-letter-templates.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/06-cover-letter-templates.md). This loop treats pagination as a runtime constraint that must be validated after each compilation attempt.

### Template-Specific Compilation Commands

When the `/apply` command executes, it invokes template-specific LaTeX engines based on document type. CVs compile with `lualatex`, while cover letters use `xelatex`, as specified in the active template manifest. Both commands execute with `-interaction=nonstopmode` to prevent interactive stops that would halt automated processing.

### Page-Count Verification with verify_pdf.py

After compilation, [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) processes the generated PDF using `pdfinfo` via Poppler to extract the exact page count. The expected limits are hard-coded in template guidelines: **2 pages for CVs** and **1 page for cover letters**. If the count deviates, the loop flags a **page-break violation** and triggers automatic remediation.

## Automatic Layout Fixes for Pagination Violations

When the verification script detects overflow or orphaned content, the framework applies surgical LaTeX modifications rather than regenerating the entire document.

### Preventing Orphaned Entries with \needspace

The primary fix involves inserting `\needspace{5\baselineskip}` immediately before offending `\cventry` commands. This forces LaTeX to reserve sufficient vertical space, keeping entry headers with their associated bullet lists on the same page. The framework automatically adds the `needspace` package to the preamble if absent.

```latex
\needspace{5\baselineskip}
\item{\cventry{2020‑2022}{Data Scientist}{Acme Corp}{Copenhagen}{}{%
  \begin{itemize}
    \item Developed a real‑time recommendation engine …
    \item Reduced latency by 30 % …
  \end{itemize}}}

```

```latex
\usepackage{needspace}

```

### Fallback Adjustments with \enlargethispage

If `\needspace` cannot resolve severe overflow—such as an entire section spilling onto a third page—the framework applies `\enlargethispage{2-3\baselineskip}` to stretch the preceding page's text block. As a last resort, it trims low-relevance content to meet the hard limit.

## Visual Inspection and the Read Tool

Beyond quantitative checks, the framework employs a built-in **Read** tool to programmatically examine the PDF layout. This step detects visual orphans that page counts might miss, such as headings separated from their content or bullet-list overflow at page boundaries.

## ATS-Readability Validation Beyond Visual Layout

Even visually correct PDFs may contain extraction issues that fail Applicant Tracking Systems (ATS). The verification script extracts the text layer—preferring `pypdf` with fallback to `pdftotext`—to validate that contact information and keywords are present, date fields use plain ASCII hyphens, and no garbled glyphs appear. This ensures LaTeX page-break fixes do not compromise machine readability.

## Template-Specific LaTeX Safety Rules

The compile-and-inspect loop enforces syntax rules documented in the template guides to prevent compilation failures:

- **Unescaped `%` characters** are automatically converted to `\%` to prevent comment truncation that silently drops content.
- **`&` symbols** inside `\cventry` are escaped to `\&` to avoid alignment-tab errors.
- **Bullets beginning with `[`** are wrapped in braces as `\item {[text]}` to prevent LaTeX from interpreting the bracketed text as optional labels.

## Implementation in the /apply Command

All pagination handling is orchestrated by the `/apply` command logic in [`.claude/commands/apply.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/commands/apply.md). The following pseudo-code illustrates the iterative validation cycle:

```python
while True:
    run_compile_command()               # lualatex / xelatex

    pages = get_page_count(pdf_path)
    if pages != EXPECTED_PAGES:
        insert_needspace_before_orphans()
        continue                        # re‑compile

    if has_orphaned_entries(pdf_path):
        insert_needspace_before_orphans()
        continue                        # re‑compile

    break                               # layout is clean

```

This deterministic loop continues until the document satisfies both page-count constraints and visual integrity requirements.

## Summary

- The framework handles LaTeX page breaks through a **compile-and-inspect loop** that validates output after each compilation.
- **[`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py)** uses `pdfinfo` to enforce hard page limits: 2 pages for CVs and 1 page for cover letters.
- **`\needspace{5\baselineskip}`** is automatically inserted before orphaned `\cventry` blocks to keep content together.
- The **Read tool** provides visual inspection capabilities to catch layout issues that page counts miss.
- **ATS-readability checks** ensure pagination fixes do not break text extraction for applicant tracking systems.
- Template-specific escaping rules prevent common LaTeX syntax errors that could cause compilation failures.

## Frequently Asked Questions

### How does the framework decide where to insert \needspace commands?

The framework identifies orphaned entries by analyzing the PDF structure after compilation. When `has_orphaned_entries()` detects a `\cventry` header separated from its bullet list across pages, it calculates the appropriate `\needspace` length (typically `5\baselineskip`) and inserts the command immediately before the offending entry in the source file.

### What happens if \enlargethispage cannot fix a page overflow?

If stretching the page by `2-3\baselineskip` still results in excess pages, the framework enters a content-trimming phase. It removes low-relevance entries—such as older experience or optional sections—then re-compiles. This fallback ensures the document strictly adheres to the mandated page count defined in the template guidelines.

### Why does the framework check for unescaped % and & characters?

Unescaped percent signs initiate LaTeX comments, causing the compiler to silently ignore all subsequent text on that line. Ampersands inside `\cventry` commands trigger alignment-tab errors because `&` is reserved for table environments. The framework escapes these characters automatically to prevent compilation failures that would interrupt the automated workflow.

### How does verify_pdf.py ensure ATS compatibility?

The script extracts the PDF's text layer using `pypdf` (falling back to `pdftotext`) and validates that required contact information and keywords are present. It specifically checks that date fields use plain ASCII hyphens rather than unicode dashes, which many applicant tracking systems fail to parse correctly, ensuring the document remains machine-readable despite layout adjustments.