# ATS-Optimized PDF Generation in Career-Ops: Keyword Injection and Rendering Pipeline

> Career-Ops creates ATS friendly PDFs by injecting keywords and rendering via Chromium. Learn how our pipeline ensures exact keyword matches from Markdown for optimal resume parsing.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Career-Ops generates ATS-friendly PDFs by sanitizing Unicode characters, inlining fonts, and rendering through Chromium while preserving exact keyword matches from the source Markdown.**

The Career-Ops repository (`santifer/career-ops`) provides a robust pipeline for converting CV Markdown files into **ATS-optimized PDF generation** outputs that pass through applicant tracking systems without character corruption or keyword loss. The entire workflow is orchestrated by `generate-pdf.mjs`, which transforms rendered HTML into parser-friendly PDF documents through a five-stage sanitization and rendering process.

## The Five-Stage ATS-Optimized PDF Pipeline

The `generate-pdf.mjs` script implements a deterministic pipeline that ensures visual fidelity while maximizing ATS compatibility. Each stage targets specific failure modes that cause traditional PDFs to fail automated parsing.

### Stage 1: Input Handling and Section Validation

The process begins with **CLI argument parsing** in the `generatePDF` function (lines **44‑63**), which accepts `node generate-pdf.mjs <input.html> <output.pdf>` along with optional flags like `--format=a4` and `--report=018`. 

Before rendering, the system validates document integrity through `validateCvSectionOrder` (lines **58‑78**). This function compares the rendered HTML section order against the original [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md) source to ensure no content drift occurred during template processing. If the sections mismatch, the script exits with a validation error, preventing corrupted PDFs from entering the job application pipeline.

### Stage 2: ATS-Safe Text Sanitization

The core **ATS optimization** occurs in `normalizeTextForATS` (lines **40‑94**), which traverses the HTML body and replaces Unicode characters known to break parser logic. The function specifically targets:

- **Smart quotes** (curly apostrophes) → straight ASCII quotes
- **Em and en dashes** → hyphens
- **Ellipsis characters** → three periods
- **Zero-width spaces and non-breaking spaces** → standard spaces
- **Arrow glyphs, bullet points, and currency symbols** (€, £) → ASCII equivalents or text descriptions

Each replacement is counted and logged for diagnostic purposes, ensuring that keywords adjacent to these symbols remain intact and readable by ATS algorithms.

### Stage 3: Font Inlining for Visual Consistency

To prevent **missing font issues** that can cause text extraction failures, the `inlineLocalFonts` function (lines **30‑55**) processes all `url('./fonts/…')` references in the CSS. It reads the local font files, base64-encodes them, and embeds them as `data:` URLs directly into the HTML. This guarantees that the PDF retains its visual typography even when the document is processed by systems that strip external font references.

### Stage 4: PDF Rendering with Chromium

The sanitized HTML (now containing inlined fonts) is written to a temporary file and loaded in a headless Chromium instance via `renderHtmlToPdf` (lines **73‑115**). The function calls `page.goto(file://…)` and renders to PDF with strict **0.6 inch margins** to maintain tight layout control for ATS scanning zones. This stage supports both `a4` and `letter` format specifications via the `--format` CLI flag.

### Stage 5: Manifest Bookkeeping

After successful rendering, `updatePDFManifest` (lines **95‑134**) appends a tracking entry to `data/pdf-index.tsv`. This tab-separated manifest maps report numbers (from `--report=`) to their corresponding PDF paths, original HTML sources, and chosen formats. The manifest enables downstream tools such as the TUI dashboard and tracker merging systems to locate and reference generated documents programmatically.

## Understanding Keyword Injection in Career-Ops

In the context of Career-Ops, **keyword injection** does not involve hidden text or metadata stuffing. Instead, it refers to the pipeline's preservation strategy that ensures the **exact keywords** supplied in your [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md) appear unchanged in the final PDF output.

The mechanism works by normalizing only the surrounding punctuation and special characters that typically cause ATS parsers to drop or mangle adjacent words. By converting smart quotes, currency symbols, and non-standard dashes into ASCII-safe equivalents, the system prevents the parser from discarding the keywords themselves. The result is a PDF containing the same keyword set as your Markdown source, expressed in a clean, extraction-friendly format without visual degradation.

## Practical Usage and Code Examples

Generate an ATS-ready PDF from your rendered CV template:

```bash
node generate-pdf.mjs output/cv.html output/cv.pdf --format=a4 --report=018

```

The internal execution flow follows this simplified pattern:

```js
// 1️⃣ Load the HTML and optional cv.md source
let html = await readFile('output/cv.html', 'utf-8');
let cvMd = await readFile('cv.md', 'utf-8');

// 2️⃣ Verify section order matches the Markdown source
validateCvSectionOrder(html, cvMd);

// 3️⃣ Apply ATS-safe normalisation (smart quotes → ", em-dash → - …)
const { html: cleanHtml, replacements } = normalizeTextForATS(html);

// 4️⃣ Inline any local fonts so the PDF keeps the visual design
const finalHtml = await inlineLocalFonts(cleanHtml);

// 5️⃣ Render the HTML to a PDF with Chromium
const { outputPath } = await renderHtmlToPdf(finalHtml, 'output/cv.pdf', {
  format: 'a4',
  reportNum: '018',
  inputPath: 'output/cv.html',
});

```

Upon completion, the manifest at `data/pdf-index.tsv` receives a new entry linking the report to the generated file:

```

018	output/pdf/cv.pdf	output/cv.html	a4	2026-07-03

```

## Summary

- **ATS-optimized PDF generation** in Career-Ops occurs through `generate-pdf.mjs`, which implements a five-stage pipeline from input validation to manifest tracking.
- The `normalizeTextForATS` function (lines **40‑94**) sanitizes Unicode characters like smart quotes and em-dashes to prevent keyword corruption during ATS parsing.
- **Font inlining** via `inlineLocalFonts` ensures visual consistency without relying on external font references that might break text extraction.
- The system validates document structure using `validateCvSectionOrder` (lines **58‑78**) to ensure the rendered output matches the source Markdown sequence.
- **Keyword injection** refers to preserving exact keyword matches by normalizing surrounding punctuation rather than adding hidden text, ensuring [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md) content survives the PDF conversion process intact.

## Frequently Asked Questions

### What is ATS-optimized PDF generation?

**ATS-optimized PDF generation** is the process of creating resume PDFs that can be accurately parsed by Applicant Tracking Systems without character encoding errors or text extraction failures. In Career-Ops, this involves sanitizing Unicode characters, inlining fonts, and rendering through Chromium with specific margin constraints to ensure both machine readability and visual polish.

### How does Career-Ops sanitize text for ATS compatibility?

Career-Ops uses the `normalizeTextForATS` function in `generate-pdf.mjs` (lines **40‑94**) to replace problematic Unicode entities—such as smart quotes, em-dashes, ellipsis characters, and currency symbols—with ASCII-safe equivalents. This prevents ATS parsers from dropping keywords that appear next to these special characters while maintaining the semantic content of the CV.

### What does keyword injection mean in Career-Ops?

In Career-Ops, **keyword injection** refers to the preservation of exact keyword matches from the source [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md) file through the PDF generation pipeline. Rather than inserting hidden strings, the system ensures that keywords remain intact by normalizing only the surrounding punctuation and special symbols, preventing ATS parsers from discarding or mangling the terms during text extraction.

### Where are generated PDFs tracked in the repository?

Generated PDFs are tracked in `data/pdf-index.tsv`, which is updated by the `updatePDFManifest` function (lines **95‑134**) in `generate-pdf.mjs`. This tab-separated manifest maps report numbers (specified via `--report=`) to their output paths, source HTML files, and chosen page formats, enabling integration with the repository's TUI dashboard and automated tracking systems.