How career-ops Generates PDF Resumes and Reports: The Complete Technical Guide

career-ops transforms Markdown CVs and evaluation reports into print-ready PDFs through a tightly coupled three-stage pipeline that combines HTML template assembly, Playwright Chromium rendering, and automated manifest indexing to enforce page budgets and workspace security.

The santifer/career-ops repository automates the creation of professional job application documents by converting Markdown source files into ATS-compatible PDFs. This open-source toolchain generates both resumes and evaluation reports using a unified HTML-to-PDF engine that prioritizes typographic safety and strict page-count enforcement. Understanding how career-ops generates PDF resumes and reports reveals a sophisticated workflow that balances flexibility with recruiter-facing constraints.

The Three-Stage PDF Generation Architecture

The system processes every document through three distinct phases, each handled by specific ECMAScript modules in the repository root.

Stage 1: HTML Assembly and ATS Sanitization

In build-cv-html.mjs, the pipeline begins by reading cv.md (or report markdown) and merging it with templates/cv-template.html. The module calls injectThemeStyle and injectPrintPageCss to embed CSS defining @page { size: …; margin: … } rules. It then executes normalizeTextForATS(), which replaces smart quotes, em-dashes, and zero-width characters with ASCII-safe equivalents to prevent ATS parser failures. Finally, validateCvSectionOrder() ensures the rendered HTML respects the section sequence defined in config/profile.yml, throwing an error unless --allow-reorder is specified.

Stage 2: Playwright Chromium Rendering

The generate-pdf.mjs module handles conversion using Playwright's headless Chromium browser. After asserting workspace safety via assertInsideWorkspace() and isWorkspaceOutputPath(), the script loads the assembled HTML and invokes page.pdf({path: outputPath, format, margin…}). Post-generation, countRenderedPdfPages() extracts the /Count value from the PDF's /Pages dictionary to determine actual page length. The enforcePageBudget() function compares this against the --max-pages argument (default 2), emitting warnings or hard errors (with --strict-pages) when content exceeds limits.

Stage 3: Manifest Indexing

Each successful generation triggers updatePDFManifest(reportNum, pdfPath, htmlPath, format), which appends a tab-separated record to data/pdf-index.tsv. The manifest stores tracker/report numbers, relative paths, page format, and generation timestamps, enabling the dashboard TUI to locate specific PDFs instantly. If a reportNum is provided, previous entries for that report are purged to maintain index accuracy.

Template Injection and Print Styling

The theme-style.mjs module manages design tokens, feeding build-cv-html.mjs with CSS variables that control typography and layout. During HTML assembly, the system injects a page-size style block directly into the document head, ensuring the subsequent PDF rendering respects physical constraints like A4 or Letter dimensions. This tight coupling between the HTML builder and the PDF generator guarantees that @media print rules render identically in both the browser preview and the final output.

Workspace Security and Path Validation

Before any file operations, generate-pdf.mjs validates that both input HTML and output PDF paths reside within the tracker workspace. The assertInsideWorkspace() and isWorkspaceOutputPath() functions prevent path-traversal attacks by rejecting absolute paths or directory escapes. This security model ensures that CLI invocations cannot overwrite system files or read sensitive data outside the designated directory structure.

LaTeX Alternative Pipeline

For users preferring typeset quality, generate-latex.mjs provides a parallel conversion path. This module transforms cv.md into LaTeX source using templates/cv-template.tex, then compiles to PDF via the system pdflatex binary. The validateLatexContent() function performs content filtering specific to LaTeX syntax, while the module ultimately calls the same updatePDFManifest logic to maintain consistency with the HTML-based workflow.

Command-Line Usage Examples

Generate a CV PDF with strict page limits:

career-ops pdf \
   --input output/cv.html \
   --output output/cv.pdf \
   --format a4 \
   --max-pages 2 \
   --strict-pages \
   --allow-reorder

Link a report PDF to tracker entry #042:

career-ops pdf \
   --input output/report-042.html \
   --output output/report-042.pdf \
   --report 042 \
   --format letter

Build HTML only (preprocessing step):

career-ops build-cv-html

Summary

  • Three-stage pipeline: HTML assembly → Playwright rendering → Manifest indexing
  • build-cv-html.mjs handles template merging and ATS-safe text normalization via normalizeTextForATS()
  • generate-pdf.mjs uses Playwright's page.pdf() with strict workspace validation through assertInsideWorkspace()
  • Page budget enforcement via --max-pages and --strict-pages flags prevents exceeding recruiter length expectations
  • Automatic manifest updates in data/pdf-index.tsv enable dashboard integration and report lookup
  • LaTeX alternative available through generate-latex.mjs using system pdflatex for typographic precision

Frequently Asked Questions

How does career-ops prevent ATS parsing errors in generated PDFs?

The system runs normalizeTextForATS() during HTML assembly to replace typographic characters like smart quotes and em-dashes with ASCII equivalents. This sanitization occurs in build-cv-html.mjs before PDF conversion, ensuring downstream applicant tracking systems receive clean, parseable text without zero-width characters or Unicode punctuation that typically cause parsing failures.

What limits the page count of generated resumes?

The enforcePageBudget() function in generate-pdf.mjs reads the actual page count from the PDF catalog's /Count dictionary after Playwright renders the document. By default, the system allows a --max-pages value of 2, emitting warnings or hard errors (with --strict-pages) when content exceeds this budget, preventing accidentally lengthy application documents that violate recruiter expectations.

How does the system track which PDF belongs to which job application?

Every successful PDF generation updates data/pdf-index.tsv via updatePDFManifest(). This tab-separated file records the report number, relative PDF path, source HTML path, format, and date, enabling the dashboard TUI to retrieve the correct document for any tracked application using the report ID as a unique lookup key.

Can I use LaTeX instead of HTML for PDF generation?

Yes. The generate-latex.mjs module provides an alternative pipeline that converts cv.md to LaTeX using templates/cv-template.tex, then compiles to PDF using the system's pdflatex binary. While the rendering engine differs from the Playwright approach, it performs similar content validation through validateLatexContent() and updates the same manifest index to maintain compatibility with the tracker workflow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →