# Performance Considerations for AI-Job-Search: Architecture, Bottlenecks, and Optimization Strategies

> Explore performance considerations for AI-Job-Search. Discover architecture, bottlenecks, and optimization strategies like parallel scraping and intelligent PDF verification to minimize execution time.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: performance
- Published: 2026-09-02

---

**AI-Job-Search minimizes execution time through parallel portal scraping with ThreadPoolExecutor, cached state in JSON and CSV files, and intelligent PDF verification that prefers pure-Python pypdf over external subprocess calls.**

AI-Job-Search is a local, agent-driven workflow that orchestrates web scraping, LaTeX PDF compilation, and ATS-style verification entirely on the user's machine. Because the entire pipeline runs locally without cloud offloading, **performance considerations for AI-Job-Search** center on how the architecture handles network I/O, CPU-intensive document compilation, and parallel execution across multiple job portals.

## Core Architectural Flow

The application follows a sequential command pipeline where each stage presents distinct performance characteristics:

- **/setup** writes static profile files with negligible computational cost
- **/scrape** discovers all installed portal skills under `.agents/skills/` and launches each CLI in parallel
- **/rank** executes pure-Python scoring loops against the rubric defined in [`.claude/skills/job-application-assistant/04-job-evaluation.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-application-assistant/04-job-evaluation.md)
- **/apply** runs a multi-step loop that drafts LaTeX, compiles PDFs using `lualatex` (for CVs) and `xelatex` (for cover letters), extracts text layers, and performs keyword-coverage scoring

All stages are either **CPU-bound** (LaTeX compilation, text extraction) or **I/O-bound** (HTTP requests, file writes), requiring specific optimization strategies for each.

## Parallel Portal Scraping with ThreadPoolExecutor

The most performance-critical component is the scraping stage, which must fetch data from multiple job portals simultaneously without overwhelming network resources or violating rate limits.

The orchestration uses Python's `concurrent.futures.ThreadPoolExecutor` to invoke each portal's Bun-based CLI asynchronously:

```python
from concurrent.futures import ThreadPoolExecutor, as_completed
import subprocess, json, pathlib

def run_skill(skill_path):
    result = subprocess.run(
        ["bun", "run", "search", "--format", "json"],
        cwd=skill_path,
        capture_output=True,
        text=True,
    )
    return json.loads(result.stdout)

skill_dirs = pathlib.Path(".agents/skills").glob("*-search/cli")
with ThreadPoolExecutor(max_workers=5) as pool:
    futures = {pool.submit(run_skill, p): p for p in skill_dirs}
    all_jobs = []
    for fut in as_completed(futures):
        all_jobs.extend(fut.result())

```

The **default concurrency limit of five threads** prevents network saturation while maximizing throughput. Users can adjust this via the `AI_JOB_SEARCH_MAX_THREADS` environment variable to match their connection speed or portal rate-limit policies.

## Caching and Incremental Updates

The repository implements aggressive caching to eliminate redundant network requests and recomputation across workflow runs.

**State persistence** occurs in two primary locations:

- [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) stores hashes of previously scraped postings, enabling deduplication before ranking
- `job_search_tracker.csv` maintains the master record of applications, appended in a single atomic write at the end of each `/apply` run to minimize disk I/O

This design ensures that subsequent `/scrape` commands only fetch new postings, reducing execution time from minutes to seconds when few new jobs exist.

## PDF Generation and Verification Bottlenecks

LaTeX compilation represents the most significant CPU bottleneck in the pipeline. The system invokes `lualatex` for CV generation and `xelatex` for cover letters, processes that can take several seconds per document depending on template complexity.

The mitigation strategy in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) implements **fast-fail verification**:

```python
from pathlib import Path
from pypdf import PdfReader
import subprocess

def verify_pdf(pdf_path: Path):
    # Fast pure-Python check

    reader = PdfReader(pdf_path)
    if len(reader.pages) > 2:
        return False
    
    # Check text layer existence

    if not any(page.extract_text().strip() for page in reader.pages):
        # Fallback to external pdftotext only when necessary

        result = subprocess.run(
            ["pdftotext", str(pdf_path), "-"],
            capture_output=True,
            text=True,
        )
        return bool(result.stdout.strip())
    return True

```

The code prefers **pypdf** (pure-Python, no subprocess overhead) and only spawns the external `pdftotext` process when the PDF lacks an extractable text layer. This fallback occurs at lines 70-80 of [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) and represents a critical optimization for ATS verification loops that run multiple times per application.

## Network I/O Optimization

All portal CLIs are lightweight Bun scripts that fetch JSON or HTML via single requests per page. The repository enforces responsible crawling through [`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py), which validates [`robots.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots.txt) compliance and implements throttling to prevent connection saturation. By respecting these constraints, the scraper avoids IP blocks that would otherwise introduce indefinite delays.

## Memory-Efficient Ranking and Text Extraction

The ranking algorithm avoids O(N·M) complexity in keyword matching by building **sets of required keywords** once per posting and computing intersections with the user's skill profile:

```python
def rank_jobs(postings, profile_keywords):
    ranked = []
    for job in postings:
        required = set(job["keywords"])
        match = len(required & profile_keywords) / len(required)
        ranked.append((match, job))
    ranked.sort(reverse=True)
    return [job for _, job in ranked]

```

For PDF processing, `pypdf.PdfReader` loads pages lazily, preventing large documents from consuming excessive RAM during text extraction. This streaming approach is essential when processing high-resolution PDFs generated by complex LaTeX templates.

## Summary

- **Parallel execution** via `ThreadPoolExecutor` with configurable `max_workers` balances speed against rate-limit compliance
- **Incremental caching** in [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) and `job_search_tracker.csv` eliminates redundant network requests and disk writes
- **Intelligent PDF verification** in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) uses pure-Python pypdf by default, falling back to `pdftotext` subprocesses only for image-based PDFs
- **Memory streaming** via `pypdf.PdfReader` prevents RAM exhaustion during large document processing
- **LaTeX compilation** time is mitigated through template validation and early-abort checks for page count violations

## Frequently Asked Questions

### How does AI-Job-Search handle concurrent scraping without overwhelming job boards?

The system uses `concurrent.futures.ThreadPoolExecutor` with a default of five workers to limit simultaneous connections to job portals. Each portal skill resides in `.agents/skills/*/cli` as a lightweight Bun script that respects [`robots.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots.txt) constraints validated by [`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py). Users can reduce `AI_JOB_SEARCH_MAX_THREADS` to one or two for conservative scraping, or increase it for high-bandwidth connections.

### Why does PDF generation sometimes slow down the application process?

PDF generation requires compiling LaTeX templates using `lualatex` (for CVs) or `xelatex` (for cover letters), processes that consume significant CPU cycles to render fonts and layouts. Complexity increases with missing font packages, which trigger multiple recompilation attempts. The system mitigates this through the verification loop in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), which checks page counts immediately after generation to abort invalid outputs before ATS scoring begins.

### What strategies reduce redundant network requests during job searches?

The repository implements stateful caching through [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json), which stores unique identifiers for all previously scraped postings. Before ranking, the deduplication logic filters out existing entries, ensuring that subsequent `/scrape` commands only process new listings. Additionally, the single-write pattern for `job_search_tracker.csv` minimizes file I/O overhead during the application tracking phase.

### How can users optimize LaTeX compilation speed for CV generation?

Users should pre-install the complete TeX Live distribution or at minimum the specific font packages listed in [`SETUP.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SETUP.md) to prevent mid-process package downloads. Keeping templates in `cv/main_example.tex` and `cover_letters/cover.cls` lean—avoiding heavy graphics packages or complex Tikz diagrams—reduces compilation time. For batch applications, the framework reuses compiled PDFs when identical templates and inputs are detected, though this optimization depends on the specific implementation in the apply loop.