Performance Considerations for AI-Job-Search: Architecture, Bottlenecks, and Optimization Strategies

AI-Job-Search minimizes execution time through parallel portal scraping with ThreadPoolExecutor, cached state in JSON and CSV files, and intelligent PDF verification that prefers pure-Python pypdf over external subprocess calls.

AI-Job-Search is a local, agent-driven workflow that orchestrates web scraping, LaTeX PDF compilation, and ATS-style verification entirely on the user's machine. Because the entire pipeline runs locally without cloud offloading, performance considerations for AI-Job-Search center on how the architecture handles network I/O, CPU-intensive document compilation, and parallel execution across multiple job portals.

Core Architectural Flow

The application follows a sequential command pipeline where each stage presents distinct performance characteristics:

  • /setup writes static profile files with negligible computational cost
  • /scrape discovers all installed portal skills under .agents/skills/ and launches each CLI in parallel
  • /rank executes pure-Python scoring loops against the rubric defined in .claude/skills/job-application-assistant/04-job-evaluation.md
  • /apply runs a multi-step loop that drafts LaTeX, compiles PDFs using lualatex (for CVs) and xelatex (for cover letters), extracts text layers, and performs keyword-coverage scoring

All stages are either CPU-bound (LaTeX compilation, text extraction) or I/O-bound (HTTP requests, file writes), requiring specific optimization strategies for each.

Parallel Portal Scraping with ThreadPoolExecutor

The most performance-critical component is the scraping stage, which must fetch data from multiple job portals simultaneously without overwhelming network resources or violating rate limits.

The orchestration uses Python's concurrent.futures.ThreadPoolExecutor to invoke each portal's Bun-based CLI asynchronously:

from concurrent.futures import ThreadPoolExecutor, as_completed
import subprocess, json, pathlib

def run_skill(skill_path):
    result = subprocess.run(
        ["bun", "run", "search", "--format", "json"],
        cwd=skill_path,
        capture_output=True,
        text=True,
    )
    return json.loads(result.stdout)

skill_dirs = pathlib.Path(".agents/skills").glob("*-search/cli")
with ThreadPoolExecutor(max_workers=5) as pool:
    futures = {pool.submit(run_skill, p): p for p in skill_dirs}
    all_jobs = []
    for fut in as_completed(futures):
        all_jobs.extend(fut.result())

The default concurrency limit of five threads prevents network saturation while maximizing throughput. Users can adjust this via the AI_JOB_SEARCH_MAX_THREADS environment variable to match their connection speed or portal rate-limit policies.

Caching and Incremental Updates

The repository implements aggressive caching to eliminate redundant network requests and recomputation across workflow runs.

State persistence occurs in two primary locations:

  • job_scraper/seen_jobs.json stores hashes of previously scraped postings, enabling deduplication before ranking
  • job_search_tracker.csv maintains the master record of applications, appended in a single atomic write at the end of each /apply run to minimize disk I/O

This design ensures that subsequent /scrape commands only fetch new postings, reducing execution time from minutes to seconds when few new jobs exist.

PDF Generation and Verification Bottlenecks

LaTeX compilation represents the most significant CPU bottleneck in the pipeline. The system invokes lualatex for CV generation and xelatex for cover letters, processes that can take several seconds per document depending on template complexity.

The mitigation strategy in tools/verify_pdf.py implements fast-fail verification:

from pathlib import Path
from pypdf import PdfReader
import subprocess

def verify_pdf(pdf_path: Path):
    # Fast pure-Python check

    reader = PdfReader(pdf_path)
    if len(reader.pages) > 2:
        return False
    
    # Check text layer existence

    if not any(page.extract_text().strip() for page in reader.pages):
        # Fallback to external pdftotext only when necessary

        result = subprocess.run(
            ["pdftotext", str(pdf_path), "-"],
            capture_output=True,
            text=True,
        )
        return bool(result.stdout.strip())
    return True

The code prefers pypdf (pure-Python, no subprocess overhead) and only spawns the external pdftotext process when the PDF lacks an extractable text layer. This fallback occurs at lines 70-80 of tools/verify_pdf.py and represents a critical optimization for ATS verification loops that run multiple times per application.

Network I/O Optimization

All portal CLIs are lightweight Bun scripts that fetch JSON or HTML via single requests per page. The repository enforces responsible crawling through tools/robots_check.py, which validates robots.txt compliance and implements throttling to prevent connection saturation. By respecting these constraints, the scraper avoids IP blocks that would otherwise introduce indefinite delays.

Memory-Efficient Ranking and Text Extraction

The ranking algorithm avoids O(N·M) complexity in keyword matching by building sets of required keywords once per posting and computing intersections with the user's skill profile:

def rank_jobs(postings, profile_keywords):
    ranked = []
    for job in postings:
        required = set(job["keywords"])
        match = len(required & profile_keywords) / len(required)
        ranked.append((match, job))
    ranked.sort(reverse=True)
    return [job for _, job in ranked]

For PDF processing, pypdf.PdfReader loads pages lazily, preventing large documents from consuming excessive RAM during text extraction. This streaming approach is essential when processing high-resolution PDFs generated by complex LaTeX templates.

Summary

  • Parallel execution via ThreadPoolExecutor with configurable max_workers balances speed against rate-limit compliance
  • Incremental caching in job_scraper/seen_jobs.json and job_search_tracker.csv eliminates redundant network requests and disk writes
  • Intelligent PDF verification in tools/verify_pdf.py uses pure-Python pypdf by default, falling back to pdftotext subprocesses only for image-based PDFs
  • Memory streaming via pypdf.PdfReader prevents RAM exhaustion during large document processing
  • LaTeX compilation time is mitigated through template validation and early-abort checks for page count violations

Frequently Asked Questions

How does AI-Job-Search handle concurrent scraping without overwhelming job boards?

The system uses concurrent.futures.ThreadPoolExecutor with a default of five workers to limit simultaneous connections to job portals. Each portal skill resides in .agents/skills/*/cli as a lightweight Bun script that respects robots.txt constraints validated by tools/robots_check.py. Users can reduce AI_JOB_SEARCH_MAX_THREADS to one or two for conservative scraping, or increase it for high-bandwidth connections.

Why does PDF generation sometimes slow down the application process?

PDF generation requires compiling LaTeX templates using lualatex (for CVs) or xelatex (for cover letters), processes that consume significant CPU cycles to render fonts and layouts. Complexity increases with missing font packages, which trigger multiple recompilation attempts. The system mitigates this through the verification loop in tools/verify_pdf.py, which checks page counts immediately after generation to abort invalid outputs before ATS scoring begins.

What strategies reduce redundant network requests during job searches?

The repository implements stateful caching through job_scraper/seen_jobs.json, which stores unique identifiers for all previously scraped postings. Before ranking, the deduplication logic filters out existing entries, ensuring that subsequent /scrape commands only process new listings. Additionally, the single-write pattern for job_search_tracker.csv minimizes file I/O overhead during the application tracking phase.

How can users optimize LaTeX compilation speed for CV generation?

Users should pre-install the complete TeX Live distribution or at minimum the specific font packages listed in SETUP.md to prevent mid-process package downloads. Keeping templates in cv/main_example.tex and cover_letters/cover.cls lean—avoiding heavy graphics packages or complex Tikz diagrams—reduces compilation time. For batch applications, the framework reuses compiled PDFs when identical templates and inputs are detected, though this optimization depends on the specific implementation in the apply loop.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →