# AI Job Search Framework: Key Files and Their Purposes Explained

> Explore the AI Job Search Framework's core files like CLAUDE.md and .agents/skills. Understand how these components manage profiles, scraping, and application lifecycle for automated job searching.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-09-03

---

**The AI Job Search Framework organizes job application automation through three architectural layers—Profile & Rules (CLAUDE.md, .claude/), Search & Automation (.agents/skills/, tools/), and Artifacts & State (documents/, job_search_tracker.csv)—that orchestrate candidate profiles, multi-portal scraping, and application lifecycle management.**

The **MadsLorentzen/ai-job-search** repository implements a thin-pointer design where declarative specifications drive CLI-based automation. This open-source framework separates immutable workflow logic from mutable personal data, enabling safe fork updates while maintaining a centralized **job_search_tracker.csv** as the source of truth for all application activities.

## Profile and Configuration Layer

### The Canonical Candidate Profile (CLAUDE.md)

At the root of the repository, **CLAUDE.md** serves as the single source of truth for candidate identity. This markdown file consolidates personal data, education history, work experience, skills matrices, career goals, and evaluation rubrics (structured as [`01-candidate-profile.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/01-candidate-profile.md) through [`07-interview-prep.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/07-interview-prep.md)). During initialization, the `/setup` command reads either this file or the `documents/` folder to construct a structured profile that feeds into subsequent fit evaluation logic.

The profile stored in **CLAUDE.md** directly powers the ranking algorithms in [`04-job-evaluation.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/04-job-evaluation.md), enabling commands like `/rank` and `/apply` to score job postings against explicit candidate criteria and tailor CV content dynamically.

### Claude Code Commands and Skills (.claude/ directory)

The **.claude/commands/** directory contains declarative markdown specifications that Claude Code reads at runtime. These files implement the primary workflow commands:

- `/setup` – Initializes or updates the candidate profile
- `/apply` – Generates tailored application materials
- `/scrape` – Invokes job portal agents
- `/rank` – Batch-scores postings against the profile
- `/outcome` – Records application results
- `/gmail-sync` – Updates tracker from email status

Each command file follows a strict contract defining inputs, outputs, and tool invocations. The **.claude/skills/** directory houses core skill definitions including `job-application-assistant`, `job-scraper`, and `upskill`, with each skill containing a **SKILL.md** descriptor that specifies its contract, enabled flags, and execution parameters.

### Permission Settings (.claude/settings.json)

The **.claude/settings.json** file maintains the permission allowlist for Claude Code execution. This JSON configuration explicitly enumerates which tools may be executed and defines shared scopes, enforcing security boundaries between the AI agent and local system resources.

## Search and Automation Layer

### Job Portal Scrapers (.agents/skills/)

The **.agents/skills/** directory contains modular CLI tools for specific job portals. Each portal skill (e.g., `jobbank-search`, `jobindex-search`, `linkedin-search`, `freehire-search`) follows a standardized contract requiring three components:

1. A `search` command returning JSON listings
2. A `detail` command for single posting retrieval
3. A **SKILL.md** descriptor defining the interface

These scrapers are built on Bun/TypeScript CLI implementations residing in each skill's `cli/` subdirectory. The framework's `/scrape` command iterates through enabled skills, aggregating results into the central tracker while respecting **robots.txt** compliance through enforcement tools.

Installation requires zero runtime dependencies for most portals:

```bash

# Install all bundled portal tools

for tool in jobbank-search jobdanmark-search jobindex-search jobnet-search \
            linkedin-search freehire-search; do
  (cd .agents/skills/$tool/cli && bun install)
done

# Execute a search on Jobindex

.jobbank-search/cli/search --query "data scientist" --location "Copenhagen" --format json

```

### CI and Utility Scripts (tools/)

The **tools/** directory contains Python-based infrastructure scripts that maintain framework integrity:

- **tools/check_framework_version.py** – Verifies that **AGENTS.md** reflects version bumps when skill files change, preventing stale releases
- **tools/check_upstream_updates.py** – Previews which personalized files upstream updates will modify, facilitating safe fork synchronization
- **tools/upstream_triage.py** – Categorizes upstream commits into "review-required" versus "safe-to-skip" for weekly maintenance workflows
- **tools/robots_check.py** – Validates robots.txt compliance before scraper retries
- **tools/lint_skills.py** – CI linter validating **SKILL.md** file shapes, settings syntax, and manifest completeness

### Salary Benchmarking (salary_lookup.py)

The **salary_lookup.py** utility provides compensation analysis by reading user-supplied **salary_data.json** (documented in **tools/README_SALARY_TOOL.md**). When invoked during the ranking phase, this script scores posting compensation against market data, outputting normalized benchmarks that influence the fit scoring algorithm.

## Artifacts and State Management

### Document Storage and Archives (documents/)

The **documents/** directory implements a strict folder hierarchy defined in **documents/README.md**. This hierarchy stores raw career materials (PDF CVs, LinkedIn exports, diplomas, reference letters) and application archives following the naming convention `<company>_<role>`.

When `/apply` generates materials for a specific posting, it creates a subdirectory under **documents/applications/** containing:

- [`posting.html`](https://github.com/MadsLorentzen/ai-job-search/blob/main/posting.html) – Raw job description (optional)
- `cv.pdf` – Compiled CV tailored to the role
- `cover.pdf` – Compiled cover letter
- [`outcome.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/outcome.md) – Status tracking (drafted, interviewed, offer, rejected)

The "Subfolder naming" rule in **documents/README.md** automatically derives these paths from company and role metadata, ensuring consistent archive organization.

### CV and Cover Letter Templates (cv/ and cover_letters/)

Default document generation relies on LaTeX templates:

- **cv/main_example.tex** – Stock **moderncv** template compiled during `/apply`
- **cover_letters/cover.cls** – Custom LaTeX class defining cover letter formatting
- **cover_letters/cover_example.tex** – Default cover letter template

The **templates/** directory supports custom template registration via `/add-template`, accepting LaTeX, Typst, or other compile-to-PDF formats. **templates/README.md** documents the registration protocol for extending output formats beyond the default moderncv implementation.

### Application Tracking (job_search_tracker.csv and Runtime State)

**job_search_tracker.csv** functions as the central state database, recording every scraped posting, computed fit scores, and application status transitions. This CSV is the primary interface for `/rank`, `/outcome`, and `/html-report` commands.

Runtime state persists in specialized directories:

- **job_scraper/** – Contains [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) and cached results enabling incremental scraping and deduplication
- **gmail_sync/** – Stores processed message IDs and last sync timestamps for the `/gmail-sync` command
- **upskill/** – Receives markdown reports from `/upskill` skill-gap analyses

## Framework Workflow Integration

A typical job search session orchestrates these files through Claude Code commands:

```bash

# 1. Initialize or update candidate profile

claude
/setup    # Interactive profile construction from documents/ or CLAUDE.md

# 2. Search across all configured portals

/scrape   # Invokes every skill in .agents/skills/ and populates tracker

# 3. Rank opportunities against profile (optional)

/rank     # Batch-scores new postings, updates job_search_tracker.csv

# 4. Generate application materials

/apply https://example.com/job/1234567

# Compiles cv/main_example.tex and cover_letters/cover_example.tex

# Archives to documents/applications/Company_Role/

# 5. Record interview outcomes

/outcome  # Updates outcome.md and tracker status

# 6. Sync status from Gmail

/gmail-sync   # Processes mailbox for application status changes

```

Additional configuration files include **SETUP.md** (installation prerequisites and TeX setup), **SECURITY.md** (threat model and safe execution guidelines), and **AGENTS.md** (agent runtime selection for Claude Code, Codex, or Gemini CLI).

## Summary

- **CLAUDE.md** stores the canonical candidate profile and evaluation rubrics that drive all personalization logic
- **.claude/commands/** and **.claude/skills/** contain the declarative specifications that Claude Code executes for setup, scraping, ranking, and application workflows
- **.agents/skills/** houses modular, Bun-based job portal scrapers following standardized JSON output contracts
- **tools/** provides CI infrastructure including version checking, upstream triage, robots.txt validation, and skill linting
- **salary_lookup.py** enables compensation benchmarking against user-supplied market data
- **documents/** maintains the archive hierarchy for raw career materials and application-specific outputs
- **cv/** and **cover_letters/** contain the default LaTeX templates compiled by the `/apply` command
- **job_search_tracker.csv** serves as the central state database for all job postings, scores, and application statuses

## Frequently Asked Questions

### What is the purpose of CLAUDE.md in the AI Job Search Framework?

**CLAUDE.md** acts as the master candidate profile containing personal data, work history, skills, goals, and structured evaluation rubrics. During `/setup`, Claude Code either parses this file or imports data from the **documents/** folder to build a structured profile. All subsequent commands—`/rank`, `/apply`, and `/outcome`—reference this profile to score job fits and tailor application materials.

### How does the framework handle multiple job portal integrations?

The framework implements a modular skill architecture in **.agents/skills/** where each portal (Jobindex, LinkedIn, Jobbank, etc.) maintains its own subdirectory with a **SKILL.md** descriptor and a Bun-based CLI. Each skill exposes standardized `search` and `detail` commands returning JSON, allowing the `/scrape` command to aggregate listings uniformly regardless of source portal.

### Where does the framework store application history and tracker data?

Application state persists in **job_search_tracker.csv** at the repository root, which records posting metadata, fit scores, and status transitions. Individual application materials archive to **documents/applications/<company>_<role>/** subdirectories containing compiled PDFs, posting HTML, and **outcome.md** status files. Scraper deduplication states live in **job_scraper/seen_jobs.json**, while Gmail synchronization tokens reside in **gmail_sync/**.

### What security measures exist for executing scraper tools?

The framework implements defense-in-depth through **.claude/settings.json**, which explicitly allowlists permitted tools and scopes for Claude Code execution. Additionally, **tools/robots_check.py** enforces robots.txt compliance before HTTP requests, and **SECURITY.md** documents the comprehensive threat model and safe execution guidelines for generated automation scripts.