AI Job Search Framework: Key Files and Their Purposes Explained
The AI Job Search Framework organizes job application automation through three architectural layers—Profile & Rules (CLAUDE.md, .claude/), Search & Automation (.agents/skills/, tools/), and Artifacts & State (documents/, job_search_tracker.csv)—that orchestrate candidate profiles, multi-portal scraping, and application lifecycle management.
The MadsLorentzen/ai-job-search repository implements a thin-pointer design where declarative specifications drive CLI-based automation. This open-source framework separates immutable workflow logic from mutable personal data, enabling safe fork updates while maintaining a centralized job_search_tracker.csv as the source of truth for all application activities.
Profile and Configuration Layer
The Canonical Candidate Profile (CLAUDE.md)
At the root of the repository, CLAUDE.md serves as the single source of truth for candidate identity. This markdown file consolidates personal data, education history, work experience, skills matrices, career goals, and evaluation rubrics (structured as 01-candidate-profile.md through 07-interview-prep.md). During initialization, the /setup command reads either this file or the documents/ folder to construct a structured profile that feeds into subsequent fit evaluation logic.
The profile stored in CLAUDE.md directly powers the ranking algorithms in 04-job-evaluation.md, enabling commands like /rank and /apply to score job postings against explicit candidate criteria and tailor CV content dynamically.
Claude Code Commands and Skills (.claude/ directory)
The .claude/commands/ directory contains declarative markdown specifications that Claude Code reads at runtime. These files implement the primary workflow commands:
/setup– Initializes or updates the candidate profile/apply– Generates tailored application materials/scrape– Invokes job portal agents/rank– Batch-scores postings against the profile/outcome– Records application results/gmail-sync– Updates tracker from email status
Each command file follows a strict contract defining inputs, outputs, and tool invocations. The .claude/skills/ directory houses core skill definitions including job-application-assistant, job-scraper, and upskill, with each skill containing a SKILL.md descriptor that specifies its contract, enabled flags, and execution parameters.
Permission Settings (.claude/settings.json)
The .claude/settings.json file maintains the permission allowlist for Claude Code execution. This JSON configuration explicitly enumerates which tools may be executed and defines shared scopes, enforcing security boundaries between the AI agent and local system resources.
Search and Automation Layer
Job Portal Scrapers (.agents/skills/)
The .agents/skills/ directory contains modular CLI tools for specific job portals. Each portal skill (e.g., jobbank-search, jobindex-search, linkedin-search, freehire-search) follows a standardized contract requiring three components:
- A
searchcommand returning JSON listings - A
detailcommand for single posting retrieval - A SKILL.md descriptor defining the interface
These scrapers are built on Bun/TypeScript CLI implementations residing in each skill's cli/ subdirectory. The framework's /scrape command iterates through enabled skills, aggregating results into the central tracker while respecting robots.txt compliance through enforcement tools.
Installation requires zero runtime dependencies for most portals:
# Install all bundled portal tools
for tool in jobbank-search jobdanmark-search jobindex-search jobnet-search \
linkedin-search freehire-search; do
(cd .agents/skills/$tool/cli && bun install)
done
# Execute a search on Jobindex
.jobbank-search/cli/search --query "data scientist" --location "Copenhagen" --format json
CI and Utility Scripts (tools/)
The tools/ directory contains Python-based infrastructure scripts that maintain framework integrity:
- tools/check_framework_version.py – Verifies that AGENTS.md reflects version bumps when skill files change, preventing stale releases
- tools/check_upstream_updates.py – Previews which personalized files upstream updates will modify, facilitating safe fork synchronization
- tools/upstream_triage.py – Categorizes upstream commits into "review-required" versus "safe-to-skip" for weekly maintenance workflows
- tools/robots_check.py – Validates robots.txt compliance before scraper retries
- tools/lint_skills.py – CI linter validating SKILL.md file shapes, settings syntax, and manifest completeness
Salary Benchmarking (salary_lookup.py)
The salary_lookup.py utility provides compensation analysis by reading user-supplied salary_data.json (documented in tools/README_SALARY_TOOL.md). When invoked during the ranking phase, this script scores posting compensation against market data, outputting normalized benchmarks that influence the fit scoring algorithm.
Artifacts and State Management
Document Storage and Archives (documents/)
The documents/ directory implements a strict folder hierarchy defined in documents/README.md. This hierarchy stores raw career materials (PDF CVs, LinkedIn exports, diplomas, reference letters) and application archives following the naming convention <company>_<role>.
When /apply generates materials for a specific posting, it creates a subdirectory under documents/applications/ containing:
posting.html– Raw job description (optional)cv.pdf– Compiled CV tailored to the rolecover.pdf– Compiled cover letteroutcome.md– Status tracking (drafted, interviewed, offer, rejected)
The "Subfolder naming" rule in documents/README.md automatically derives these paths from company and role metadata, ensuring consistent archive organization.
CV and Cover Letter Templates (cv/ and cover_letters/)
Default document generation relies on LaTeX templates:
- cv/main_example.tex – Stock moderncv template compiled during
/apply - cover_letters/cover.cls – Custom LaTeX class defining cover letter formatting
- cover_letters/cover_example.tex – Default cover letter template
The templates/ directory supports custom template registration via /add-template, accepting LaTeX, Typst, or other compile-to-PDF formats. templates/README.md documents the registration protocol for extending output formats beyond the default moderncv implementation.
Application Tracking (job_search_tracker.csv and Runtime State)
job_search_tracker.csv functions as the central state database, recording every scraped posting, computed fit scores, and application status transitions. This CSV is the primary interface for /rank, /outcome, and /html-report commands.
Runtime state persists in specialized directories:
- job_scraper/ – Contains
seen_jobs.jsonand cached results enabling incremental scraping and deduplication - gmail_sync/ – Stores processed message IDs and last sync timestamps for the
/gmail-synccommand - upskill/ – Receives markdown reports from
/upskillskill-gap analyses
Framework Workflow Integration
A typical job search session orchestrates these files through Claude Code commands:
# 1. Initialize or update candidate profile
claude
/setup # Interactive profile construction from documents/ or CLAUDE.md
# 2. Search across all configured portals
/scrape # Invokes every skill in .agents/skills/ and populates tracker
# 3. Rank opportunities against profile (optional)
/rank # Batch-scores new postings, updates job_search_tracker.csv
# 4. Generate application materials
/apply https://example.com/job/1234567
# Compiles cv/main_example.tex and cover_letters/cover_example.tex
# Archives to documents/applications/Company_Role/
# 5. Record interview outcomes
/outcome # Updates outcome.md and tracker status
# 6. Sync status from Gmail
/gmail-sync # Processes mailbox for application status changes
Additional configuration files include SETUP.md (installation prerequisites and TeX setup), SECURITY.md (threat model and safe execution guidelines), and AGENTS.md (agent runtime selection for Claude Code, Codex, or Gemini CLI).
Summary
- CLAUDE.md stores the canonical candidate profile and evaluation rubrics that drive all personalization logic
- .claude/commands/ and .claude/skills/ contain the declarative specifications that Claude Code executes for setup, scraping, ranking, and application workflows
- .agents/skills/ houses modular, Bun-based job portal scrapers following standardized JSON output contracts
- tools/ provides CI infrastructure including version checking, upstream triage, robots.txt validation, and skill linting
- salary_lookup.py enables compensation benchmarking against user-supplied market data
- documents/ maintains the archive hierarchy for raw career materials and application-specific outputs
- cv/ and cover_letters/ contain the default LaTeX templates compiled by the
/applycommand - job_search_tracker.csv serves as the central state database for all job postings, scores, and application statuses
Frequently Asked Questions
What is the purpose of CLAUDE.md in the AI Job Search Framework?
CLAUDE.md acts as the master candidate profile containing personal data, work history, skills, goals, and structured evaluation rubrics. During /setup, Claude Code either parses this file or imports data from the documents/ folder to build a structured profile. All subsequent commands—/rank, /apply, and /outcome—reference this profile to score job fits and tailor application materials.
How does the framework handle multiple job portal integrations?
The framework implements a modular skill architecture in .agents/skills/ where each portal (Jobindex, LinkedIn, Jobbank, etc.) maintains its own subdirectory with a SKILL.md descriptor and a Bun-based CLI. Each skill exposes standardized search and detail commands returning JSON, allowing the /scrape command to aggregate listings uniformly regardless of source portal.
Where does the framework store application history and tracker data?
Application state persists in job_search_tracker.csv at the repository root, which records posting metadata, fit scores, and status transitions. Individual application materials archive to documents/applications/_/ subdirectories containing compiled PDFs, posting HTML, and outcome.md status files. Scraper deduplication states live in job_scraper/seen_jobs.json, while Gmail synchronization tokens reside in gmail_sync/.
What security measures exist for executing scraper tools?
The framework implements defense-in-depth through .claude/settings.json, which explicitly allowlists permitted tools and scopes for Claude Code execution. Additionally, tools/robots_check.py enforces robots.txt compliance before HTTP requests, and SECURITY.md documents the comprehensive threat model and safe execution guidelines for generated automation scripts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →