How the ai-job-search Project Handles Data: Architecture and File Management
The ai-job-search project treats all persistent data as read-only state files stored in git-ignored JSON and CSV artifacts, accessed exclusively through encapsulated helper functions that enforce strict validation and prevent accidental version control exposure.
The MadsLorentzen/ai-job-search repository implements a deterministic data handling strategy that isolates personal mutable state from code and configuration. By storing sensitive job search data in git-ignored files and exposing them only via audited commands, the project ensures user privacy while maintaining reproducible workflows. This article examines the specific file structures, validation mechanisms, and security controls that govern how the ai-job-search project handles data.
Core Data Files and Their Purpose
The project maintains five distinct data artifacts, each serving a specific function within the job search pipeline. All files reside within the repository tree but are explicitly excluded from commits via .gitignore.
salary_data.json
salary_data.json stores company-wise salary benchmarks used by the /apply workflow to provide salary negotiation guidance. The file is loaded by salary_lookup.py through a strict validation pipeline: load_data() → validate_data() → read_raw_data().
The collect_validation_issues function inspects the JSON structure and aborts with a clear error if the file is missing or malformed salary_lookup.py#L40-L51. This file is listed in .gitignore alongside **/job_scraper/seen_jobs.json to ensure salary data remains private and never enters version control.
seen_jobs.json
job_scraper/seen_jobs.json serves as the canonical store for every scraped job posting, containing deduplication keys, status fields, ranking scores, deadlines, language gates, and gap/strength arrays. The file is populated exclusively by the /scrape command and updated only by /rank, which adds rank_score, rank_verdict, strengths, and gaps fields .claude/commands/rank.md#L101-L163.
Crucially, all modifications are additive—the schema is never restructured, guaranteeing that downstream commands like /upskill and /apply can reliably depend on existing fields. The file is deliberately git-ignored (defined at .gitignore#L26) to prevent personal job search history from being committed.
job_search_tracker.csv
This CSV tracks submitted applications, recording dates, channels, statuses, notes, and CV/cover-letter paths. Unlike the JSON caches, job_search_tracker.csv is version-controlled to persist application history in Git. Most commands treat it as read-only, while /apply appends new rows and /notion-sync creates read-only views combining this data with seen_jobs.json .claude/commands/notion-sync.md#L1-L8.
Auxiliary State Files
notion_sync.json holds disposable snapshots of Notion sync state and is never committed or read back into the main workflow. Similarly, company_research/*.json files serve as optional caches for external research data. Both are protected by the security guard's required ignore rules tools/security_guards.py#L105-L111.
Data Access Patterns and Validation
All data interactions occur through well-encapsulated helper functions that enforce normalization, fuzzy matching, and schema validation.
Encapsulated Helper Functions
The salary_lookup.py CLI tool provides robust fuzzy-matching capabilities via the search_company function. This utility normalizes company names using normalize, anglicize, and extract_core_words before calculating relevance through match_score_optimized salary_lookup.py#L63-L90. The tool supports multiple output modes: listing all entries (--list-all), JSON export (--json), and dataset validation (--validate) salary_lookup.py#L91-L100.
Validation and Error Handling
Before any command processes salary data, the validate_data() function runs collect_validation_issues to verify JSON integrity. If salary_data.json is malformed or missing, the CLI aborts immediately with a descriptive error message, preventing corrupted data from propagating through the /apply workflow.
Command Workflows and Data Flow
The project's command architecture ensures that data flows in one direction: scraped data is enriched, then referenced, but never circularly modified.
The /scrape Command
The /scrape command (defined in .claude/skills/job-scraper/SKILL.md) reads seen_jobs.json, creating the file if absent, and writes new job entries. Deduplication logic scans both job_scraper/seen_jobs.json and job_search_tracker.csv to prevent duplicate entries .claude/skills/job-scraper/SKILL.md#L120-L145.
The /rank Command
Rank loads seen_jobs.json, computes ranking fields, and writes them back without altering existing keys. This additive approach ensures that rank_score, gaps, and strengths are appended to existing job objects while preserving fields required by subsequent commands .claude/commands/rank.md#L15-L24.
The /apply Command
Apply reads the enriched seen_jobs.json entry for a specific job, copies relevant fields (CV paths, deadlines, source URLs) into job_search_tracker.csv, but explicitly never modifies seen_jobs.json .claude/commands/apply.md#L353-L356. This separation of concerns ensures the job cache remains immutable while the application history grows.
Security and Privacy Controls
The repository implements automated safeguards to prevent accidental data exposure.
Git Ignore Enforcement
The tools/security_guards.py script actively enforces git-ignore rules for all data files. It validates that salary_data.json, seen_jobs.json, and company_research/*.json are excluded from version control tools/security_guards.py#L61-L111. Additionally, tools/lint_skills.py verifies that .claude/settings.json contains a proper allowlist before any command executes tools/lint_skills.py#L12-L23.
Read-Only Architecture Benefits
By treating state files as read-only for most operations and restricting writes to specific audited commands, the architecture guarantees that personal data remains isolated from code changes. This design enables reproducibility—the job search workflow behaves deterministically regardless of user-specific data—while protecting sensitive information from accidental commits.
Practical Code Examples
Validate your salary data before running application workflows:
# Validate salary data structure
python salary_lookup.py --validate
# List all companies in the benchmark
python salary_lookup.py --list-all
# Search with city filtering
python salary_lookup.py "Danske Bank" --city "København"
# Output JSON for scripting
python salary_lookup.py "Novo Nordisk" --json
Execute the standard data pipeline:
# Scrape jobs (creates/extends job_scraper/seen_jobs.json)
/claude/skills/job-scraper/SKILL.md
# Rank and enrich data (adds scores, gaps, strengths)
/claude/commands/rank.md
# Apply (reads seen_jobs.json, writes to job_search_tracker.csv)
/claude/commands/apply.md <job-url>
Summary
- Data isolation: Personal state files (
salary_data.json,seen_jobs.json) live in the repository but are git-ignored to prevent accidental commits. - Validation-first approach:
salary_lookup.pyvalidates JSON structure viacollect_validation_issuesbefore processing, aborting on malformed data. - Additive modifications: The /rank command enriches
seen_jobs.jsonwithout restructuring schemas, ensuring downstream command compatibility. - Security automation:
tools/security_guards.pyenforces ignore rules, whiletools/lint_skills.pyvalidates command configurations. - Immutable job cache: /apply reads from
seen_jobs.jsonbut only writes tojob_search_tracker.csv, maintaining clean separation between scraped data and application history.
Frequently Asked Questions
How does ai-job-search prevent sensitive data from being committed to Git?
The project uses a multi-layered defense: .gitignore explicitly lists salary_data.json and **/job_scraper/seen_jobs.json, while tools/security_guards.py actively scans for these files during pre-commit checks. Any attempt to stage ignored data files triggers an error, ensuring personal salary benchmarks and job histories remain local.
What happens if salary_data.json is corrupted or missing?
The salary_lookup.py tool runs validate_data() which calls collect_validation_issues to inspect the JSON structure. If validation fails or the file is absent, the CLI aborts immediately with a descriptive error message, preventing the /apply command from proceeding with incomplete salary negotiation data.
Can multiple commands modify seen_jobs.json simultaneously?
No. The architecture enforces a strict write pattern: only /scrape creates entries, and only /rank enriches them. Both commands use additive updates that append fields without removing existing keys. This design prevents race conditions and ensures that commands like /upskill and /apply can reliably access fields added by previous pipeline stages.
Why is job_search_tracker.csv version-controlled while seen_jobs.json is not?
job_search_tracker.csv contains intentional application history that users may want to backup or audit over time, making it suitable for Git. Conversely, seen_jobs.json contains transient scraped data and personal job search state that changes frequently and may contain sensitive company information, warranting exclusion from version control via .gitignore.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →