# How the ai-job-search Project Handles Data: Architecture and File Management

> Discover how the ai-job-search project manages data. Learn about its architecture and file management, focusing on read-only state files and strict validation.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: architecture
- Published: 2026-09-02

---

**The ai-job-search project treats all persistent data as read-only state files stored in git-ignored JSON and CSV artifacts, accessed exclusively through encapsulated helper functions that enforce strict validation and prevent accidental version control exposure.**

The `MadsLorentzen/ai-job-search` repository implements a deterministic data handling strategy that isolates personal mutable state from code and configuration. By storing sensitive job search data in git-ignored files and exposing them only via audited commands, the project ensures user privacy while maintaining reproducible workflows. This article examines the specific file structures, validation mechanisms, and security controls that govern how the ai-job-search project handles data.

## Core Data Files and Their Purpose

The project maintains five distinct data artifacts, each serving a specific function within the job search pipeline. All files reside within the repository tree but are explicitly excluded from commits via `.gitignore`.

### salary_data.json

[`salary_data.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_data.json) stores company-wise salary benchmarks used by the **/apply** workflow to provide salary negotiation guidance. The file is loaded by [`salary_lookup.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_lookup.py) through a strict validation pipeline: `load_data()` → `validate_data()` → `read_raw_data()`. 

The `collect_validation_issues` function inspects the JSON structure and aborts with a clear error if the file is missing or malformed `salary_lookup.py#L40-L51`. This file is listed in `.gitignore` alongside `**/job_scraper/seen_jobs.json` to ensure salary data remains private and never enters version control.

### seen_jobs.json

[`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) serves as the canonical store for every scraped job posting, containing deduplication keys, status fields, ranking scores, deadlines, language gates, and gap/strength arrays. The file is populated exclusively by the **/scrape** command and updated only by **/rank**, which adds `rank_score`, `rank_verdict`, `strengths`, and `gaps` fields `.claude/commands/rank.md#L101-L163`.

Crucially, all modifications are additive—the schema is never restructured, guaranteeing that downstream commands like **/upskill** and **/apply** can reliably depend on existing fields. The file is deliberately git-ignored (defined at `.gitignore#L26`) to prevent personal job search history from being committed.

### job_search_tracker.csv

This CSV tracks submitted applications, recording dates, channels, statuses, notes, and CV/cover-letter paths. Unlike the JSON caches, `job_search_tracker.csv` is version-controlled to persist application history in Git. Most commands treat it as read-only, while **/apply** appends new rows and **/notion-sync** creates read-only views combining this data with [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) `.claude/commands/notion-sync.md#L1-L8`.

### Auxiliary State Files

[`notion_sync.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/notion_sync.json) holds disposable snapshots of Notion sync state and is never committed or read back into the main workflow. Similarly, `company_research/*.json` files serve as optional caches for external research data. Both are protected by the security guard's required ignore rules `tools/security_guards.py#L105-L111`.

## Data Access Patterns and Validation

All data interactions occur through well-encapsulated helper functions that enforce normalization, fuzzy matching, and schema validation.

### Encapsulated Helper Functions

The [`salary_lookup.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_lookup.py) CLI tool provides robust fuzzy-matching capabilities via the `search_company` function. This utility normalizes company names using `normalize`, `anglicize`, and `extract_core_words` before calculating relevance through `match_score_optimized` `salary_lookup.py#L63-L90`. The tool supports multiple output modes: listing all entries (`--list-all`), JSON export (`--json`), and dataset validation (`--validate`) `salary_lookup.py#L91-L100`.

### Validation and Error Handling

Before any command processes salary data, the `validate_data()` function runs `collect_validation_issues` to verify JSON integrity. If [`salary_data.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_data.json) is malformed or missing, the CLI aborts immediately with a descriptive error message, preventing corrupted data from propagating through the **/apply** workflow.

## Command Workflows and Data Flow

The project's command architecture ensures that data flows in one direction: scraped data is enriched, then referenced, but never circularly modified.

### The /scrape Command

The **/scrape** command (defined in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md)) reads [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json), creating the file if absent, and writes new job entries. Deduplication logic scans both [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) and `job_search_tracker.csv` to prevent duplicate entries `.claude/skills/job-scraper/SKILL.md#L120-L145`.

### The /rank Command

**Rank** loads [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json), computes ranking fields, and writes them back without altering existing keys. This additive approach ensures that `rank_score`, `gaps`, and `strengths` are appended to existing job objects while preserving fields required by subsequent commands `.claude/commands/rank.md#L15-L24`.

### The /apply Command

**Apply** reads the enriched [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) entry for a specific job, copies relevant fields (CV paths, deadlines, source URLs) into `job_search_tracker.csv`, but explicitly never modifies [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) `.claude/commands/apply.md#L353-L356`. This separation of concerns ensures the job cache remains immutable while the application history grows.

## Security and Privacy Controls

The repository implements automated safeguards to prevent accidental data exposure.

### Git Ignore Enforcement

The [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) script actively enforces git-ignore rules for all data files. It validates that [`salary_data.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_data.json), [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json), and `company_research/*.json` are excluded from version control `tools/security_guards.py#L61-L111`. Additionally, [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py) verifies that [`.claude/settings.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/settings.json) contains a proper allowlist before any command executes `tools/lint_skills.py#L12-L23`.

### Read-Only Architecture Benefits

By treating state files as read-only for most operations and restricting writes to specific audited commands, the architecture guarantees that personal data remains isolated from code changes. This design enables reproducibility—the job search workflow behaves deterministically regardless of user-specific data—while protecting sensitive information from accidental commits.

## Practical Code Examples

Validate your salary data before running application workflows:

```bash

# Validate salary data structure

python salary_lookup.py --validate

# List all companies in the benchmark

python salary_lookup.py --list-all

# Search with city filtering

python salary_lookup.py "Danske Bank" --city "København"

# Output JSON for scripting

python salary_lookup.py "Novo Nordisk" --json

```

Execute the standard data pipeline:

```bash

# Scrape jobs (creates/extends job_scraper/seen_jobs.json)

/claude/skills/job-scraper/SKILL.md

# Rank and enrich data (adds scores, gaps, strengths)

/claude/commands/rank.md

# Apply (reads seen_jobs.json, writes to job_search_tracker.csv)

/claude/commands/apply.md <job-url>

```

## Summary

- **Data isolation**: Personal state files ([`salary_data.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_data.json), [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json)) live in the repository but are git-ignored to prevent accidental commits.
- **Validation-first approach**: [`salary_lookup.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_lookup.py) validates JSON structure via `collect_validation_issues` before processing, aborting on malformed data.
- **Additive modifications**: The **/rank** command enriches [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) without restructuring schemas, ensuring downstream command compatibility.
- **Security automation**: [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) enforces ignore rules, while [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py) validates command configurations.
- **Immutable job cache**: **/apply** reads from [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) but only writes to `job_search_tracker.csv`, maintaining clean separation between scraped data and application history.

## Frequently Asked Questions

### How does ai-job-search prevent sensitive data from being committed to Git?

The project uses a multi-layered defense: `.gitignore` explicitly lists [`salary_data.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_data.json) and `**/job_scraper/seen_jobs.json`, while [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) actively scans for these files during pre-commit checks. Any attempt to stage ignored data files triggers an error, ensuring personal salary benchmarks and job histories remain local.

### What happens if salary_data.json is corrupted or missing?

The [`salary_lookup.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/salary_lookup.py) tool runs `validate_data()` which calls `collect_validation_issues` to inspect the JSON structure. If validation fails or the file is absent, the CLI aborts immediately with a descriptive error message, preventing the **/apply** command from proceeding with incomplete salary negotiation data.

### Can multiple commands modify seen_jobs.json simultaneously?

No. The architecture enforces a strict write pattern: only **/scrape** creates entries, and only **/rank** enriches them. Both commands use additive updates that append fields without removing existing keys. This design prevents race conditions and ensures that commands like **/upskill** and **/apply** can reliably access fields added by previous pipeline stages.

### Why is job_search_tracker.csv version-controlled while seen_jobs.json is not?

`job_search_tracker.csv` contains intentional application history that users may want to backup or audit over time, making it suitable for Git. Conversely, [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) contains transient scraped data and personal job search state that changes frequently and may contain sensitive company information, warranting exclusion from version control via `.gitignore`.