# What Data Sources Does ai-job-search Use for Job Postings?

> Discover the data sources ai-job-search uses for job postings. Learn how portal-specific CLIs and web-search queries power its job aggregation.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: data-sources
- Published: 2026-09-02

---

**The ai-job-search repository aggregates job postings through portal-specific CLIs located in the hidden `.agents/skills/` directory, complemented by fallback web-search queries defined in [`.claude/skills/job-scraper/search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/search-queries.md).**

The MadsLorentzen/ai-job-search project implements a modular scraping pipeline that decouples data acquisition from processing. Understanding what data sources ai-job-search uses for job postings requires examining the hidden `.agents` directory structure and the orchestration logic within the `.claude/skills/` folder.

## Portal-Specific CLIs in `.agents/skills/`

The primary data sources are **portal-specific command-line interfaces** (CLIs) residing in `.agents/skills/`. Each subdirectory represents a distinct job board integration (e.g., Indeed, LinkedIn, StackOverflow Jobs) and contains a [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file defining the contract for data retrieval.

### How Portal CLIs Work

Each portal skill implements a standardized interface. According to the source code structure, these CLIs return raw job postings as JSON objects containing fields such as `title`, `company`, `location`, `url`, and `posted_date`. The scraper executes these scripts to ingest data directly from their respective platforms.

### CLI Discovery and Execution

The scraper dynamically discovers available portals by scanning the `.agents/skills/` directory for [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files. The implementation locates associated [`cli.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/cli.py) scripts and executes them to retrieve postings:

```python
from pathlib import Path
import subprocess
import sys
import json

# Discover portal CLIs by scanning .agents/skills/

portal_cli_paths = Path(".agents/skills").rglob("SKILL.md")
portals = [p.parent for p in portal_cli_paths]

# Execute each portal CLI and collect postings

all_postings = []
for portal in portals:
    cli = portal / "cli.py"
    result = subprocess.run(
        [sys.executable, cli, "--json"], 
        capture_output=True, 
        text=True
    )
    postings = json.loads(result.stdout)
    all_postings.extend(postings)

```

This dynamic discovery mechanism allows the pipeline to support new job boards simply by adding new skill directories without modifying core logic.

## Fallback Web-Search Queries

When a portal CLI is unavailable or fails, the system falls back to generic web-search queries. These queries are defined in [`.claude/skills/job-scraper/search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/search-queries.md), which contains personalized search templates used to construct queries for external search engines.

The fallback mechanism ensures the pipeline remains resilient when specific job board APIs or scraping interfaces become inaccessible. The search patterns stored in this markdown file are parsed and executed as alternative data sources.

## Data Normalization and Provenance

After retrieval, all postings undergo normalization and deduplication. The system records provenance metadata identifying which portal CLI or fallback query supplied each posting, enabling downstream components like `/rank` and `/apply` to filter results based on source reliability.

Processed data persists to [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json), which maintains the accumulated dataset of previously encountered positions. This storage layer tracks the source of each entry, allowing the system to avoid reprocessing duplicate postings across multiple scraping sessions.

## Summary

- **Primary sources**: Portal-specific CLIs located in `.agents/skills/`, with each skill containing a [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) and [`cli.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/cli.py) for specific job boards like Indeed and LinkedIn.
- **Fallback mechanism**: Web-search queries defined in [`.claude/skills/job-scraper/search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/search-queries.md) trigger when portal CLIs are unavailable.
- **Discovery**: The scraper dynamically discovers data sources by scanning the `.agents/skills/` directory structure at runtime.
- **Storage**: Normalized postings with provenance metadata are stored in [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) for downstream processing.

## Frequently Asked Questions

### What job boards does ai-job-search support?

The repository supports any job board implemented as a portal skill in the `.agents/skills/` directory. Each job board requires a [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) specification and corresponding [`cli.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/cli.py) implementation. The modular architecture allows users to add support for additional platforms by creating new skill directories following the established contract.

### How does the scraper handle missing portal CLIs?

When a portal CLI is unavailable or fails to execute, the pipeline automatically falls back to generic web-search queries. These alternative data source patterns are stored in [`.claude/skills/job-scraper/search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/search-queries.md) and provide resilience against individual job board API changes or connectivity issues.

### Where are scraped job postings stored?

All retrieved job postings are persisted to [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) after normalization and deduplication. This JSON file maintains the complete dataset of encountered positions along with provenance metadata identifying which specific portal CLI or search query generated each entry.

### Can I add custom data sources to ai-job-search?

Yes. You can extend the data sources by creating new skill directories under `.agents/skills/`. Each new source requires a [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file defining the CLI contract and a [`cli.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/cli.py) script that outputs job postings in the expected JSON format. The scraper will automatically discover and execute your custom CLI on the next run.