What Data Sources Does ai-job-search Use for Job Postings?
The ai-job-search repository aggregates job postings through portal-specific CLIs located in the hidden .agents/skills/ directory, complemented by fallback web-search queries defined in .claude/skills/job-scraper/search-queries.md.
The MadsLorentzen/ai-job-search project implements a modular scraping pipeline that decouples data acquisition from processing. Understanding what data sources ai-job-search uses for job postings requires examining the hidden .agents directory structure and the orchestration logic within the .claude/skills/ folder.
Portal-Specific CLIs in .agents/skills/
The primary data sources are portal-specific command-line interfaces (CLIs) residing in .agents/skills/. Each subdirectory represents a distinct job board integration (e.g., Indeed, LinkedIn, StackOverflow Jobs) and contains a SKILL.md file defining the contract for data retrieval.
How Portal CLIs Work
Each portal skill implements a standardized interface. According to the source code structure, these CLIs return raw job postings as JSON objects containing fields such as title, company, location, url, and posted_date. The scraper executes these scripts to ingest data directly from their respective platforms.
CLI Discovery and Execution
The scraper dynamically discovers available portals by scanning the .agents/skills/ directory for SKILL.md files. The implementation locates associated cli.py scripts and executes them to retrieve postings:
from pathlib import Path
import subprocess
import sys
import json
# Discover portal CLIs by scanning .agents/skills/
portal_cli_paths = Path(".agents/skills").rglob("SKILL.md")
portals = [p.parent for p in portal_cli_paths]
# Execute each portal CLI and collect postings
all_postings = []
for portal in portals:
cli = portal / "cli.py"
result = subprocess.run(
[sys.executable, cli, "--json"],
capture_output=True,
text=True
)
postings = json.loads(result.stdout)
all_postings.extend(postings)
This dynamic discovery mechanism allows the pipeline to support new job boards simply by adding new skill directories without modifying core logic.
Fallback Web-Search Queries
When a portal CLI is unavailable or fails, the system falls back to generic web-search queries. These queries are defined in .claude/skills/job-scraper/search-queries.md, which contains personalized search templates used to construct queries for external search engines.
The fallback mechanism ensures the pipeline remains resilient when specific job board APIs or scraping interfaces become inaccessible. The search patterns stored in this markdown file are parsed and executed as alternative data sources.
Data Normalization and Provenance
After retrieval, all postings undergo normalization and deduplication. The system records provenance metadata identifying which portal CLI or fallback query supplied each posting, enabling downstream components like /rank and /apply to filter results based on source reliability.
Processed data persists to job_scraper/seen_jobs.json, which maintains the accumulated dataset of previously encountered positions. This storage layer tracks the source of each entry, allowing the system to avoid reprocessing duplicate postings across multiple scraping sessions.
Summary
- Primary sources: Portal-specific CLIs located in
.agents/skills/, with each skill containing aSKILL.mdandcli.pyfor specific job boards like Indeed and LinkedIn. - Fallback mechanism: Web-search queries defined in
.claude/skills/job-scraper/search-queries.mdtrigger when portal CLIs are unavailable. - Discovery: The scraper dynamically discovers data sources by scanning the
.agents/skills/directory structure at runtime. - Storage: Normalized postings with provenance metadata are stored in
job_scraper/seen_jobs.jsonfor downstream processing.
Frequently Asked Questions
What job boards does ai-job-search support?
The repository supports any job board implemented as a portal skill in the .agents/skills/ directory. Each job board requires a SKILL.md specification and corresponding cli.py implementation. The modular architecture allows users to add support for additional platforms by creating new skill directories following the established contract.
How does the scraper handle missing portal CLIs?
When a portal CLI is unavailable or fails to execute, the pipeline automatically falls back to generic web-search queries. These alternative data source patterns are stored in .claude/skills/job-scraper/search-queries.md and provide resilience against individual job board API changes or connectivity issues.
Where are scraped job postings stored?
All retrieved job postings are persisted to job_scraper/seen_jobs.json after normalization and deduplication. This JSON file maintains the complete dataset of encountered positions along with provenance metadata identifying which specific portal CLI or search query generated each entry.
Can I add custom data sources to ai-job-search?
Yes. You can extend the data sources by creating new skill directories under .agents/skills/. Each new source requires a SKILL.md file defining the CLI contract and a cli.py script that outputs job postings in the expected JSON format. The scraper will automatically discover and execute your custom CLI on the next run.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →