Portal Search Skills in AI Job Search: How Declarative Specifications Drive Modular Job Scraping
Portal search skills are self‑contained Markdown specifications that define how to query a specific job portal and transform its raw output into a standardized format, enabling the scraper orchestrator to support any portal without core code changes.
In the MadsLorentzen/ai-job-search repository, portal search skills implement a thin‑pointer design pattern that keeps agent runtimes synchronized through shared, declarative configuration files. This architecture separates portal‑specific query logic from the core scraping pipeline, making the system highly extensible.
What Portal Search Skills Are
A portal search skill is a lightweight specification file—typically named SKILL.md—that lives under the .agents/skills/<portal-name>/ directory. Each skill acts as a contract between the job search workflow and a particular job listing source (LinkedIn, Indeed, StackOverflow Jobs, etc.).
According to the repository's agent architecture documentation in [AGENTS.md](https://github.com/MadsLorentzen/ai-job-search/blob/master/AGENTS.md), these skills are organized under:
.agents/skills/– Portable skill definitions in the "Agent Skills" format.claude/skills/job-scraper/– The canonical scraper workflow that consumes these skills
Every skill file contains four essential components:
- CLI command or script – The executable that performs the actual search
- Argument schema – Required and optional parameters (keywords, location, pagination)
- Output mapping – Rules for parsing the portal's response into the canonical job object schema
- Metadata – Rate limits, authentication requirements, and other operational constraints
This design ensures that all agent implementations (Claude, Codex, Antigravity, or others) share the same definition, eliminating duplication and drift between runtimes.
How Portal Search Skills Work
The scraper orchestrator interacts with portal search skills through a three‑stage pipeline:
Stage 1: Skill Discovery and Loading
When the workflow needs to search a specific portal, it locates the corresponding skill directory and parses the SKILL.md file. The skill identifier is typically the portal name (lowercase, hyphenated).
from pathlib import Path
import yaml
def load_skill(skill_name: str) -> dict:
skill_path = Path(f".agents/skills/{skill_name}/SKILL.md")
# Parse frontmatter and Markdown body
return parse_skill_markdown(skill_path.read_text())
Stage 2: Command Execution with Parameter Binding
The orchestrator extracts the CLI template from the skill and binds the workflow's search parameters to create an executable command.
Given this skill definition in .agents/skills/linkedin/SKILL.md:
# LinkedIn Job Search Skill
## CLI
```bash
linkedin-search --query "{keywords}" --location "{location}" --page {page}
Arguments
- keywords – Search term(s) (string, required)
- location – City or region (string, required)
- page – Pagination index (integer, default 1)
Output Mapping
{
"jobs": "$.results[*]",
"title": "$.results[*].title",
"company": "$.results[*].companyName",
"location": "$.results[*].location",
"url": "$.results[*].applyUrl",
"postedDate": "$.results[*].datePosted"
}
Rate Limit
- 10 requests per minute
- Requires
LINKEDIN_SESSIONenvironment variable
The orchestrator formats and executes:
```python
import subprocess
import json
def execute_skill(skill: dict, params: dict) -> list:
# Bind parameters to CLI template
command = skill["cli"].format(**params)
result = subprocess.run(
command,
shell=True,
capture_output=True,
text=True,
timeout=skill.get("timeout", 30)
)
if result.returncode != 0:
raise SkillExecutionError(result.stderr)
# Parse JSON response
raw_data = json.loads(result.stdout)
# Apply output mapping to transform to canonical schema
return apply_output_mapping(raw_data, skill["output_mapping"])
Stage 3: Normalization and Pipeline Integration
The output mapping uses JSONPath or similar selectors to extract fields from the portal's native response structure. This produces job objects that conform to the repository's internal schema, regardless of source:
from jsonpath_ng import parse
def apply_output_mapping(raw_data: dict, mapping: dict) -> list:
jobs = []
# Extract array of job entries
job_entries = [match.value for match in parse(mapping["jobs"]).find(raw_data)]
for entry in job_entries:
job = {
"title": extract_field(entry, mapping["title"]),
"company": extract_field(entry, mapping["company"]),
"location": extract_field(entry, mapping["location"]),
"url": extract_field(entry, mapping["url"]),
"postedDate": extract_field(entry, mapping["postedDate"]),
"source": skill["name"], # e.g., "linkedin"
}
jobs.append(job)
return jobs
The normalized jobs flow into downstream stages—ranking, filtering, and application generation—entirely agnostic to their original portal.
Adding a New Portal Search Skill
The thin‑pointer architecture enables portal support without modifying core scraper code:
- Create the skill directory:
mkdir .agents/skills/remoteok - Write
SKILL.mdwith CLI, arguments, output mapping, and metadata - Test locally:
python -m job_scraper --skill remoteok --params '{"keywords": "python"}' - Commit and deploy – All runtimes automatically recognize the new portal
# RemoteOK Job Search Skill
## CLI
```bash
curl -s "https://remoteok.com/api?tag={keywords}&page={page}"
Arguments
- keywords – Job tag or search term (string, required)
- page – Page offset (integer, default 1)
Output Mapping
{
"jobs": "$[*].",
"title": "$[*].position",
"company": "$[*].company",
"location": "$[*].location or 'Remote'",
"url": "$[*].url",
"postedDate": "$[*].date"
}
Rate Limit
- 1 request per 2 seconds
- Respect
X-RateLimit-Remainingheader
---
## Summary
- **Portal search skills** are declarative Markdown specifications in `.agents/skills/<portal>/SKILL.md` that define how to query and parse job listings from specific sources.
- The **scraper orchestrator** loads skills dynamically, binds workflow parameters to CLI templates, executes commands, and applies output mappings to produce normalized job objects.
- **Thin‑pointer design** ensures all agent runtimes share identical skill definitions, eliminating duplication across Claude, Codex, and other implementations.
- **Zero core code changes** are required to add new portals—simply create a new skill directory with the appropriate specification.
---
## Frequently Asked Questions
### What file format are portal search skills written in?
Portal search skills are written in **Markdown with YAML frontmatter**. The [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files use standard Markdown structure with fenced code blocks for CLI templates and JSON mappings, making them human-readable and parseable by any runtime. This format choice aligns with the repository's goal of being "readable, editable, and cross-runtime compatible" as specified in [`AGENTS.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/AGENTS.md).
### How does the scraper handle rate limits defined in skills?
The scraper extracts **rate limit metadata** from each skill's [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) and enforces them through a throttling layer. For example, a skill specifying "10 requests per minute" causes the orchestrator to queue or delay subsequent calls to that portal. Some skills also define headers like `X-RateLimit-Remaining` for dynamic adjustment, which the scraper monitors during execution.
### Can multiple portal search skills run in parallel?
Yes, the **canonical workflow** in `.claude/skills/job-scraper/` can dispatch searches across multiple portals concurrently because each skill is **self-contained and stateless**. The orchestrator manages thread pools or async tasks per portal, respecting individual rate limits while maximizing throughput. Job objects from all sources converge into a single normalized stream for downstream processing.
### What happens if a portal changes its API response format?
Only the affected **portal search skill** requires updating—specifically the `Output Mapping` section of its [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md). The core scraper logic remains unchanged. This isolation is the primary benefit of the thin‑pointer architecture: portal-specific adaptations are localized to declarative configuration files rather than scattered across Python/TypeScript implementations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →