# Portal Search Skills in AI Job Search: How Declarative Specifications Drive Modular Job Scraping

> Discover portal search skills in AI job search. Learn how declarative Markdown specifications enable modular job scraping for any portal without core code changes.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: deep-dive
- Published: 2026-08-30

---

**Portal search skills are self‑contained Markdown specifications that define how to query a specific job portal and transform its raw output into a standardized format, enabling the scraper orchestrator to support any portal without core code changes.**

In the `MadsLorentzen/ai-job-search` repository, portal search skills implement a thin‑pointer design pattern that keeps agent runtimes synchronized through shared, declarative configuration files. This architecture separates portal‑specific query logic from the core scraping pipeline, making the system highly extensible.

---

## What Portal Search Skills Are

A **portal search skill** is a lightweight specification file—typically named [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md)—that lives under the `.agents/skills/<portal-name>/` directory. Each skill acts as a contract between the job search workflow and a particular job listing source (LinkedIn, Indeed, StackOverflow Jobs, etc.).

According to the repository's agent architecture documentation in [[`AGENTS.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/AGENTS.md)](https://github.com/MadsLorentzen/ai-job-search/blob/master/AGENTS.md), these skills are organized under:

- **`.agents/skills/`** – Portable skill definitions in the "Agent Skills" format
- **`.claude/skills/job-scraper/`** – The canonical scraper workflow that consumes these skills

Every skill file contains four essential components:

1. **CLI command or script** – The executable that performs the actual search
2. **Argument schema** – Required and optional parameters (keywords, location, pagination)
3. **Output mapping** – Rules for parsing the portal's response into the canonical job object schema
4. **Metadata** – Rate limits, authentication requirements, and other operational constraints

This design ensures that **all agent implementations** (Claude, Codex, Antigravity, or others) share the same definition, eliminating duplication and drift between runtimes.

---

## How Portal Search Skills Work

The scraper orchestrator interacts with portal search skills through a three‑stage pipeline:

### Stage 1: Skill Discovery and Loading

When the workflow needs to search a specific portal, it locates the corresponding skill directory and parses the [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file. The skill identifier is typically the portal name (lowercase, hyphenated).

```python
from pathlib import Path
import yaml

def load_skill(skill_name: str) -> dict:
    skill_path = Path(f".agents/skills/{skill_name}/SKILL.md")
    # Parse frontmatter and Markdown body

    return parse_skill_markdown(skill_path.read_text())

```

### Stage 2: Command Execution with Parameter Binding

The orchestrator extracts the CLI template from the skill and binds the workflow's search parameters to create an executable command.

Given this skill definition in [`.agents/skills/linkedin/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.agents/skills/linkedin/SKILL.md):

```markdown

# LinkedIn Job Search Skill

## CLI

```bash
linkedin-search --query "{keywords}" --location "{location}" --page {page}

```

## Arguments

- **keywords** – Search term(s) (string, required)
- **location** – City or region (string, required)
- **page** – Pagination index (integer, default 1)

## Output Mapping

```json
{
  "jobs": "$.results[*]",
  "title": "$.results[*].title",
  "company": "$.results[*].companyName",
  "location": "$.results[*].location",
  "url": "$.results[*].applyUrl",
  "postedDate": "$.results[*].datePosted"
}

```

## Rate Limit

- 10 requests per minute
- Requires `LINKEDIN_SESSION` environment variable

```

The orchestrator formats and executes:

```python
import subprocess
import json

def execute_skill(skill: dict, params: dict) -> list:
    # Bind parameters to CLI template

    command = skill["cli"].format(**params)
    
    result = subprocess.run(
        command,
        shell=True,
        capture_output=True,
        text=True,
        timeout=skill.get("timeout", 30)
    )
    
    if result.returncode != 0:
        raise SkillExecutionError(result.stderr)
    
    # Parse JSON response

    raw_data = json.loads(result.stdout)
    
    # Apply output mapping to transform to canonical schema

    return apply_output_mapping(raw_data, skill["output_mapping"])

```

### Stage 3: Normalization and Pipeline Integration

The **output mapping** uses JSONPath or similar selectors to extract fields from the portal's native response structure. This produces job objects that conform to the repository's internal schema, regardless of source:

```python
from jsonpath_ng import parse

def apply_output_mapping(raw_data: dict, mapping: dict) -> list:
    jobs = []
    
    # Extract array of job entries

    job_entries = [match.value for match in parse(mapping["jobs"]).find(raw_data)]
    
    for entry in job_entries:
        job = {
            "title": extract_field(entry, mapping["title"]),
            "company": extract_field(entry, mapping["company"]),
            "location": extract_field(entry, mapping["location"]),
            "url": extract_field(entry, mapping["url"]),
            "postedDate": extract_field(entry, mapping["postedDate"]),
            "source": skill["name"],  # e.g., "linkedin"

        }
        jobs.append(job)
    
    return jobs

```

The normalized jobs flow into downstream stages—ranking, filtering, and application generation—entirely agnostic to their original portal.

---

## Adding a New Portal Search Skill

The thin‑pointer architecture enables portal support without modifying core scraper code:

1. **Create the skill directory**: `mkdir .agents/skills/remoteok`
2. **Write [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md)** with CLI, arguments, output mapping, and metadata
3. **Test locally**: `python -m job_scraper --skill remoteok --params '{"keywords": "python"}'`
4. **Commit and deploy** – All runtimes automatically recognize the new portal

```markdown

# RemoteOK Job Search Skill

## CLI

```bash
curl -s "https://remoteok.com/api?tag={keywords}&page={page}"

```

## Arguments

- **keywords** – Job tag or search term (string, required)
- **page** – Page offset (integer, default 1)

## Output Mapping

```json
{
  "jobs": "$[*].",
  "title": "$[*].position",
  "company": "$[*].company",
  "location": "$[*].location or 'Remote'",
  "url": "$[*].url",
  "postedDate": "$[*].date"
}

```

## Rate Limit

- 1 request per 2 seconds
- Respect `X-RateLimit-Remaining` header

```

---

## Summary

- **Portal search skills** are declarative Markdown specifications in `.agents/skills/<portal>/SKILL.md` that define how to query and parse job listings from specific sources.
- The **scraper orchestrator** loads skills dynamically, binds workflow parameters to CLI templates, executes commands, and applies output mappings to produce normalized job objects.
- **Thin‑pointer design** ensures all agent runtimes share identical skill definitions, eliminating duplication across Claude, Codex, and other implementations.
- **Zero core code changes** are required to add new portals—simply create a new skill directory with the appropriate specification.

---

## Frequently Asked Questions

### What file format are portal search skills written in?

Portal search skills are written in **Markdown with YAML frontmatter**. The [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files use standard Markdown structure with fenced code blocks for CLI templates and JSON mappings, making them human-readable and parseable by any runtime. This format choice aligns with the repository's goal of being "readable, editable, and cross-runtime compatible" as specified in [`AGENTS.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/AGENTS.md).

### How does the scraper handle rate limits defined in skills?

The scraper extracts **rate limit metadata** from each skill's [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) and enforces them through a throttling layer. For example, a skill specifying "10 requests per minute" causes the orchestrator to queue or delay subsequent calls to that portal. Some skills also define headers like `X-RateLimit-Remaining` for dynamic adjustment, which the scraper monitors during execution.

### Can multiple portal search skills run in parallel?

Yes, the **canonical workflow** in `.claude/skills/job-scraper/` can dispatch searches across multiple portals concurrently because each skill is **self-contained and stateless**. The orchestrator manages thread pools or async tasks per portal, respecting individual rate limits while maximizing throughput. Job objects from all sources converge into a single normalized stream for downstream processing.

### What happens if a portal changes its API response format?

Only the affected **portal search skill** requires updating—specifically the `Output Mapping` section of its [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md). The core scraper logic remains unchanged. This isolation is the primary benefit of the thin‑pointer architecture: portal-specific adaptations are localized to declarative configuration files rather than scattered across Python/TypeScript implementations.