# How the AI Job Search Framework Auto-Discovers Installed Portal CLIs for Scraping

> Learn how the AI Job Search Framework auto-discovers installed portal CLIs. It parses SKILL.md files to extract entry points and flags without manual registration.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-09-01

---

**The AI Job Search Framework treats every job portal as a self-contained skill under `.agents/skills/`, automatically discovering installed CLIs by glob-matching [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files and parsing their frontmatter to extract entry points, flags, and enabled status without requiring manual registration.**

The `MadsLorentzen/ai-job-search` repository implements a plug-and-play architecture for job scraping that eliminates boilerplate configuration. When users invoke the `/scrape` command, the framework's orchestrator dynamically locates and executes portal-specific CLIs by reading structured metadata files rather than relying on hardcoded registries. This auto-discovery mechanism enables seamless integration of new job portals without modifying core framework code.

## The Skill-Based Architecture for Portal Integrations

### Self-Contained Portal Skills

Each job portal integration resides as an isolated skill within the `.agents/skills/` directory. Every skill includes a [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file that declares the CLI entry point, supported command-line flags, and an optional `enabled:` boolean flag that controls runtime visibility.

### Core vs. Agent Skill Directories

The discovery mechanism searches across two distinct locations to build the complete skill inventory:

- `.claude/skills/` — Core framework skills including the `job-scraper` orchestrator itself
- `.agents/skills/` — User-installed or third-party portal integrations

## How Auto-Discovery Works

### Glob-Matching SKILL.md Files

When the `/scrape` command executes, the system initiates filesystem discovery by locating all [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) manifests. According to [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py) at line 99, the framework builds the skill inventory using glob patterns that scan both core and agent directories:

```python
skills = sorted(ROOT.glob(".claude/skills/*/SKILL.md")) + \
         sorted(ROOT.glob(".agents/skills/*/SKILL.md"))

```

### Parsing CLI Metadata and Enabled Flags

The scraper orchestrator reads each [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) to extract the CLI invocation pattern and check the enabled status. As documented in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md) at line 61, the framework "discovers all installed portal CLI skills by reading every [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) found under `.agents/skills/*/SKILL.md`."

Line 63 of the same file specifies that the framework "honors the `enabled` toggle," treating portals as enabled unless the frontmatter explicitly sets `enabled: false`. This allows developers to keep portal code installed while excluding it from execution.

### Validating Against Security Guards

Before execution, discovered CLI patterns undergo validation against the allowed-tools registry defined in [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) (lines 46-51). This security layer ensures only explicitly permitted command patterns—such as `bun run .agents/skills/linkedin-search/cli/src/cli.ts *`—are eligible for subprocess execution.

### Executing Discovered CLIs

For each enabled portal, the scraper constructs the final command using the `bun run` pattern defined in the portal's [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md). The framework translates user search queries into portal-specific flags, executes the CLI via subprocess, and collects JSON output. Post-processing logic, such as client-side date filtering, applies as described in [`job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job-scraper/SKILL.md) lines 68-73.

### CI Pipeline Auto-Discovery

The discovery mechanism extends to continuous integration. The `discover-clis` job in [`.github/workflows/ci.yml`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.github/workflows/ci.yml) (lines 185-211) automatically builds a test matrix by scanning for [`package.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/package.json) files under each portal's `cli/` directory. This ensures new portals undergo automated testing without requiring manual updates to the workflow configuration.

## Implementation Examples

### Listing Discovered Portal Skills

The `/add-portal --list` command leverages the same glob logic found in the linter to display available integrations:

```python
from pathlib import Path
import yaml

ROOT = Path('.')
skill_files = sorted(ROOT.glob('.agents/skills/*/SKILL.md'))

for f in skill_files:
    meta = yaml.safe_load(f.read_text())
    print(f"{meta['name']}\t{meta.get('enabled', True)}")

```

### Running a Discovered CLI

The scraper executes enabled portal CLIs by parsing the `allowed-tools` pattern from [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) frontmatter:

```python
import subprocess
import json
import pathlib
import yaml

def run_portal_cli(skill_md_path, query):
    meta = yaml.safe_load(pathlib.Path(skill_md_path).read_text())
    cli_pattern = meta['allowed-tools'][0]
    # Extract command from pattern like "bun run path/to/cli.ts *"

    cli_cmd = cli_pattern.split('(')[1].split(')')[0]
    
    cmd = f"{cli_cmd} search --query \"{query}\" --format json".split()
    result = subprocess.run(cmd, capture_output=True, text=True, check=True)
    return json.loads(result.stdout)

# Aggregate postings from all enabled portals

for skill in discovered_skills:
    if skill.get('enabled', True):
        postings = run_portal_cli(skill['path'], user_query)

```

### CI Matrix Generation

The GitHub Actions workflow dynamically constructs the test matrix using `jq` to scan for CLI entry points:

```yaml
discover-clis:
  runs-on: ubuntu-latest
  outputs:
    tools: ${{ steps.discover.outputs.tools }}
  steps:
    - name: Discover portal CLIs
      id: discover
      run: |
        tools=$(jq -nc '
          [paths(scalars) as $p | select($p | test("^\\.agents/skills/.+/cli/package\\.json$"))]
          | map({"tool": "bun run " + ($p | join("/")) + " *"})')
        echo "::set-output name=tools::$tools"

```

## Key Files in the Discovery Pipeline

- **[`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py)** — Performs filesystem globbing for all [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) manifests (line 99), establishing the canonical list of available skills.
- **[`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md)** — Documents the auto-discovery protocol, the `enabled` flag behavior (lines 61-63), and post-processing requirements (lines 68-73).
- **[`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py)** — Validates CLI invocation patterns against an allowlist (lines 46-51), preventing unauthorized command execution.
- **[`.github/workflows/ci.yml`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.github/workflows/ci.yml)** — Contains the `discover-clis` job (lines 185-211) that auto-generates CI test matrices for portal CLIs.
- **`.agents/skills/*/SKILL.md`** — Individual portal configurations (e.g., [`linkedin-search/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/linkedin-search/SKILL.md)) defining entry points and capabilities.

## Summary

- The framework implements **zero-configuration portal discovery** by glob-matching [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files across `.claude/skills/` and `.agents/skills/` directories.
- Each portal's [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) frontmatter specifies the **CLI entry point**, **supported flags**, and an **optional `enabled` toggle** that defaults to true.
- The `job-scraper` skill orchestrates execution by validating commands against [`security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/security_guards.py) and invoking CLIs via the `bun run` pattern.
- **Disabled portals are automatically skipped** when their [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) contains `enabled: false`, eliminating runtime overhead for inactive integrations.
- The **CI pipeline mirrors runtime discovery**, automatically testing new portals by detecting [`cli/package.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/cli/package.json) files without manual workflow updates.

## Frequently Asked Questions

### What file triggers the auto-discovery of a new portal CLI?

The presence of a [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file within an `.agents/skills/<portal-name>/` directory triggers discovery. The framework specifically glob-matches `**/*/SKILL.md` patterns to identify candidate portals, then parses the YAML frontmatter to extract the CLI configuration.

### How does the framework prevent disabled portals from running?

The scraper checks the `enabled` key in each [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md)'s frontmatter before execution. As implemented in the discovery logic, portals default to enabled unless explicitly set to `enabled: false`, at which point they are omitted from the execution list entirely, incurring no runtime cost.

### Can I add a new portal without modifying the core framework code?

Yes. Adding a new portal requires only creating a subdirectory under `.agents/skills/` containing a properly formatted [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) and the CLI implementation. No changes to [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py), [`job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job-scraper/SKILL.md), or CI configuration are necessary—the framework discovers and integrates the new portal automatically upon the next `/scrape` invocation.

### How does the CI pipeline know which portal CLIs to test?

The `discover-clis` job in [`.github/workflows/ci.yml`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.github/workflows/ci.yml) independently scans the repository for [`package.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/package.json) files within portal CLI directories. Using `jq` to construct a JSON matrix of discovered tools, the pipeline automatically includes new portals in the test suite without requiring manual updates to the workflow file.