How the AI Job Search Framework Auto-Discovers Installed Portal CLIs for Scraping

The AI Job Search Framework treats every job portal as a self-contained skill under .agents/skills/, automatically discovering installed CLIs by glob-matching SKILL.md files and parsing their frontmatter to extract entry points, flags, and enabled status without requiring manual registration.

The MadsLorentzen/ai-job-search repository implements a plug-and-play architecture for job scraping that eliminates boilerplate configuration. When users invoke the /scrape command, the framework's orchestrator dynamically locates and executes portal-specific CLIs by reading structured metadata files rather than relying on hardcoded registries. This auto-discovery mechanism enables seamless integration of new job portals without modifying core framework code.

The Skill-Based Architecture for Portal Integrations

Self-Contained Portal Skills

Each job portal integration resides as an isolated skill within the .agents/skills/ directory. Every skill includes a SKILL.md file that declares the CLI entry point, supported command-line flags, and an optional enabled: boolean flag that controls runtime visibility.

Core vs. Agent Skill Directories

The discovery mechanism searches across two distinct locations to build the complete skill inventory:

  • .claude/skills/ — Core framework skills including the job-scraper orchestrator itself
  • .agents/skills/ — User-installed or third-party portal integrations

How Auto-Discovery Works

Glob-Matching SKILL.md Files

When the /scrape command executes, the system initiates filesystem discovery by locating all SKILL.md manifests. According to tools/lint_skills.py at line 99, the framework builds the skill inventory using glob patterns that scan both core and agent directories:

skills = sorted(ROOT.glob(".claude/skills/*/SKILL.md")) + \
         sorted(ROOT.glob(".agents/skills/*/SKILL.md"))

Parsing CLI Metadata and Enabled Flags

The scraper orchestrator reads each SKILL.md to extract the CLI invocation pattern and check the enabled status. As documented in .claude/skills/job-scraper/SKILL.md at line 61, the framework "discovers all installed portal CLI skills by reading every SKILL.md found under .agents/skills/*/SKILL.md."

Line 63 of the same file specifies that the framework "honors the enabled toggle," treating portals as enabled unless the frontmatter explicitly sets enabled: false. This allows developers to keep portal code installed while excluding it from execution.

Validating Against Security Guards

Before execution, discovered CLI patterns undergo validation against the allowed-tools registry defined in tools/security_guards.py (lines 46-51). This security layer ensures only explicitly permitted command patterns—such as bun run .agents/skills/linkedin-search/cli/src/cli.ts *—are eligible for subprocess execution.

Executing Discovered CLIs

For each enabled portal, the scraper constructs the final command using the bun run pattern defined in the portal's SKILL.md. The framework translates user search queries into portal-specific flags, executes the CLI via subprocess, and collects JSON output. Post-processing logic, such as client-side date filtering, applies as described in job-scraper/SKILL.md lines 68-73.

CI Pipeline Auto-Discovery

The discovery mechanism extends to continuous integration. The discover-clis job in .github/workflows/ci.yml (lines 185-211) automatically builds a test matrix by scanning for package.json files under each portal's cli/ directory. This ensures new portals undergo automated testing without requiring manual updates to the workflow configuration.

Implementation Examples

Listing Discovered Portal Skills

The /add-portal --list command leverages the same glob logic found in the linter to display available integrations:

from pathlib import Path
import yaml

ROOT = Path('.')
skill_files = sorted(ROOT.glob('.agents/skills/*/SKILL.md'))

for f in skill_files:
    meta = yaml.safe_load(f.read_text())
    print(f"{meta['name']}\t{meta.get('enabled', True)}")

Running a Discovered CLI

The scraper executes enabled portal CLIs by parsing the allowed-tools pattern from SKILL.md frontmatter:

import subprocess
import json
import pathlib
import yaml

def run_portal_cli(skill_md_path, query):
    meta = yaml.safe_load(pathlib.Path(skill_md_path).read_text())
    cli_pattern = meta['allowed-tools'][0]
    # Extract command from pattern like "bun run path/to/cli.ts *"

    cli_cmd = cli_pattern.split('(')[1].split(')')[0]
    
    cmd = f"{cli_cmd} search --query \"{query}\" --format json".split()
    result = subprocess.run(cmd, capture_output=True, text=True, check=True)
    return json.loads(result.stdout)

# Aggregate postings from all enabled portals

for skill in discovered_skills:
    if skill.get('enabled', True):
        postings = run_portal_cli(skill['path'], user_query)

CI Matrix Generation

The GitHub Actions workflow dynamically constructs the test matrix using jq to scan for CLI entry points:

discover-clis:
  runs-on: ubuntu-latest
  outputs:
    tools: ${{ steps.discover.outputs.tools }}
  steps:
    - name: Discover portal CLIs
      id: discover
      run: |
        tools=$(jq -nc '
          [paths(scalars) as $p | select($p | test("^\\.agents/skills/.+/cli/package\\.json$"))]
          | map({"tool": "bun run " + ($p | join("/")) + " *"})')
        echo "::set-output name=tools::$tools"

Key Files in the Discovery Pipeline

  • tools/lint_skills.py — Performs filesystem globbing for all SKILL.md manifests (line 99), establishing the canonical list of available skills.
  • .claude/skills/job-scraper/SKILL.md — Documents the auto-discovery protocol, the enabled flag behavior (lines 61-63), and post-processing requirements (lines 68-73).
  • tools/security_guards.py — Validates CLI invocation patterns against an allowlist (lines 46-51), preventing unauthorized command execution.
  • .github/workflows/ci.yml — Contains the discover-clis job (lines 185-211) that auto-generates CI test matrices for portal CLIs.
  • .agents/skills/*/SKILL.md — Individual portal configurations (e.g., linkedin-search/SKILL.md) defining entry points and capabilities.

Summary

  • The framework implements zero-configuration portal discovery by glob-matching SKILL.md files across .claude/skills/ and .agents/skills/ directories.
  • Each portal's SKILL.md frontmatter specifies the CLI entry point, supported flags, and an optional enabled toggle that defaults to true.
  • The job-scraper skill orchestrates execution by validating commands against security_guards.py and invoking CLIs via the bun run pattern.
  • Disabled portals are automatically skipped when their SKILL.md contains enabled: false, eliminating runtime overhead for inactive integrations.
  • The CI pipeline mirrors runtime discovery, automatically testing new portals by detecting cli/package.json files without manual workflow updates.

Frequently Asked Questions

What file triggers the auto-discovery of a new portal CLI?

The presence of a SKILL.md file within an .agents/skills/<portal-name>/ directory triggers discovery. The framework specifically glob-matches **/*/SKILL.md patterns to identify candidate portals, then parses the YAML frontmatter to extract the CLI configuration.

How does the framework prevent disabled portals from running?

The scraper checks the enabled key in each SKILL.md's frontmatter before execution. As implemented in the discovery logic, portals default to enabled unless explicitly set to enabled: false, at which point they are omitted from the execution list entirely, incurring no runtime cost.

Can I add a new portal without modifying the core framework code?

Yes. Adding a new portal requires only creating a subdirectory under .agents/skills/ containing a properly formatted SKILL.md and the CLI implementation. No changes to tools/lint_skills.py, job-scraper/SKILL.md, or CI configuration are necessary—the framework discovers and integrates the new portal automatically upon the next /scrape invocation.

How does the CI pipeline know which portal CLIs to test?

The discover-clis job in .github/workflows/ci.yml independently scans the repository for package.json files within portal CLI directories. Using jq to construct a JSON matrix of discovered tools, the pipeline automatically includes new portals in the test suite without requiring manual updates to the workflow file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →