How the AI Job Search Framework Discovers Job Portal CLIs: Agent-Skills Architecture Explained
The AI Job Search Framework discovers job portal CLIs by scanning the .agents/skills/ directory for SKILL.md files, respecting the enabled: metadata flag, and invoking the executable defined in each portal's cli/package.json.
The AI Job Search Framework eliminates hard-coded portal integrations through a portable Agent-Skills convention. By treating every job portal as a self-contained skill module, the framework dynamically discovers and executes CLIs at both runtime and CI time. This architecture allows new portals to be added simply by creating a folder under .agents/skills/ without modifying central registries or hard-coded lists.
Directory Scanning and SKILL.md Resolution
Discovery begins with a filesystem scan targeting the pattern *.agents/skills/*/SKILL.md. As documented in README.md at line 325, the framework walks this directory structure to identify potential portal integrations. Each skill folder must contain a SKILL.md file that declares an enabled: boolean flag defaulting to true when unspecified.
The system loads only skills marked as enabled, looking for the CLI entry point inside the skill’s cli/ subdirectory. This structure ensures every portal remains self-contained, bundling both its metadata declaration and runtime logic in a predictable location that the framework can inspect without prior configuration.
Contract Enforcement for Standardized CLIs
Every discovered portal must adhere to a strict contract defined in AGENTS.md at line 19. This contract requires each CLI to expose two primary subcommands:
search– Execute job searches with query parametersdetail– Retrieve specific job listing details
Additionally, all portal CLIs must accept a --format argument supporting json, table, or plain output modes. The SKILL.md file serves as the contract manifest, declaring the enabled: status and portal metadata. Because the scraper invokes these CLIs uniformly without further wiring, any portal following this contract integrates automatically without framework modifications.
Runtime Discovery via the /scrape Workflow
Actual runtime discovery occurs within the /scrape command implementation. As specified in .claude/skills/job-scraper/search-queries.md at line 7, the job-scraper workflow explicitly discovers every portal skill under .agents/skills/*/SKILL.md and executes the associated CLI found in that folder’s cli/ directory.
When a user triggers /scrape, the framework performs a live filesystem scan, identifies enabled skills, and launches their respective CLIs with standardized arguments. The real implementation resides in the .claude/skills/job-scraper module, which treats all portal CLIs uniformly regardless of the underlying implementation language or API.
CI-Time Auto-Matrix with discover-clis
Discovery extends beyond runtime into continuous integration. The GitHub Actions workflow defined in .github/workflows/ci.yml at line 185 implements a discover-clis job that automatically builds test matrices for every portal CLI.
This job globs for *.agents/skills/*/cli/package.json files, converting the discovered paths into a JSON matrix using jq. By generating this matrix dynamically, the CI pipeline guarantees that every portal skill—including newly added ones—undergoes type-checking and testing on every push without requiring updates to the workflow file itself.
Practical Discovery Implementation
The discovery logic follows a straightforward pattern. This Python pseudo-code illustrates how the framework identifies and invokes portal CLIs:
import pathlib, yaml, subprocess, json
def discover_portal_clis(base=".agents/skills"):
# Find every SKILL.md under the skills directory
for skill_path in pathlib.Path(base).rglob("SKILL.md"):
meta = yaml.safe_load(skill_path.read_text())
if not meta.get("enabled", True):
continue # skip disabled portals
cli_dir = skill_path.parent / "cli"
pkg = json.loads((cli_dir / "package.json").read_text())
# The CLI entry point is defined in package.json → "bin" or "main"
cli_cmd = pkg.get("bin") or pkg.get("main")
# Run the portal's `search` command (example)
subprocess.run([str(cli_dir / cli_cmd), "search", "--format", "json"])
The framework reads the package.json to resolve the executable path—typically specified in the bin or main field—then invokes it with the required arguments.
Automated Testing Matrix Generation
The CI pipeline implements dynamic matrix generation through shell scripting. As defined in the discover-clis job in .github/workflows/ci.yml:
discover-clis:
runs-on: ubuntu-latest
outputs:
tools: ${{ steps.discover.outputs.tools }}
steps:
- name: Find portal CLIs
id: discover
run: |
tools=$(jq -nc '[inputs | select(test("\\.agents/skills/.*/cli/package\\.json$"))]' \
<(find . -path "./.agents/skills/*/cli/package.json"))
echo "tools=$tools" >> $GITHUB_OUTPUT
This configuration ensures that adding a new portal automatically includes it in the CI test matrix without workflow modifications.
Adding New Portals via /add-portal
New portals integrate into the discovery system through the /add-portal command. When executed, this helper generates a new folder structure under .agents/skills/ complete with a SKILL.md file pre-configured with enabled: true. For example:
# In the framework shell
/add-portal linkedin-search
The new skill immediately becomes discoverable by both the /scrape command and the CI pipeline. The framework will locate the new skill at .agents/skills/linkedin-search/SKILL.md and execute its CLI entry point defined in .agents/skills/linkedin-search/cli/package.json on the next run.
Summary
- Filesystem-based discovery: The framework scans
.agents/skills/*/SKILL.mdto identify portal integrations at runtime and CI time without maintaining a central registry. - Enablement flags: Each portal controls its own activation through the
enabled:field inSKILL.md, defaulting to active when unspecified. - Standardized contracts: All portal CLIs must implement
searchanddetailcommands with--format json|table|plainsupport, ensuring uniform treatment by the scraper. - Dual-context execution: The same discovery logic runs during interactive
/scrapecommands and automated CI validation via thediscover-clisjob. - Zero-registry addition: New portals integrate by creating a folder under
.agents/skills/, making the system truly plug-and-play.
Frequently Asked Questions
How does the AI Job Search Framework know which portals to scrape?
The framework scans the .agents/skills/ directory for SKILL.md files at both runtime and CI time. Each skill folder contains metadata including an enabled: flag that determines whether the portal participates in scraping. This eliminates the need for hard-coded portal lists in the core framework code.
What happens if a portal CLI is disabled or misconfigured?
If a portal's SKILL.md contains enabled: false or lacks the required contract implementation, the framework simply skips that directory during discovery. The modular design ensures that one broken or disabled portal does not affect the operation of others or crash the scraping process.
How does adding a new portal affect the CI pipeline?
The discover-clis job in .github/workflows/ci.yml automatically detects new portals by globbing for package.json files within .agents/skills/*/cli/. When you run /add-portal to create a new skill, the CI pipeline includes it in the test matrix on the next push without requiring workflow file modifications.
Where is the actual CLI entry point defined for each portal?
The entry point is declared in the portal's cli/package.json file, typically under the bin or main field. The discovery logic reads this file to determine the executable path relative to the skill directory, then invokes it with standardized arguments like search --format json.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →