# Portal Health Monitoring System in the AI Job Search Framework: Architecture and Usage

> Discover the portal health monitoring system in the AI Job Search Framework. Automatically detect broken job portal skills using failure analysis and sentinel probes.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: architecture
- Published: 2026-09-01

---

**The portal health monitoring system automatically detects degraded or broken job portal CLI skills by analyzing scrape results for failure patterns and executing sentinel probes when data quality issues are suspected.**

The AI Job Search Framework by MadsLorentzen embeds a robust portal health monitoring system to prevent silent data corruption when third-party job sites change their markup or block requests. Defined in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md) at step 4.75, this system operates in two distinct modes to balance diagnostic thoroughness against API rate limits. It inspects both real-time scrape results and historical data stored in [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) to determine whether a portal CLI skill is functioning correctly.

## How the Portal Health Monitoring System Operates

The system runs as an integrated component of the **Job Scraper** skill, evaluating each enabled portal after normal execution or upon explicit user request. Rather than assuming portal availability, it applies data quality heuristics and bounded probe requests to generate definitive health verdicts.

### Free-Pass Health Check Mode (Default)

In the default **free-pass** mode, the system analyzes results already returned during steps 0-4 of the scraping workflow without issuing additional network requests. For each enabled portal, it inspects the result set for obvious failure patterns including:

- Null `company` fields on every result
- Empty job titles or titles containing raw HTML entities  
- URLs that do not belong to the target portal domain

If a portal returns zero results, the system cross-references the historic [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) file (keyed by the `portal` field or URL domain) to verify whether the portal previously produced valid jobs. When historical data exists but current results are absent, the portal is flagged as **suspect** and escalates to the probe-only workflow.

### Probe-Only Health Check Mode (/scrape health)

When a user explicitly requests diagnostics via `/scrape health` or `/scrape health <portal>`, the system enters **probe-only** mode and skips normal search steps. Each portal receives a **sentinel probe**—a single search using the example query defined in its own [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file—limited to three results with JSON output formatting.

If the initial probe returns empty results, the framework retries once using a generic common word to distinguish between content availability and parsing failures. HTTP 429 responses (rate limiting) or block pages generate an **inconclusive (rate-limited)** verdict rather than marking the portal as broken, preventing false positives during temporary IP restrictions.

## Detailed Health Check Workflow

The portal health monitoring system follows a strict escalation path documented in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md) to minimize unnecessary network traffic while ensuring accurate diagnosis:

1. **Initial Inspection (Free-Pass)** – For each portal participating in step 1b, the framework applies degraded criteria to the existing result set as defined in lines 99-104 of the skill definition. Zero-result portals trigger historic comparison against [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) entries to identify suspect status.

2. **Escalation Probing** – Suspect portals undergo sentinel probe execution with `--format json` and a limit of 3 results, as specified in lines 106-107. Empty probe results trigger a single retry with a broad keyword.

3. **Verdict Assignment** – Persistent failures after retry mark the portal **broken**; 429 responses mark it **inconclusive**. The system generates a `health:` line in Step 5 output only for non-healthy states, listing the portal name, verdict, and root cause.

## Verdicts and User Interaction

The framework implements four distinct health states:

- **Healthy**: Portals passing all checks produce no health line in output (silence indicates success)
- **Degraded**: Portals returning results with data quality issues (e.g., null companies) receive this status  
- **Broken**: Portals failing sentinel probes after retry are marked broken
- **Inconclusive (Rate-Limited)**: Portals returning 429 status or block pages receive this temporary status

After presenting health results, the system offers to toggle the portal's `enabled` flag to `false` in its configuration. This allows the scraper to skip broken portals in future runs while maintaining the ability to fall back to WebSearch when necessary.

## CLI Commands and Usage Examples

Trigger the portal health monitoring system directly from the Claude CLI interface to diagnose specific or all installed portals.

### Triggering Health Checks

Check all installed portals:

```bash
/scrape health

```

Diagnose a specific portal regardless of its enabled status:

```bash
/scrape health jobnet

```

### Interpreting Health Output

The framework outputs health status in Step 5 using a standardized format:

```

health: jobnet - degraded (company null on all 12 results); parsing anchors in .agents/skills/jobnet/url-reference.md
health: jobindex - broken (0 results for the SKILL.md test query and a broader retry); parsing anchors in .agents/skills/jobindex/url-reference.md

```

Each line identifies the portal, the assigned verdict, and the specific technical cause triggering the classification.

### Programmatic Access Pattern

While the framework handles health checks internally, the logical structure follows this pattern:

```python
from job_scraper import Scraper

scraper = Scraper()
health_report = scraper.run_health_check(portal="jobnet")
print(health_report)  # Contains verdict, details, and suggested actions

```

## Key Configuration Files

The portal health monitoring system relies on these specific files within the `MadsLorentzen/ai-job-search` repository:

- **[`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md)**: Core skill definition containing the health check logic at step 4.75
- **`.agents/skills/*/SKILL.md`**: Individual portal definitions (e.g., [`jobnet-search/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/jobnet-search/SKILL.md)) containing example queries used for sentinel probes  
- **[`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json)**: Runtime database of historic job entries used to detect portals that previously produced results but now return none

## Summary

- The portal health monitoring system prevents silent failures by detecting when job portal CLI skills degrade or break
- **Free-pass mode** analyzes existing scrape results for data quality issues without additional network requests
- **Probe-only mode** executes sentinel probes with bounded retries when users request explicit health diagnostics via `/scrape health`
- The system distinguishes between permanent failures (**broken**), data quality issues (**degraded**), and temporary blocks (**inconclusive**)
- Historic data in [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) enables detection of portals that stop returning results after previous success
- Users can automatically disable broken portals while preserving WebSearch fallback capabilities

## Frequently Asked Questions

### What triggers a portal to be marked as "suspect" in the health monitoring system?

A portal receives **suspect** status when the free-pass health check detects zero current results but finds historical entries for that portal in [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json). This comparison indicates the portal previously produced valid job data but has now stopped returning results, signaling potential markup changes or access restrictions that require further investigation through sentinel probes.

### How does the framework distinguish between a broken portal and rate limiting?

The sentinel probe workflow treats HTTP 429 responses and block pages as **inconclusive (rate-limited)** rather than **broken**. This distinction prevents false positives during temporary IP restrictions. Only persistent failures after retry—where the probe returns zero results for both the specific example query and a generic common word—result in a **broken** verdict according to the logic in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md).

### Can I check the health of a disabled portal?

Yes. Using the `/scrape health <portal>` command allows you to diagnose specific portals regardless of their `enabled` status. This is useful for verifying whether a previously disabled portal has been fixed before re-enabling it for regular scraping operations.

### Where does the framework store the example queries used for health probes?

Each portal's example query is defined in its respective `.agents/skills/<portal-name>/SKILL.md` file. The health monitoring system reads these definitions to construct **sentinel probes**—test searches with `--format json` and a limit of 3 results—that verify whether the portal's CLI skill can successfully parse live data.