# How the Health Check Mechanism Detects Parser Rot in Portal Skills

> Discover how the Health Check mechanism detects parser rot in portal skills. Learn how it validates output against JSON schema and fails pipelines for invalid results.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-08-31

---

**The Health Check mechanism detects parser rot by executing every portal skill against a version-controlled fixture, validating the output against a canonical JSON schema, and failing the CI pipeline when parsers return invalid or null results.**

The `MadsLorentzen/ai-job-search` repository treats each job portal as a modular **skill** located in `/.agents/skills/`. Because external websites periodically refactor their HTML and API responses, parsers silently break over time—a degradation the project calls **parser rot**. The Health Check mechanism guards against this by continuously verifying that each skill can still extract structured job data from known-good samples.

## What Is Parser Rot?

Parser rot occurs when a job portal modifies its DOM structure, CSS selectors, or JSON response format without warning. A skill that previously extracted `title`, `company`, and `location` nodes suddenly returns `None` or raises exceptions. Without automated detection, these failures accumulate undetected until the entire aggregation pipeline produces empty datasets. The Health Check mechanism solves this by treating parser validation as a first-class testing concern.

## The Five-Step Health Check Pipeline

The detection logic resides primarily in [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py), with CI enforcement handled by [`tests/test_lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_lint_skills.py). The pipeline executes five discrete phases:

### Step 1: Discovering Portal Skills

The linter recursively scans the `/.agents/skills/` directory for Python modules exporting a `Skill` class. The `discover_skills()` generator yields tuples of `(skill_name, module)` for every file matching the pattern.

```python

# tools/lint_skills.py

SKILL_ROOT = Path(__file__).parent.parent / ".agents" / "skills"

def discover_skills():
    """Yield (skill_name, module) for every skill file."""
    for py_path in SKILL_ROOT.rglob("*.py"):
        mod_name = f".agents.skills.{py_path.stem}"
        module = importlib.import_module(mod_name)
        yield py_path.stem, module

```

### Step 2: Loading Fixture Data

For each discovered skill, the system loads a static **fixture** representing a recent real-world response from that portal. Fixtures live in `tests/fixtures/<portal_id>/sample.html` (or `.json`) and are committed to version control. This ensures the Health Check always parses a snapshot of data known to be valid at a previous point in time.

```python
fixture = Path("tests/fixtures") / skill.portal_id / "sample.html"
raw = fixture.read_text()

```

### Step 3: Executing the Parser

The linter instantiates the skill and invokes its `parse()` method with the raw fixture content. Any exception raised during execution—whether from missing DOM nodes or changed CSS selectors—immediately signals rot.

```python
skill = module.Skill()
parsed = skill.parse(raw)  # Rot detected if this raises

```

### Step 4: Schema Validation

Valid output must conform to the canonical job record schema defined in [`schemas/job_record.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/schemas/job_record.json). The `validate_job_record()` function in [`tools/validate_schema.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/validate_schema.py) asserts that required keys (`title`, `company`, `location`) exist and contain correct data types. Missing keys or type mismatches trigger validation failures.

```python
from tools.validate_schema import validate_job_record

try:
    validate_job_record(parsed)  # Schema enforcement

except ValidationError as exc:
    failures[name] = str(exc)

```

### Step 5: Reporting Rot

The `run_health_check()` function aggregates all failures into a structured report. If any skill fails parsing or validation, the script logs the specific portal and error message, then exits with a non-zero status code. This halts the CI pipeline and forces immediate remediation.

```python
def run_health_check():
    failures = {}
    for name, module in discover_skills():
        # ... parsing logic ...

        if failures:
            print("⚠️  Parser rot detected in the following skills:")
            for s, err in failures.items():
                print(f"  • {s}: {err}")
            raise SystemExit(1)

```

## Continuous Integration Enforcement

The [`tests/test_lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_lint_skills.py) file provides a thin wrapper that invokes [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py) during every pull request. Because the test suite aborts on any Health Check failure, **parser rot cannot reach the main branch**. Developers must update the fixture files or fix the skill logic before merging, ensuring the aggregation pipeline remains operational.

## Summary

- **Parser rot** is the silent degradation of portal parsers caused by external website changes.
- The **Health Check mechanism** lives in [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py) and validates every skill in `/.agents/skills/`.
- Validation uses **static fixtures** stored in `tests/fixtures/` to provide deterministic inputs.
- The **schema validator** in [`tools/validate_schema.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/validate_schema.py) enforces output conformity against [`schemas/job_record.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/schemas/job_record.json).
- **CI enforcement** via [`tests/test_lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_lint_skills.py) prevents broken parsers from being deployed.

## Frequently Asked Questions

### How does the Health Check know which fixture to use for each skill?

The linter maps each skill to its fixture using the `portal_id` attribute defined on the `Skill` class. It constructs the path `tests/fixtures/{portal_id}/sample.html` dynamically, ensuring every skill tests against its own representative data snapshot.

### What happens if a portal changes its layout and the parser breaks?

The next CI run will execute the Health Check, the `skill.parse()` call will raise an exception or return invalid data, and `validate_job_record()` will fail the schema check. The build logs identify the specific portal and error, alerting developers to update the CSS selectors or API parsing logic.

### Can I run the Health Check locally without triggering CI?

Yes. Execute `python -m tools.lint_skills` from the repository root. This runs the full discovery, parsing, and validation pipeline against local fixture files, allowing you to verify skill health before committing changes.

### Where is the canonical schema defined that validates parser output?

The schema is defined in [`schemas/job_record.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/schemas/job_record.json) at the repository root. The `validate_job_record()` function in [`tools/validate_schema.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/validate_schema.py) loads this JSON Schema and validates the Python dictionary returned by each skill's `parse()` method against it.