# How Internal Skill Filtering and Categorization Works in Antigravity Awesome Skills

> Discover how Antigravity Awesome Skills filters and categorizes internal skills using a two-stage pipeline with options for skill filtering, deterministic categorization, and confidence scoring.

- Repository: [sickn33/antigravity-awesome-skills](https://github.com/sickn33/antigravity-awesome-skills)
- Tags: internals
- Published: 2026-03-18

---

**The Antigravity Awesome Skills repository implements a two-stage pipeline where `SkillScanner` first discovers skills from the filesystem and optionally filters them by name via the `skill_filter` parameter, then `generate_index` categorizes each skill through a deterministic four-tier priority chain (explicit front-matter → folder name → keyword matching → dynamic ID parsing) with confidence scoring and provenance tracking.**

The `sickn33/antigravity-awesome-skills` repository manages modular capabilities through a sophisticated internal skill filtering and categorization system. This architecture separates discovery and filtering concerns from classification logic, enabling both targeted audits of individual skills and automated organization of the entire skill library. Understanding these mechanisms is essential for contributors maintaining the index or running the Sentinel audit workflow against specific capabilities.

## Skill Filtering in the Sentinel Audit Workflow

The **Sentinel audit workflow** limits processing scope through deterministic in-process filtering that operates after initial filesystem discovery. This ensures that targeted audits never perform redundant disk operations.

### The SkillScanner Discovery Phase

The `SkillScanner` class in [`skills/skill-sentinel/scripts/scanner.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/skills/skill-sentinel/scripts/scanner.py) (lines 53‑61) walks the `skills/` tree and reads each [`SKILL.md`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/SKILL.md) front-matter. The `discover_all()` method returns a complete list of skill dictionaries containing metadata such as `name`, `id`, and `description`.

### Applying the skill_filter Parameter

In [`skills/skill-sentinel/scripts/run_audit.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/skills/skill-sentinel/scripts/run_audit.py) (lines 96‑100), the optional `skill_filter` parameter (passed from the CLI via `--skill`) applies a list comprehension to the discovered set:

```python
if skill_filter:
    all_skills = [s for s in all_skills if s["name"] == skill_filter]
    if not all_skills:
        return {"error": f"Skill '{skill_filter}' nao encontrada."}

```

This filtering stage removes every entry whose `name` does not match the filter, returning an error if no matches exist.

### Direct Skill Discovery with discover_skill

For single-skill lookups, `SkillScanner` offers the `discover_skill(name)` helper (lines 63‑66 in [`scanner.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/scanner.py)). This method bypasses full enumeration by returning a single skill record directly, optimizing performance when only one target is required.

## Skill Categorization and Index Generation

Categorization occurs during index generation via [`skill_categorization/tools/scripts/generate_index.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/skill_categorization/tools/scripts/generate_index.py). The system assigns categories through a deterministic priority chain that records confidence scores and provenance reasons for every decision.

### The Four-Tier Categorization Priority Chain

The `infer_category` function implements a cascading logic:

1. **Explicit front-matter** (confidence 1.0)
2. **Folder-based parent** (confidence 0.95)
3. **Keyword matching** (confidence up to 0.92)
4. **Dynamic inference** (confidence 0.42‑0.20)

Each tier is evaluated only if the previous tier returns no result.

### Explicit Front-Matter Categories

When a [`SKILL.md`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/SKILL.md) file contains a `category:` field, `normalize_category` converts it to lower-case kebab-case. This yields confidence **1.0** with the reason `frontmatter:category`.

### Folder-Based Parent Categories

If no explicit category exists, the system uses the parent directory name (e.g., `skills/web-engineering/...`). This produces confidence **0.95** with the reason `path:folder`.

### Keyword-Based Inference

The `infer_category` function (lines 41‑51) searches concatenated skill text (ID, name, description, body) against `CATEGORY_KEYWORDS`. Scoring rules:

- Exact word matches add **3 points**
- Longer substring matches add **1 point**

The best-scoring category receives confidence `min(0.92, 0.45 + 0.05*score)` with reason `keyword-match:term1,term2`.

### Dynamic ID-Based Inference

If keyword matching fails, `infer_dynamic_category` (lines 90‑114) tokenizes the skill ID:

- **Prefix-based** rules (e.g., `aws-lambda`) yield confidence **0.42**
- **Last token** fallback yields confidence **0.34**
- Ultimate fallback to `"general"` yields confidence **0.20**

Reasons are recorded as `derived-from-id-prefix:...`, `derived-from-id-suffix:...`, or `default:general`.

### Normalizing Category Names

The `normalize_category` helper (lines 76‑87) ensures consistent identifiers:

```python
def normalize_category(value):
    if value is None: return None
    text = str(value).strip().lower()
    text = text.replace("_", "-")
    text = re.sub(r"\s+", "-", text)
    text = re.sub(r"[^a-z0-9-]", "", text)
    text = re.sub(r"-+", "-", text).strip("-")
    return text or None

```

This converts any input to clean kebab-case for uniform indexing.

## End-to-End Pipeline Flow

The complete internal skill filtering and categorization pipeline operates as follows:

1. **Discovery** – `SkillScanner` walks the `skills/` tree, reading each [`SKILL.md`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/SKILL.md) front-matter via `discover_all()`.

2. **Optional Filter** – `run_audit` applies `skill_filter` (if supplied via `--skill`) using the list comprehension filter, or calls `discover_skill()` for direct lookups.

3. **Index Generation** – `generate_index` parses each skill, merges metadata, and invokes `infer_category` to produce a final `category`, confidence score, and provenance reason.

4. **Result** – The JSON index ([`skills_index.json`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/skills_index.json)) contains structured entries such as:

```json
{
  "id": "web-search",
  "name": "Web Search",
  "category": "ai-ml",
  "category_confidence": 0.92,
  "category_reason": "keyword-match:search,web,api"
}

```

## Practical Code Examples

### Filtering Skills Programmatically

To filter skills by name using the same logic as the Sentinel audit:

```python
from scanner import SkillScanner

scanner = SkillScanner()
all_skills = scanner.discover_all()                 # → full list

print(f"Total skills discovered: {len(all_skills)}")

# Manual filter (same logic used by run_audit)

name_to_find = "instagram"
filtered = [s for s in all_skills if s["name"] == name_to_find]
print(filtered)   # empty list → not found, or one dict → found

```

*Relevant code:* `SkillScanner.discover_all` ([`scanner.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/scanner.py) lines 53‑61) and the filtering snippet in [`run_audit.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/run_audit.py) (lines 96‑99).

### Running a Targeted Audit via CLI

Execute the Sentinel audit for a single skill without processing the entire repository:

```bash
python -m skills.skill-sentinel.scripts.run_audit --skill instagram --format json

```

The CLI translates `--skill instagram` into the `skill_filter` argument, invoking the same list-comprehension filter shown above. The command returns a JSON payload with `run_id`, `overall_score`, and a single `snapshots` entry for the *instagram* skill.

*Relevant code:* CLI parsing in [`run_audit.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/run_audit.py) (lines 30‑35) and the `run_audit` call (lines 44‑51).

### Generating and Inspecting the Skill Index

To generate the skill index and inspect inferred categories with confidence scores:

```python
from skill_categorization.tools.scripts.generate_index import generate_index

skills_dir = "skills"
output_file = "skills_index.json"
generate_index(skills_dir, output_file)

# Load and print a few inferred categories

import json
with open(output_file) as f:
    index = json.load(f)

for entry in index[:5]:                     # first 5 skills

    print(entry["name"], "→", entry["category"],
          f"(confidence {entry['category_confidence']})")

```

Running this script shows categories derived from front-matter, folder names, or keyword matches, together with the numeric confidence and the reason string (e.g., `keyword-match:react,frontend`).

*Relevant code:* `generate_index` orchestrates parsing and categorization (lines 98‑124), calls `infer_category` (line 49) for each skill.

## Summary

- **Skill filtering** occurs in-process via the `skill_filter` parameter in [`run_audit.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/run_audit.py), using a list comprehension to match skill names after `SkillScanner` completes filesystem discovery.
- **Direct discovery** bypasses full enumeration through `SkillScanner.discover_skill(name)`, optimizing single-skill lookups.
- **Categorization** follows a strict four-tier priority: explicit front-matter (confidence 1.0) → parent folder name (0.95) → keyword matching (up to 0.92) → dynamic ID parsing (0.42‑0.20).
- **Provenance tracking** records the decision reason (e.g., `keyword-match:search,web`) and confidence score in the final [`skills_index.json`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/skills_index.json) output.
- **Normalization** ensures consistent kebab-case category identifiers via `normalize_category`, handling spaces, underscores, and special characters.

## Frequently Asked Questions

### How does the skill_filter parameter work in the Sentinel audit workflow?

The `skill_filter` parameter is an optional string passed via the `--skill` CLI argument in [`run_audit.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/run_audit.py). After `SkillScanner.discover_all()` builds the complete list of skills from the filesystem, the system applies a list comprehension—`[s for s in all_skills if s["name"] == skill_filter]`—to retain only matching entries. If no match is found, the workflow returns an error dictionary indicating the skill was not found, preventing unnecessary processing of unrelated capabilities.

### What is the priority order for skill categorization in the index generator?

The categorization logic in [`generate_index.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/generate_index.py) follows a deterministic four-tier priority chain. First, it checks for an explicit `category` field in the SKILL.md front-matter (confidence 1.0). If absent, it uses the parent directory name (confidence 0.95). If neither exists, it scans the skill text for keywords defined in `CATEGORY_KEYWORDS`, scoring exact matches higher (confidence up to 0.92). Finally, it falls back to dynamic inference from the skill ID itself, using prefix rules or the last token (confidence 0.42‑0.20).

### How does the system handle category name normalization?

All category values pass through the `normalize_category` function defined in [`generate_index.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/generate_index.py) (lines 76‑87). This helper converts inputs to lower-case, replaces underscores with hyphens, collapses whitespace and multiple hyphens into single delimiters, and strips non-alphanumeric characters except hyphens. The result is a consistent **kebab-case** identifier (e.g., `"AI_ML"` becomes `"ai-ml"`, `"Web Engineering"` becomes `"web-engineering"`) that ensures uniform indexing and filtering across the repository.

### Can I retrieve a single skill without loading the entire skill list?

Yes. While `SkillScanner.discover_all()` enumerates every skill in the `skills/` directory, the class provides the `discover_skill(name)` helper (lines 63‑66 in [`scanner.py`](https://github.com/sickn33/antigravity-awesome-skills/blob/main/scanner.py)) for targeted lookups. This method bypasses the full enumeration by directly locating the specific skill record, optimizing performance when you need to audit or categorize a single capability rather than the entire repository.