How Internal Skill Filtering and Categorization Works in Antigravity Awesome Skills
The Antigravity Awesome Skills repository implements a two-stage pipeline where SkillScanner first discovers skills from the filesystem and optionally filters them by name via the skill_filter parameter, then generate_index categorizes each skill through a deterministic four-tier priority chain (explicit front-matter → folder name → keyword matching → dynamic ID parsing) with confidence scoring and provenance tracking.
The sickn33/antigravity-awesome-skills repository manages modular capabilities through a sophisticated internal skill filtering and categorization system. This architecture separates discovery and filtering concerns from classification logic, enabling both targeted audits of individual skills and automated organization of the entire skill library. Understanding these mechanisms is essential for contributors maintaining the index or running the Sentinel audit workflow against specific capabilities.
Skill Filtering in the Sentinel Audit Workflow
The Sentinel audit workflow limits processing scope through deterministic in-process filtering that operates after initial filesystem discovery. This ensures that targeted audits never perform redundant disk operations.
The SkillScanner Discovery Phase
The SkillScanner class in skills/skill-sentinel/scripts/scanner.py (lines 53‑61) walks the skills/ tree and reads each SKILL.md front-matter. The discover_all() method returns a complete list of skill dictionaries containing metadata such as name, id, and description.
Applying the skill_filter Parameter
In skills/skill-sentinel/scripts/run_audit.py (lines 96‑100), the optional skill_filter parameter (passed from the CLI via --skill) applies a list comprehension to the discovered set:
if skill_filter:
all_skills = [s for s in all_skills if s["name"] == skill_filter]
if not all_skills:
return {"error": f"Skill '{skill_filter}' nao encontrada."}
This filtering stage removes every entry whose name does not match the filter, returning an error if no matches exist.
Direct Skill Discovery with discover_skill
For single-skill lookups, SkillScanner offers the discover_skill(name) helper (lines 63‑66 in scanner.py). This method bypasses full enumeration by returning a single skill record directly, optimizing performance when only one target is required.
Skill Categorization and Index Generation
Categorization occurs during index generation via skill_categorization/tools/scripts/generate_index.py. The system assigns categories through a deterministic priority chain that records confidence scores and provenance reasons for every decision.
The Four-Tier Categorization Priority Chain
The infer_category function implements a cascading logic:
- Explicit front-matter (confidence 1.0)
- Folder-based parent (confidence 0.95)
- Keyword matching (confidence up to 0.92)
- Dynamic inference (confidence 0.42‑0.20)
Each tier is evaluated only if the previous tier returns no result.
Explicit Front-Matter Categories
When a SKILL.md file contains a category: field, normalize_category converts it to lower-case kebab-case. This yields confidence 1.0 with the reason frontmatter:category.
Folder-Based Parent Categories
If no explicit category exists, the system uses the parent directory name (e.g., skills/web-engineering/...). This produces confidence 0.95 with the reason path:folder.
Keyword-Based Inference
The infer_category function (lines 41‑51) searches concatenated skill text (ID, name, description, body) against CATEGORY_KEYWORDS. Scoring rules:
- Exact word matches add 3 points
- Longer substring matches add 1 point
The best-scoring category receives confidence min(0.92, 0.45 + 0.05*score) with reason keyword-match:term1,term2.
Dynamic ID-Based Inference
If keyword matching fails, infer_dynamic_category (lines 90‑114) tokenizes the skill ID:
- Prefix-based rules (e.g.,
aws-lambda) yield confidence 0.42 - Last token fallback yields confidence 0.34
- Ultimate fallback to
"general"yields confidence 0.20
Reasons are recorded as derived-from-id-prefix:..., derived-from-id-suffix:..., or default:general.
Normalizing Category Names
The normalize_category helper (lines 76‑87) ensures consistent identifiers:
def normalize_category(value):
if value is None: return None
text = str(value).strip().lower()
text = text.replace("_", "-")
text = re.sub(r"\s+", "-", text)
text = re.sub(r"[^a-z0-9-]", "", text)
text = re.sub(r"-+", "-", text).strip("-")
return text or None
This converts any input to clean kebab-case for uniform indexing.
End-to-End Pipeline Flow
The complete internal skill filtering and categorization pipeline operates as follows:
-
Discovery –
SkillScannerwalks theskills/tree, reading eachSKILL.mdfront-matter viadiscover_all(). -
Optional Filter –
run_auditappliesskill_filter(if supplied via--skill) using the list comprehension filter, or callsdiscover_skill()for direct lookups. -
Index Generation –
generate_indexparses each skill, merges metadata, and invokesinfer_categoryto produce a finalcategory, confidence score, and provenance reason. -
Result – The JSON index (
skills_index.json) contains structured entries such as:
{
"id": "web-search",
"name": "Web Search",
"category": "ai-ml",
"category_confidence": 0.92,
"category_reason": "keyword-match:search,web,api"
}
Practical Code Examples
Filtering Skills Programmatically
To filter skills by name using the same logic as the Sentinel audit:
from scanner import SkillScanner
scanner = SkillScanner()
all_skills = scanner.discover_all() # → full list
print(f"Total skills discovered: {len(all_skills)}")
# Manual filter (same logic used by run_audit)
name_to_find = "instagram"
filtered = [s for s in all_skills if s["name"] == name_to_find]
print(filtered) # empty list → not found, or one dict → found
Relevant code: SkillScanner.discover_all (scanner.py lines 53‑61) and the filtering snippet in run_audit.py (lines 96‑99).
Running a Targeted Audit via CLI
Execute the Sentinel audit for a single skill without processing the entire repository:
python -m skills.skill-sentinel.scripts.run_audit --skill instagram --format json
The CLI translates --skill instagram into the skill_filter argument, invoking the same list-comprehension filter shown above. The command returns a JSON payload with run_id, overall_score, and a single snapshots entry for the instagram skill.
Relevant code: CLI parsing in run_audit.py (lines 30‑35) and the run_audit call (lines 44‑51).
Generating and Inspecting the Skill Index
To generate the skill index and inspect inferred categories with confidence scores:
from skill_categorization.tools.scripts.generate_index import generate_index
skills_dir = "skills"
output_file = "skills_index.json"
generate_index(skills_dir, output_file)
# Load and print a few inferred categories
import json
with open(output_file) as f:
index = json.load(f)
for entry in index[:5]: # first 5 skills
print(entry["name"], "→", entry["category"],
f"(confidence {entry['category_confidence']})")
Running this script shows categories derived from front-matter, folder names, or keyword matches, together with the numeric confidence and the reason string (e.g., keyword-match:react,frontend).
Relevant code: generate_index orchestrates parsing and categorization (lines 98‑124), calls infer_category (line 49) for each skill.
Summary
- Skill filtering occurs in-process via the
skill_filterparameter inrun_audit.py, using a list comprehension to match skill names afterSkillScannercompletes filesystem discovery. - Direct discovery bypasses full enumeration through
SkillScanner.discover_skill(name), optimizing single-skill lookups. - Categorization follows a strict four-tier priority: explicit front-matter (confidence 1.0) → parent folder name (0.95) → keyword matching (up to 0.92) → dynamic ID parsing (0.42‑0.20).
- Provenance tracking records the decision reason (e.g.,
keyword-match:search,web) and confidence score in the finalskills_index.jsonoutput. - Normalization ensures consistent kebab-case category identifiers via
normalize_category, handling spaces, underscores, and special characters.
Frequently Asked Questions
How does the skill_filter parameter work in the Sentinel audit workflow?
The skill_filter parameter is an optional string passed via the --skill CLI argument in run_audit.py. After SkillScanner.discover_all() builds the complete list of skills from the filesystem, the system applies a list comprehension—[s for s in all_skills if s["name"] == skill_filter]—to retain only matching entries. If no match is found, the workflow returns an error dictionary indicating the skill was not found, preventing unnecessary processing of unrelated capabilities.
What is the priority order for skill categorization in the index generator?
The categorization logic in generate_index.py follows a deterministic four-tier priority chain. First, it checks for an explicit category field in the SKILL.md front-matter (confidence 1.0). If absent, it uses the parent directory name (confidence 0.95). If neither exists, it scans the skill text for keywords defined in CATEGORY_KEYWORDS, scoring exact matches higher (confidence up to 0.92). Finally, it falls back to dynamic inference from the skill ID itself, using prefix rules or the last token (confidence 0.42‑0.20).
How does the system handle category name normalization?
All category values pass through the normalize_category function defined in generate_index.py (lines 76‑87). This helper converts inputs to lower-case, replaces underscores with hyphens, collapses whitespace and multiple hyphens into single delimiters, and strips non-alphanumeric characters except hyphens. The result is a consistent kebab-case identifier (e.g., "AI_ML" becomes "ai-ml", "Web Engineering" becomes "web-engineering") that ensures uniform indexing and filtering across the repository.
Can I retrieve a single skill without loading the entire skill list?
Yes. While SkillScanner.discover_all() enumerates every skill in the skills/ directory, the class provides the discover_skill(name) helper (lines 63‑66 in scanner.py) for targeted lookups. This method bypasses the full enumeration by directly locating the specific skill record, optimizing performance when you need to audit or categorize a single capability rather than the entire repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →