What the build_context Node Gathers in NVIDIA SkillSpector
The build_context node in NVIDIA SkillSpector generates a comprehensive ScanContext dictionary containing file inventories, content caches, executable script detection, and parsed YAML manifest data from SKILL.md.
The build_context node serves as the data foundation for the entire SkillSpector analysis pipeline. Located in src/skillspector/nodes/build_context.py, this node executes at the start of the workflow to transform raw skill directories into structured metadata that downstream static scanners and LLM-based policies consume to evaluate skill safety and compliance.
The Eight Context Keys Produced by build_context
When invoked via build_context(state), the node returns a flat dictionary containing eight specific keys that describe every aspect of the local skill directory.
File Inventory and Content Caching
components: A sorted list of relative file paths discovered under the skill directory. The_walk_skill_filesfunction (lines 78-101) generates these paths as POSIX strings while excluding Git repositories, virtual environments,node_modules, and hidden files via the internal_SKIP_DIRSset (lines 35-38).file_cache: A dictionary mapping each relative path to its UTF-8 encoded contents (with invalid character replacement), populated by_read_file_cacheat lines 54-67.
Metadata and Security Indicators
component_metadata: A list of dictionaries generated by_build_component_metadata(lines 21-52), where each entry contains the file's relative path, inferred type (e.g.,python,markdown), line count, size in bytes, and whether the file extension indicates an executable script.has_executable_scripts: A boolean flag set toTrueif any component matches executable extensions such as.py,.sh, or.js, enabling security policies to quickly identify potentially dangerous code (lines 33-38).
Skill Manifest and Configuration
manifest: Structured data extracted fromSKILL.mdorskill.mdfront-matter using_parse_manifest(lines 70-98), including the skill's name, description, triggers, required permissions, and parameters.previous_manifest: Initialized asNone(lines 39-41) to support comparison logic when re-scanning existing skills.model_config: The global LLM configuration dictionary imported fromskillspector.constants.MODEL_CONFIG(lines 29-32) and injected unchanged into the context (lines 41-44), containing model names, temperature, and token limits.
AST Placeholder
ast_cache: An empty dictionary{}reserved for downstream analysis nodes to populate with abstract syntax tree representations (lines 38-40).
Input Validation and File Walking
Before gathering data, _resolve_skill_dir (lines 64-75) validates that the provided skill_path exists and is a directory. The node then walks the filesystem using _walk_skill_files, which applies aggressive filtering to exclude version control directories, Python virtual environments, Node.js modules, and hidden files according to the _SKIP_DIRS definition.
Practical Usage Examples
Generating a ScanContext
from skillspector.state import SkillspectorState
from skillspector.nodes.build_context import build_context
# Initialise a state with the path to the skill we want to scan
state = SkillspectorState()
state["skill_path"] = "/path/to/my_skill"
# Build the scan context
scan_context = build_context(state)
print("Found components:", len(scan_context["components"]))
print("Executable scripts present:", scan_context["has_executable_scripts"])
print("Manifest:", scan_context["manifest"])
Accessing Component Metadata
metadata = scan_context["component_metadata"]
for comp in metadata:
print(f"{comp['path']:40} {comp['type']:12} "
f"{comp['lines']:4} lines "
f"{'exec' if comp['executable'] else 'non-exec'}")
This prints a concise table of each file’s type, line count, and executable status as detected during the initial scan.
Summary
- The
build_contextnode insrc/skillspector/nodes/build_context.pycreates a flat metadata dictionary calledScanContextthat powers the entire SkillSpector pipeline. - It captures file inventories (
components), content caches (file_cache), and security indicators (has_executable_scripts,component_metadata). - The node extracts structured skill manifests from
SKILL.mdYAML front-matter and exposes global model configurations for consistent LLM usage downstream. - All file walking excludes development artifacts like Git directories and virtual environments via the
_SKIP_DIRSfilter set.
Frequently Asked Questions
Where is the build_context node implemented in SkillSpector?
The core implementation resides in src/skillspector/nodes/build_context.py, specifically the build_context(state) function that assembles the dictionary at lines 21-45 according to the source code.
What files does build_context exclude when scanning a skill directory?
The _walk_skill_files function respects the _SKIP_DIRS set defined at lines 35-38, which excludes .git repositories, Python virtual environments, node_modules, and hidden files to ensure only relevant skill code is analyzed.
How does build_context extract skill metadata and permissions?
The _parse_manifest function (lines 70-98) locates SKILL.md or skill.md and parses its YAML front-matter to populate the manifest key with the skill's name, description, triggers, and required permissions.
Why is ast_cache empty when build_context finishes?
The ast_cache key is intentionally initialized as an empty dictionary {} at lines 38-40 to serve as a placeholder that downstream AST analysis nodes populate after the initial context build, keeping the build_context node focused on file system metadata only.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →