What Is the `graph.py` File in SkillSpector? DAG Orchestration and Pipeline Execution
The graph.py file in SkillSpector implements a lightweight directed acyclic graph (DAG) engine that orchestrates analysis nodes, manages dependencies via topological sorting, and executes the pipeline that transforms raw input into final SARIF reports.
The graph.py file in NVIDIA/SkillSpector serves as the central nervous system of the codebase, coordinating how individual analysis components interact to produce skill assessments. Located at src/skillspector/graph.py, this module implements the AnalysisGraph class—a flexible execution framework that allows static analyzers, LLM-based meta-analyzers, and reporting modules to run in a deterministic, dependency-respecting order.
Core Responsibilities of graph.py
According to the SkillSpector source code, graph.py handles seven critical responsibilities that turn independent analysis components into a cohesive pipeline:
-
Define the graph model: Implements a lightweight DAG where each node represents a processing step (e.g., input resolution, static analysis, LLM-based meta-analysis, deduplication, report generation).
-
Register nodes and edges: Provides an API (
add_node,add_edge) that other modules use to declare dependencies. For example, the resolve-input node must run before any analyzer nodes, and the deduplicate node runs after all findings have been collected. -
Topological ordering: Performs a topological sort to compute a safe execution order that respects all declared dependencies, ensuring each node runs only after its upstream nodes have produced their output.
-
Execution engine: Offers a
runmethod that walks the sorted node list, invokes each node'sprocessmethod, and propagates a shared context (a mutable dictionary) that carries intermediate results. -
Error handling and short-circuiting: Catches exceptions from individual nodes, records them in the context, and can abort the graph early if a critical failure occurs (e.g., input validation error).
-
Result aggregation: After the graph finishes, it collects the final payload (usually a SARIF report or a synthesized skill-assessment) from the designated sink node(s) and returns it to the CLI or API caller.
-
Extensibility: Because the graph is built at runtime, new analysis nodes (static pattern checkers, custom LLM prompts, etc.) can be plugged in without modifying the central execution flow.
The DAG Architecture in SkillSpector
The AnalysisGraph class treats the analysis pipeline as a directed acyclic graph where nodes represent discrete processing units and edges represent data dependencies. This architecture ensures that the resolve-input node (defined in src/skillspector/nodes/resolve_input.py) executes before any analyzer nodes, and that the report node (defined in src/skillspector/nodes/report.py) runs only after all findings have been collected and deduplicated.
The system uses a shared context dictionary—passed between nodes—to maintain state across the pipeline. This context carries parsed source files, raw findings from static analysis, LLM responses, and intermediate metadata, allowing each node to access the outputs of its dependencies without tight coupling.
Key API Methods: Building and Running Graphs
Registering Nodes with add_node
The add_node method registers a processing unit with a unique identifier. Each node must implement a process method that accepts and mutates the shared context.
Declaring Dependencies with add_edge
The add_edge method creates directed connections between nodes, enforcing that the source node completes before the target node begins. This dependency declaration enables the topological sort to determine the correct execution sequence.
Executing the Pipeline with run
The run method triggers the execution sequence:
- Computes topological ordering of all registered nodes
- Iterates through the sorted list
- Invokes each node's
processmethod with the current context - Returns the final aggregated results
Error Handling and Short-Circuiting
The execution engine in graph.py implements robust error isolation. When a node raises an exception, the graph catches it, records the error in the context, and evaluates whether to continue or abort. Critical failures—such as input validation errors in the resolve-input phase—trigger early termination, preventing downstream nodes from running on invalid data.
Practical Usage Examples
Building a Simple Analysis Graph
from skillspector.graph import AnalysisGraph
from skillspector.nodes.resolve_input import ResolveInputNode
from skillspector.nodes.analyzers.static_patterns import StaticPatternsNode
from skillspector.nodes.report import ReportNode
# Create the graph
g = AnalysisGraph()
# Register nodes
g.add_node("resolve", ResolveInputNode())
g.add_node("static", StaticPatternsNode())
g.add_node("report", ReportNode())
# Declare dependencies
g.add_edge("resolve", "static") # static analysis needs resolved input
g.add_edge("static", "report") # report needs findings from static analysis
# Run the graph
result = g.run()
print(result) # -> final SARIF or JSON report
Extending the Graph with a Custom LLM Analyzer
from skillspector.graph import AnalysisGraph
from my_custom.llm_analyzer import MyLLMAnalyzerNode
graph = AnalysisGraph()
graph.add_node("llm", MyLLMAnalyzerNode())
graph.add_edge("resolve", "llm") # LLM needs resolved source code
graph.add_edge("llm", "report") # LLM output feeds the final report
graph.run()
These examples demonstrate how the graph API enables you to assemble any combination of nodes in a clear, dependency-driven way without modifying the core execution logic in src/skillspector/graph.py.
Integration with the SkillSpector Ecosystem
The graph.py module does not operate in isolation. It coordinates with several key components:
src/skillspector/cli.py: The CLI wrapper constructs the default graph configuration and invokesgraph.run()to initiate analysis.src/skillspector/nodes/resolve_input.py: The entry point node that parses command-line arguments or API payloads into the normalized context.src/skillspector/nodes/analyzers/*.py: Various analyzer implementations (static patterns, YARA rules, LLM-based meta-analysis) that plug into the graph as processing nodes.src/skillspector/nodes/report.py: The sink node that aggregates findings from all upstream analyzers and emits the final SARIF or JSON report.
This modular architecture allows new analyzers to be added to the analyzers directory and registered in the graph without changing the orchestration logic.
Summary
- The
graph.pyfile in SkillSpector implements a DAG-based execution engine that coordinates analysis pipelines through theAnalysisGraphclass. - It provides
add_node,add_edge, andrunmethods to build and execute dependency graphs with automatic topological sorting. - A shared context dictionary propagates intermediate results between nodes, enabling loose coupling between analysis components.
- The system supports runtime extensibility, allowing custom analyzers and LLM-based nodes to integrate without modifying core orchestration code.
- Error handling includes short-circuiting capabilities to abort the pipeline on critical failures while logging node-specific exceptions.
Frequently Asked Questions
What graph library does SkillSpector use?
SkillSpector does not rely on external graph libraries like NetworkX. Instead, src/skillspector/graph.py implements a lightweight, custom DAG abstraction specifically tailored for the analysis pipeline, keeping dependencies minimal and allowing fine-grained control over execution semantics.
How does graph.py handle node failures?
When a node raises an exception, the execution engine catches it, records the error in the shared context, and evaluates whether the failure is critical. Non-critical errors allow the graph to continue, while critical failures (such as invalid input resolution) trigger early termination to prevent wasted computation on downstream nodes.
Can I add custom analyzers to the SkillSpector graph?
Yes. The graph supports runtime extensibility through the add_node method. You can implement custom analyzer classes with a process method that accepts the context dictionary, then register them in the graph using add_node and connect them with add_edge to declare dependencies on existing nodes like resolve or report.
What output format does the graph produce?
The graph itself is agnostic to output formats, but the standard ReportNode (located in src/skillspector/nodes/report.py) aggregates findings into SARIF (Static Analysis Results Interchange Format) or JSON reports. The final output is returned by graph.run() and consumed by the CLI or API layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →