# What Is the `graph.py` File in SkillSpector? DAG Orchestration and Pipeline Execution

> Discover the graph.py file in SkillSpector. It orchestrates analysis nodes and executes pipelines, transforming input into SARIF reports using a DAG engine and topological sorting.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: internals
- Published: 2026-07-12

---

**The [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) file in SkillSpector implements a lightweight directed acyclic graph (DAG) engine that orchestrates analysis nodes, manages dependencies via topological sorting, and executes the pipeline that transforms raw input into final SARIF reports.**

The [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) file in NVIDIA/SkillSpector serves as the central nervous system of the codebase, coordinating how individual analysis components interact to produce skill assessments. Located at [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py), this module implements the **AnalysisGraph** class—a flexible execution framework that allows static analyzers, LLM-based meta-analyzers, and reporting modules to run in a deterministic, dependency-respecting order.

## Core Responsibilities of [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py)

According to the SkillSpector source code, [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) handles seven critical responsibilities that turn independent analysis components into a cohesive pipeline:

- **Define the graph model**: Implements a lightweight DAG where each node represents a processing step (e.g., input resolution, static analysis, LLM-based meta-analysis, deduplication, report generation).

- **Register nodes and edges**: Provides an API (`add_node`, `add_edge`) that other modules use to declare dependencies. For example, the *resolve-input* node must run before any analyzer nodes, and the *deduplicate* node runs after all findings have been collected.

- **Topological ordering**: Performs a topological sort to compute a safe execution order that respects all declared dependencies, ensuring each node runs only after its upstream nodes have produced their output.

- **Execution engine**: Offers a `run` method that walks the sorted node list, invokes each node's `process` method, and propagates a shared *context* (a mutable dictionary) that carries intermediate results.

- **Error handling and short-circuiting**: Catches exceptions from individual nodes, records them in the context, and can abort the graph early if a critical failure occurs (e.g., input validation error).

- **Result aggregation**: After the graph finishes, it collects the final payload (usually a SARIF report or a synthesized skill-assessment) from the designated sink node(s) and returns it to the CLI or API caller.

- **Extensibility**: Because the graph is built at runtime, new analysis nodes (static pattern checkers, custom LLM prompts, etc.) can be plugged in without modifying the central execution flow.

## The DAG Architecture in SkillSpector

The **AnalysisGraph** class treats the analysis pipeline as a directed acyclic graph where nodes represent discrete processing units and edges represent data dependencies. This architecture ensures that the *resolve-input* node (defined in [`src/skillspector/nodes/resolve_input.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py)) executes before any analyzer nodes, and that the *report* node (defined in [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)) runs only after all findings have been collected and deduplicated.

The system uses a shared **context** dictionary—passed between nodes—to maintain state across the pipeline. This context carries parsed source files, raw findings from static analysis, LLM responses, and intermediate metadata, allowing each node to access the outputs of its dependencies without tight coupling.

## Key API Methods: Building and Running Graphs

### Registering Nodes with `add_node`

The `add_node` method registers a processing unit with a unique identifier. Each node must implement a `process` method that accepts and mutates the shared context.

### Declaring Dependencies with `add_edge`

The `add_edge` method creates directed connections between nodes, enforcing that the source node completes before the target node begins. This dependency declaration enables the topological sort to determine the correct execution sequence.

### Executing the Pipeline with `run`

The `run` method triggers the execution sequence:
1. Computes topological ordering of all registered nodes
2. Iterates through the sorted list
3. Invokes each node's `process` method with the current context
4. Returns the final aggregated results

## Error Handling and Short-Circuiting

The execution engine in [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) implements robust error isolation. When a node raises an exception, the graph catches it, records the error in the context, and evaluates whether to continue or abort. Critical failures—such as input validation errors in the resolve-input phase—trigger early termination, preventing downstream nodes from running on invalid data.

## Practical Usage Examples

### Building a Simple Analysis Graph

```python
from skillspector.graph import AnalysisGraph
from skillspector.nodes.resolve_input import ResolveInputNode
from skillspector.nodes.analyzers.static_patterns import StaticPatternsNode
from skillspector.nodes.report import ReportNode

# Create the graph

g = AnalysisGraph()

# Register nodes

g.add_node("resolve", ResolveInputNode())
g.add_node("static", StaticPatternsNode())
g.add_node("report", ReportNode())

# Declare dependencies

g.add_edge("resolve", "static")   # static analysis needs resolved input

g.add_edge("static", "report")   # report needs findings from static analysis

# Run the graph

result = g.run()
print(result)   # -> final SARIF or JSON report

```

### Extending the Graph with a Custom LLM Analyzer

```python
from skillspector.graph import AnalysisGraph
from my_custom.llm_analyzer import MyLLMAnalyzerNode

graph = AnalysisGraph()
graph.add_node("llm", MyLLMAnalyzerNode())
graph.add_edge("resolve", "llm")   # LLM needs resolved source code

graph.add_edge("llm", "report")    # LLM output feeds the final report

graph.run()

```

These examples demonstrate how the graph API enables you to assemble any combination of nodes in a clear, dependency-driven way without modifying the core execution logic in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py).

## Integration with the SkillSpector Ecosystem

The [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) module does not operate in isolation. It coordinates with several key components:

- **[`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py)**: The CLI wrapper constructs the default graph configuration and invokes `graph.run()` to initiate analysis.
- **[`src/skillspector/nodes/resolve_input.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py)**: The entry point node that parses command-line arguments or API payloads into the normalized context.
- **`src/skillspector/nodes/analyzers/*.py`**: Various analyzer implementations (static patterns, YARA rules, LLM-based meta-analysis) that plug into the graph as processing nodes.
- **[`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)**: The sink node that aggregates findings from all upstream analyzers and emits the final SARIF or JSON report.

This modular architecture allows new analyzers to be added to the `analyzers` directory and registered in the graph without changing the orchestration logic.

## Summary

- The [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) file in SkillSpector implements a **DAG-based execution engine** that coordinates analysis pipelines through the `AnalysisGraph` class.
- It provides **`add_node`**, **`add_edge`**, and **`run`** methods to build and execute dependency graphs with automatic topological sorting.
- A **shared context dictionary** propagates intermediate results between nodes, enabling loose coupling between analysis components.
- The system supports **runtime extensibility**, allowing custom analyzers and LLM-based nodes to integrate without modifying core orchestration code.
- Error handling includes **short-circuiting capabilities** to abort the pipeline on critical failures while logging node-specific exceptions.

## Frequently Asked Questions

### What graph library does SkillSpector use?

SkillSpector does not rely on external graph libraries like NetworkX. Instead, [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) implements a lightweight, custom DAG abstraction specifically tailored for the analysis pipeline, keeping dependencies minimal and allowing fine-grained control over execution semantics.

### How does [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) handle node failures?

When a node raises an exception, the execution engine catches it, records the error in the shared context, and evaluates whether the failure is critical. Non-critical errors allow the graph to continue, while critical failures (such as invalid input resolution) trigger early termination to prevent wasted computation on downstream nodes.

### Can I add custom analyzers to the SkillSpector graph?

Yes. The graph supports runtime extensibility through the `add_node` method. You can implement custom analyzer classes with a `process` method that accepts the context dictionary, then register them in the graph using `add_node` and connect them with `add_edge` to declare dependencies on existing nodes like `resolve` or `report`.

### What output format does the graph produce?

The graph itself is agnostic to output formats, but the standard `ReportNode` (located in [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)) aggregates findings into **SARIF** (Static Analysis Results Interchange Format) or JSON reports. The final output is returned by `graph.run()` and consumed by the CLI or API layer.