How SkillSpector Aggregates Findings from Different Analyzers: The LangGraph State Pattern
SkillSpector uses LangGraph's operator.add annotation on a shared findings list to automatically concatenate results from multiple static analysis nodes into a single aggregated state before meta-analysis.
NVIDIA's SkillSpector coordinates multiple security analyzers through a LangGraph workflow that runs as a unified pipeline. The system aggregates findings from different analyzers by leveraging a shared SkillspectorState object, where the findings field uses Python's operator.add reduction mechanism to merge outputs automatically. This design eliminates manual collection logic and ensures that by the time execution reaches the meta-analysis phase, all detector results exist in a single consolidated list.
The SkillspectorState: A Shared Aggregation Container
The core of the aggregation mechanism lives in the state definition. Rather than manually merging lists after each analyzer runs, SkillSpector declares the findings field with a special annotation that tells LangGraph how to combine values when multiple nodes write to the same key.
State Definition with Automatic Concatenation
In src/skillspector/state.py, the findings field is declared using Annotated[list[Finding], operator.add]. This configuration instructs LangGraph to treat the list as a reducible state that grows by concatenation:
# src/skillspector/state.py
findings: Annotated[list[Finding], operator.add]
When any analyzer node returns a dictionary containing a findings key, LangGraph automatically appends those items to the existing list rather than replacing it. This pattern ensures thread-safe aggregation without requiring explicit synchronization code in the analyzer implementations.
Analyzer Nodes and Graph Wiring
Each security analyzer operates as an independent node in the LangGraph workflow. The system supports extensible registration of analyzers while maintaining a consistent interface for contribution to the shared state.
Analyzer Registry
All available analyzers are registered in src/skillspector/nodes/analyzers/__init__.py within the ANALYZER_NODES dictionary. Each entry maps a node identifier to a function that accepts the current state and returns a list of findings:
# src/skillspector/nodes/analyzers/__init__.py
ANALYZER_NODES = {
"static_patterns_prompt_injection": static_patterns_prompt_injection_node,
...
}
Every node function returns a typed dictionary containing a findings list, which LangGraph immediately merges into the global state using the configured operator.add behavior.
Parallel Execution in the Graph
The workflow definition in src/skillspector/graph.py wires each analyzer to execute after the context-building phase. The graph adds edges from build_context to every analyzer ID, then from each analyzer to the meta_analyzer:
# src/skillspector/graph.py
for analyzer_id in ANALYZER_NODE_IDS:
workflow.add_edge("build_context", analyzer_id)
workflow.add_edge(analyzer_id, "meta_analyzer")
This structure allows analyzers to run in parallel (where dependencies permit) while ensuring all complete before the meta-analysis phase begins.
How LangGraph Handles the Aggregation
The aggregation occurs implicitly through LangGraph's state management. As each analyzer node executes, it produces results like:
return {"findings": [Finding(rule_id="SEC-001", severity="high", ...)]}
Because the findings field uses operator.add, LangGraph concatenates this new list with existing findings rather than overwriting the state. By the time the meta_analyzer node receives control, the SkillspectorState contains the union of all findings produced by every analyzer in the pipeline.
Meta-Analysis and Filtering
The meta_analyzer node consumes the aggregated findings and applies additional LLM-based enrichment or filtering. Located in src/skillspector/nodes/meta_analyzer.py, this node reads the complete list and produces a refined output:
# src/skillspector/nodes/meta_analyzer.py
findings: list[Finding] = state.get("findings", [])
...
return {"filtered_findings": filtered}
The filtered_findings field represents the final processed collection after deduplication, severity adjustments, or false-positive removal that the LLM coordinator performs on the aggregated data.
Practical Examples
Running the Full Pipeline
To execute the complete workflow and automatically aggregate findings from all registered analyzers:
# Scan a skill directory – the workflow will invoke all analyzers and
# combine their findings into a single report.
skillspector scan ./my-skill/
Manual Invocation for Subset Analysis
You can manually invoke specific analyzers and observe the aggregation behavior:
from skillspector.graph import graph
from skillspector.state import SkillspectorState
# Initialise state with a simple input path
state: SkillspectorState = {"input_path": "./my-skill/", "mode": "scan"}
# Run only the static‑pattern analyzers
for analyzer_id in [
"static_patterns_prompt_injection",
"static_patterns_data_exfiltration",
]:
node = graph.nodes[analyzer_id]
result = node(state)
# The node returns {"findings": [...]}; LangGraph adds them to state.findings
state.update(result)
print("Aggregated findings:", len(state["findings"]))
for f in state["findings"]:
print(f.rule_id, f.severity, f.message)
Inspecting Filtered Results
After the complete graph execution, access the final filtered set:
# After the full graph execution
final_state = graph.invoke(state) # runs the compiled workflow
filtered = final_state["filtered_findings"]
print(f"Filtered down to {len(filtered)} findings after LLM review")
Summary
- LangGraph state reduction: The
findingsfield inSkillspectorStateusesAnnotated[list[Finding], operator.add]to automatically merge outputs from multiple analyzers. - Parallel analyzer execution: The graph in
src/skillspector/graph.pywires all analyzers to receive input frombuild_contextand feed intometa_analyzer. - Implicit aggregation: Analyzer nodes simply return
{"findings": [...]}dictionaries; LangGraph handles concatenation without explicit merge code. - Meta-analysis layer: The
meta_analyzernode consumes the complete aggregated list and producesfiltered_findingsafter LLM review. - Extensible design: New analyzers register in
ANALYZER_NODESand automatically participate in the aggregation pipeline.
Frequently Asked Questions
Does SkillSpector run analyzers sequentially or in parallel?
SkillSpector leverages LangGraph's parallel execution capabilities where graph topology permits. Since all analyzers in ANALYZER_NODE_IDS receive edges from build_context without dependencies on each other, LangGraph can execute them concurrently. The operator.add annotation ensures thread-safe aggregation of their findings into the shared state list.
What happens if two analyzers report the same vulnerability?
The raw aggregation preserves all findings, including potential duplicates. The meta_analyzer node in src/skillspector/nodes/meta_analyzer.py receives the complete aggregated list and can apply deduplication logic, typically using the LLM coordinator to identify and remove duplicate findings before writing the final filtered_findings to state.
Can I add custom analyzers to the aggregation pipeline?
Yes. Custom analyzers must follow the node function signature used in src/skillspector/nodes/analyzers/__init__.py: accept a SkillspectorState and return a dictionary with a findings key containing list[Finding]. Register the new node in ANALYZER_NODES, and the graph wiring in src/skillspector/graph.py will automatically include it in the aggregation workflow.
How does the meta_analyzer modify the aggregated findings?
The meta_analyzer reads the aggregated findings list from state, optionally enriches each finding with LLM-generated context or severity adjustments, applies filtering rules to remove false positives, and writes the processed results to a new state key called filtered_findings. This preserves the original aggregated data while providing a refined dataset for final reporting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →