How SpiderFoot’s YAML‑Based Correlation Engine Processes Scan Results

SpiderFoot's correlation engine transforms declarative YAML rules into actionable intelligence by loading rule definitions, querying scan result databases, enriching event data, and applying aggregation and analysis methods to identify patterns.

The SpiderFoot open-source OSINT platform uses a flexible, YAML-driven correlation system to make sense of scan results. Instead of hardcoding detection logic, the engine reads external rule files that define what data to collect, how to group it, and which analysis techniques to apply. This article walks through the complete pipeline implemented in spiderfoot/correlation.py on the master branch, with precise source references and runnable examples.


Initializing the SpiderFootCorrelator Class

The correlation workflow begins with the SpiderFootCorrelator class in spiderfoot/correlation.py. Its __init__ method (lines 49–77) receives three critical inputs: a SpiderFootDb database handle, the raw YAML ruleset, and an optional scanId to scope operations to a specific scan.

The constructor performs type validation, loads the complete list of event types from the database, and builds a type_entity_map that maps event types to their corresponding entity classifications. This foundation ensures all subsequent operations have access to both the rule logic and the database schema context needed for proper event enrichment.


Loading and Validating YAML Correlation Rules

Parsing Raw YAML

Rules are ingested through a parsing phase (lines 80–95) that uses yaml.safe_load on each rule definition. The original YAML string is preserved under the rawYaml key, and newline characters are stripped from metadata fields to ensure clean display. This preservation allows the engine to store the exact rule definition alongside correlation results for auditability.

Syntax Validation

Before any rule executes, check_ruleset_validity and check_rule_validity (lines 64–76, 84–102) enforce structural requirements:

  • Required fields: meta, collections, and headline must be present
  • Supported methods: All field comparisons must use recognized operators

Invalid rules raise SyntaxError immediately, preventing malformed definitions from producing silent failures during scan processing.


Executing Correlations with run_correlations

The run_correlations method (lines 108–131) serves as the main execution entry point. It performs a critical safety check: fetching the scan instance and aborting if the scan status indicates it is still running. Once confirmed complete, the method iterates over every validated rule, delegating individual processing to process_rule.

This design separates orchestration from implementation—run_correlations handles lifecycle management while process_rule contains the domain-specific correlation logic.


Determining Data Requirements with analyze_rule_scope

Before querying the database, the engine must know which related data to fetch. The analyze_rule_scope method (lines 90–108) inspects three rule components:

  • Collections: What event fields to match against
  • Aggregation: How events will be grouped
  • Analysis: What statistical or pattern methods to apply

From this inspection, the engine sets boolean flags—fetchChildren, fetchSources, and fetchEntities—that control whether the database query should retrieve related child events, source attribution chains, or entity classifications. This optimization prevents expensive joins when the correlation logic does not require them.


Collecting and Enriching Events from the Database

Building Database Queries

The collect_from_db method (lines 40–87) constructs query criteria through build_db_criteria, supporting matches on:

  • Event type: The classification of the data (e.g., IP_ADDRESS, EMAILADDR)
  • Module: The SpiderFoot plugin that produced the event
  • Data: The actual content value (exact or regex matching)

These criteria feed into SpiderFootDb.scanResultEvent to retrieve matching rows from the scan result database.

Event Enrichment Pipeline

Raw database rows undergo three enrichment phases defined in spiderfoot/correlation.py:

Sources — enrich_event_sources (lines 30–45): Loads the originating elements for each event, establishing provenance.

Children — enrich_event_child (lines 46–55): Loads descendant events, building the forward relationship graph.

Entities — enrich_event_entities (lines 84–104): Walks the source chain recursively until reaching an entity of type ENTITY or INTERNAL, collapsing complex attribution paths into manageable identifiers.


Filtering Collections and Aggregating Events

Refining Collection Results

When a correlation rule defines multiple collection stages, the first performs the initial database query while subsequent collections call refine_collection (lines 63–84). This method filters the working event set by applying field-level comparisons using either exact string matching or regex pattern matching. Events that fail any refinement stage are removed from consideration.

Bucket Aggregation

Rules with an aggregation block trigger aggregate_events (lines 34–77), which groups events by values extracted via event_extract. The aggregation system handles nested field references such as child.ip, automatically stripping values from non-matching sub-elements to maintain clean bucket boundaries.


Analyzing Event Buckets

The analyze_events dispatcher (lines 90–106) routes each bucket to the appropriate analysis method based on the rule's analysis configuration:

Method Function Purpose
threshold analysis_threshold (lines 108–176) Counts unique values, flags buckets exceeding or falling below defined limits
outlier analysis_outlier (lines 176–227) Removes statistically unusual buckets (abnormally small or large)
first_collection_only analysis_first_collection_only (lines 227–267) Restricts analysis to events from the initial collection stage
match_all_to_first_collection analysis_match_all_to_first_collection (lines 267–307) Requires all subsequent collections to align with first-collection values

Each analysis method may delete buckets that fail to meet criteria, progressively narrowing the result set to high-confidence correlations.


Generating Titles and Persisting Results

Human-Readable Correlation Titles

The build_correlation_title method (lines 111–127) substitutes placeholder tokens like {field} in the rule's headline template with actual values extracted from the first event in each surviving bucket. This produces context-specific titles such as "Multiple hosts share IP 192.0.2.1" rather than generic labels.

Database Persistence

Finally, create_correlation (lines 132–158) persists results through SpiderFootDb.correlationResultCreate, recording:

  • The rule identifier
  • Rule metadata
  • The raw YAML definition for audit trails
  • The generated human-readable title
  • The complete list of involved event IDs

This persistence enables the SpiderFoot web interface to display correlations, link to source events, and correlate across multiple scans.


Practical Code Examples

Running Correlations for a Completed Scan

import yaml
from spiderfoot.correlation import SpiderFootCorrelator
from spiderfoot.db import SpiderFootDb

# Load YAML rule definitions

with open("example_rules.yaml") as f:
    raw_rules = yaml.safe_load(f)

# Initialize database connection

db = SpiderFootDb()
db.open()

# Create correlator scoped to specific scan

correlator = SpiderFootCorrelator(db, raw_rules, scanId="2023-08-15-abcdef")

# Execute all correlation rules

correlator.run_correlations()

Inspecting Rule Output Programmatically


# Access first rule and process individually

rule = correlator.get_ruleset()[0]
buckets = correlator.process_rule(rule)

# Examine aggregation results

for bucket, events in buckets.items():
    print(f"Bucket {bucket!r}: {len(events)} events")
    for event in events[:3]:  # Sample first 3

        print(f"  - {event.get('type')}: {event.get('data')}")

Key Source Files

File Purpose
spiderfoot/correlation.py Core YAML correlation engine—rule parsing, event collection, enrichment, aggregation, analysis, and persistence
spiderfoot/db.py Database abstraction layer providing SpiderFootDb for scan result queries and correlation storage
spiderfoot/event.py Event object definitions used throughout correlation processing
modules/ OSINT plugin directory whose outputs populate the database that correlations analyze

Summary

SpiderFoot's YAML-based correlation engine processes scan results through a twelve-stage pipeline:

  • Initialization: SpiderFootCorrelator.__init__ establishes database connections and entity mappings
  • Rule loading: YAML parsing with yaml.safe_load and syntax validation via check_rule_validity
  • Scope analysis: analyze_rule_scope optimizes data fetching requirements
  • Data collection: collect_from_db and build_db_criteria query scan results with flexible matching
  • Enrichment: Source, child, and entity resolution provides full attribution context
  • Filtering: refine_collection applies multi-stage field-level filtering
  • Aggregation: aggregate_events groups related events into analysis buckets
  • Analysis: Statistical methods (threshold, outlier, collection constraints) identify significant patterns
  • Output generation: Dynamic title construction and database persistence via create_correlation

The declarative YAML approach allows security researchers to add new correlation capabilities without modifying Python code—simply defining new rules that the existing engine executes.


Frequently Asked Questions

What YAML fields are required for a valid SpiderFoot correlation rule?

Every correlation rule must contain three top-level keys: meta (containing descriptive metadata), collections (defining what data to gather and how to filter it), and headline (a template for the human-readable correlation title). The check_rule_validity function in spiderfoot/correlation.py enforces these requirements and raises SyntaxError for non-compliant definitions.

Can correlation rules match events using regular expressions?

Yes. The refine_collection method supports both exact and regex matching methods for field comparisons. When processing collection rules, the engine applies these methods via refine_collection (lines 63–84) to filter events against specified field values using Python's regex engine.

How does the correlation engine prevent analyzing incomplete scans?

The run_correlations method (lines 108–131) explicitly checks scan status before execution. It fetches the scan instance and aborts if the scan is still running, ensuring correlations only operate on complete datasets where all plugin modules have finished contributing results.

What analysis methods can YAML correlation rules specify?

The engine supports four analysis methods as implemented in spiderfoot/correlation.py: threshold for counting unique values against limits, outlier for statistical bucket filtering, first_collection_only for restricting analysis to initial collection events, and match_all_to_first_collection for requiring alignment across collection stages.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →