How to Create and Test Custom Correlation Rules in SpiderFoot

Create custom correlation rules in SpiderFoot by writing YAML files in the correlations/ directory, loading them via SpiderFootHelpers.loadCorrelationRulesRaw(), and executing them through SpiderFootCorrelator.run_correlations()—then verify results in the database or via unit tests.

SpiderFoot's correlation engine enables security analysts to define custom rules that combine data from multiple reconnaissance modules, aggregate findings, and surface meaningful security insights. This guide walks through the complete workflow for authoring, loading, executing, and testing custom correlation rules using the actual implementation from the smicallef/spiderfoot repository.

How SpiderFoot's Correlation Engine Works

The correlation system consists of three integrated components working together to process rules against scan data.

Component File Location Core Responsibility
Rule loader spiderfoot/helpers.py lines 174-185 Reads and parses raw YAML rule files from disk
Correlator engine spiderfoot/correlation.py lines 49-63, 108-131 Validates rule syntax, executes collection queries, applies aggregation and analysis logic
Result writer spiderfoot/correlation.py lines 897-906, 929-447 Generates human-readable titles and persists correlations to the database

Understanding this pipeline helps you debug rules that fail validation or produce unexpected results.

Step 1: Author a YAML Correlation Rule

SpiderFoot correlation rules use YAML syntax with a strict schema. Create your rule file in the repository's correlations/ directory with a filename matching the rule's id field.

Required YAML Structure

Section Purpose Required Keys
meta Human-readable metadata name, description, risk
collections Database queries to gather events collect blocks with method, field, value
aggregation Event grouping (optional) field to bucket by
analysis Thresholds and filters (optional) method, field, plus method-specific parameters
headline Result title template Text with {field} placeholders

Example: Duplicate Hostname Detection

This minimal rule detects hostnames appearing multiple times across scan results—adapted from the template in correlations/template.yaml:


# file: correlations/duplicate_host.yaml

id: duplicate_host
version: 1
meta:
  name: Duplicate Hostnames
  description: >
    Finds hostnames that appear more than once in the scan results,
    indicating potential infrastructure overlap or DNS inconsistencies.
  risk: MEDIUM
collections:
  - collect:
      - method: exact
        field: type
        value: INTERNET_NAME
analysis:
  - method: threshold
    minimum: 2        # trigger when count ≥ 2

    field: data
headline: "Duplicate hostname detected: {data}"

Critical requirement: The id field must exactly match the filename without the .yaml extension. Mismatches cause the correlator to skip the rule during validation in SpiderFootCorrelator.__init__().

Step 2: Load Rules into the Correlator

Use the helper utilities to load your custom rules before execution.

from spiderfoot.helpers import SpiderFootHelpers
from spiderfoot.correlation import SpiderFootCorrelator
from spiderfoot.db import SpiderFootDb

# 1️⃣ Load all YAML files from the correlations directory

raw_rules = SpiderFootHelpers.loadCorrelationRulesRaw(
    path="./correlations/",          # directory containing *.yaml files

    ignore_files=["template.yaml"]   # optional: skip specific files

)

# 2️⃣ Initialize database connection

db = SpiderFootDb("./spiderfoot.db")  # or existing DB handle from a scan

# 3️⃣ Instantiate correlator with loaded rules

correlator = SpiderFootCorrelator(
    dbh=db,
    ruleset=raw_rules,
    scanId="your-scan-id-here"
)

The loadCorrelationRulesRaw() function at lines 174-185 in spiderfoot/helpers.py returns a dictionary mapping rule IDs to raw YAML strings. The SpiderFootCorrelator constructor at lines 49-63 in spiderfoot/correlation.py then parses each YAML document, validates schema compliance, and builds executable rule objects.

Step 3: Execute Correlation Rules

Run all loaded rules against the scan data with a single method call:


# Execute every validated rule against the scan database

correlator.run_correlations()

The run_correlations() method (lines 108-131 in spiderfoot/correlation.py) iterates through self.rules and invokes process_rule() for each entry. Failed rules log warnings without halting execution—check logs for validation errors if results appear incomplete.

Step 4: Verify Correlation Results

Correlations matching your criteria are stored in tbl_scan_correlation_results. Query them programmatically or inspect via SpiderFoot's web interface.


# Retrieve all correlation titles for verification

summary = db.correlationSummary(scanId="your-scan-id-here")
for correlation in summary:
    print(f"Detected: {correlation['title']}")

# Access full correlation details

details = db.correlationResultGet(correlationId="specific-id")

The build_correlation_title() method (lines 897-906) substitutes {field} placeholders in your headline template with actual event data. Then create_correlation() (lines 929-447) persists the completed result to the database with proper risk scoring and metadata.

Step 5: Unit Test Custom Correlation Rules

SpiderFoot includes a test suite at test/unit/spiderfoot/test_spiderfootcorrelator.py. Follow this pattern to validate your rules programmatically.

Complete Test Example

import pytest
from spiderfoot.db import SpiderFootDb
from spiderfoot.correlation import SpiderFootCorrelator

def test_duplicate_host_correlation():
    """Verify duplicate hostname detection triggers correctly."""
    
    # 1️⃣ Create isolated in-memory database

    db = SpiderFootDb(":memory:")
    db.scanInstanceCreate(
        instanceId="test-scan-001",
        scanName="Unit Test Scan",
        scanner="pytest",
        moduleList="sfp_dnsdumpster"
    )
    
    # 2️⃣ Insert mock events matching collection criteria

    db.scanResultEventCreate(
        instanceId="test-scan-001",
        eventType="INTERNET_NAME",
        data="vulnerable.example.com",
        module="sfp_dnsdumpster"
    )
    db.scanResultEventCreate(
        instanceId="test-scan-001",
        eventType="INTERNET_NAME",
        data="vulnerable.example.com",  # duplicate entry

        module="sfp_shodan"
    )
    
    # 3️⃣ Load and execute rule

    raw_rules = {
        "duplicate_host": open("correlations/duplicate_host.yaml").read()
    }
    correlator = SpiderFootCorrelator(
        dbh=db,
        ruleset=raw_rules,
        scanId="test-scan-001"
    )
    correlator.run_correlations()
    
    # 4️⃣ Assert expected outcome

    results = db.correlationSummary(scanId="test-scan-001")
    assert len(results) == 1
    assert "vulnerable.example.com" in results[0]["title"]
    assert results[0]["risk"] == "MEDIUM"

Testing Best Practices

  • Use :memory: SQLite databases for fast, isolated tests
  • Include edge cases: zero matches, single matches, threshold boundaries
  • Verify both positive matches (rule triggers) and negative matches (rule correctly skips)
  • Test aggregation and analysis methods independently when logic grows complex

Advanced Correlation Rule Patterns

Aggregation with Field Grouping

Group events before analysis to detect patterns across data subsets:

collections:
  - collect:
      - method: exact
        field: type
        value: IP_ADDRESS
aggregation:
  field: scanId          # group by source scan

analysis:
  - method: threshold
    minimum: 5           # 5+ IPs per scanId triggers

    field: data

Multiple Collection Blocks

Combine data from different event types in a single rule:

collections:
  - collect:
      - method: exact
        field: type
        value: VULNERABILITY_CVE
  - collect:
      - method: regex
        field: type
        value: "^VULNERABILITY_.*"

Summary

  • Create correlation rules as YAML files in correlations/ with id matching filename
  • Load rules via SpiderFootHelpers.loadCorrelationRulesRaw() from spiderfoot/helpers.py
  • Validate rule syntax automatically in SpiderFootCorrelator.__init__() at spiderfoot/correlation.py lines 49-63
  • Execute with run_correlations() which processes each rule through process_rule()
  • Verify results in tbl_scan_correlation_results using db.correlationSummary() or direct queries
  • Test rules using in-memory databases with controlled mock events following the pattern in test_spiderfootcorrelator.py

Frequently Asked Questions

What file naming convention must correlation rules follow?

The YAML filename must exactly match the rule's id field with .yaml appended. A rule with id: suspicious_cert must reside in correlations/suspicious_cert.yaml. The SpiderFootCorrelator constructor validates this match during initialization and logs warnings for discrepancies.

How do I debug a correlation rule that produces no results?

First, verify the rule passes validation by checking SpiderFoot logs for YAML syntax errors. Next, confirm your collections blocks match actual event types in your scan data—query the database directly with db.scanResultEvent() to see available types. Finally, test with lowered thresholds (e.g., minimum: 1) to ensure the collection logic works before tightening constraints.

Can correlation rules reference custom event types from private modules?

Yes. The correlator operates on the database layer and does not validate event types against SpiderFoot's built-in lists. Any string value in collections.collect[].value will match corresponding entries in tbl_scan_results. Document your custom types clearly to avoid maintenance issues.

What analysis methods are available beyond threshold?

SpiderFoot's spiderfoot/correlation.py implements multiple analysis methods including outlier for statistical anomaly detection, match for pattern-based filtering, and compare for cross-field relationships. Inspect the analysis_methods dictionary in the correlator source for the complete, version-specific list and their required parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →