How to Create and Test Custom Correlation Rules in SpiderFoot
Create custom correlation rules in SpiderFoot by writing YAML files in the correlations/ directory, loading them via SpiderFootHelpers.loadCorrelationRulesRaw(), and executing them through SpiderFootCorrelator.run_correlations()—then verify results in the database or via unit tests.
SpiderFoot's correlation engine enables security analysts to define custom rules that combine data from multiple reconnaissance modules, aggregate findings, and surface meaningful security insights. This guide walks through the complete workflow for authoring, loading, executing, and testing custom correlation rules using the actual implementation from the smicallef/spiderfoot repository.
How SpiderFoot's Correlation Engine Works
The correlation system consists of three integrated components working together to process rules against scan data.
| Component | File Location | Core Responsibility |
|---|---|---|
| Rule loader | spiderfoot/helpers.py lines 174-185 |
Reads and parses raw YAML rule files from disk |
| Correlator engine | spiderfoot/correlation.py lines 49-63, 108-131 |
Validates rule syntax, executes collection queries, applies aggregation and analysis logic |
| Result writer | spiderfoot/correlation.py lines 897-906, 929-447 |
Generates human-readable titles and persists correlations to the database |
Understanding this pipeline helps you debug rules that fail validation or produce unexpected results.
Step 1: Author a YAML Correlation Rule
SpiderFoot correlation rules use YAML syntax with a strict schema. Create your rule file in the repository's correlations/ directory with a filename matching the rule's id field.
Required YAML Structure
| Section | Purpose | Required Keys |
|---|---|---|
meta |
Human-readable metadata | name, description, risk |
collections |
Database queries to gather events | collect blocks with method, field, value |
aggregation |
Event grouping (optional) | field to bucket by |
analysis |
Thresholds and filters (optional) | method, field, plus method-specific parameters |
headline |
Result title template | Text with {field} placeholders |
Example: Duplicate Hostname Detection
This minimal rule detects hostnames appearing multiple times across scan results—adapted from the template in correlations/template.yaml:
# file: correlations/duplicate_host.yaml
id: duplicate_host
version: 1
meta:
name: Duplicate Hostnames
description: >
Finds hostnames that appear more than once in the scan results,
indicating potential infrastructure overlap or DNS inconsistencies.
risk: MEDIUM
collections:
- collect:
- method: exact
field: type
value: INTERNET_NAME
analysis:
- method: threshold
minimum: 2 # trigger when count ≥ 2
field: data
headline: "Duplicate hostname detected: {data}"
Critical requirement: The id field must exactly match the filename without the .yaml extension. Mismatches cause the correlator to skip the rule during validation in SpiderFootCorrelator.__init__().
Step 2: Load Rules into the Correlator
Use the helper utilities to load your custom rules before execution.
from spiderfoot.helpers import SpiderFootHelpers
from spiderfoot.correlation import SpiderFootCorrelator
from spiderfoot.db import SpiderFootDb
# 1️⃣ Load all YAML files from the correlations directory
raw_rules = SpiderFootHelpers.loadCorrelationRulesRaw(
path="./correlations/", # directory containing *.yaml files
ignore_files=["template.yaml"] # optional: skip specific files
)
# 2️⃣ Initialize database connection
db = SpiderFootDb("./spiderfoot.db") # or existing DB handle from a scan
# 3️⃣ Instantiate correlator with loaded rules
correlator = SpiderFootCorrelator(
dbh=db,
ruleset=raw_rules,
scanId="your-scan-id-here"
)
The loadCorrelationRulesRaw() function at lines 174-185 in spiderfoot/helpers.py returns a dictionary mapping rule IDs to raw YAML strings. The SpiderFootCorrelator constructor at lines 49-63 in spiderfoot/correlation.py then parses each YAML document, validates schema compliance, and builds executable rule objects.
Step 3: Execute Correlation Rules
Run all loaded rules against the scan data with a single method call:
# Execute every validated rule against the scan database
correlator.run_correlations()
The run_correlations() method (lines 108-131 in spiderfoot/correlation.py) iterates through self.rules and invokes process_rule() for each entry. Failed rules log warnings without halting execution—check logs for validation errors if results appear incomplete.
Step 4: Verify Correlation Results
Correlations matching your criteria are stored in tbl_scan_correlation_results. Query them programmatically or inspect via SpiderFoot's web interface.
# Retrieve all correlation titles for verification
summary = db.correlationSummary(scanId="your-scan-id-here")
for correlation in summary:
print(f"Detected: {correlation['title']}")
# Access full correlation details
details = db.correlationResultGet(correlationId="specific-id")
The build_correlation_title() method (lines 897-906) substitutes {field} placeholders in your headline template with actual event data. Then create_correlation() (lines 929-447) persists the completed result to the database with proper risk scoring and metadata.
Step 5: Unit Test Custom Correlation Rules
SpiderFoot includes a test suite at test/unit/spiderfoot/test_spiderfootcorrelator.py. Follow this pattern to validate your rules programmatically.
Complete Test Example
import pytest
from spiderfoot.db import SpiderFootDb
from spiderfoot.correlation import SpiderFootCorrelator
def test_duplicate_host_correlation():
"""Verify duplicate hostname detection triggers correctly."""
# 1️⃣ Create isolated in-memory database
db = SpiderFootDb(":memory:")
db.scanInstanceCreate(
instanceId="test-scan-001",
scanName="Unit Test Scan",
scanner="pytest",
moduleList="sfp_dnsdumpster"
)
# 2️⃣ Insert mock events matching collection criteria
db.scanResultEventCreate(
instanceId="test-scan-001",
eventType="INTERNET_NAME",
data="vulnerable.example.com",
module="sfp_dnsdumpster"
)
db.scanResultEventCreate(
instanceId="test-scan-001",
eventType="INTERNET_NAME",
data="vulnerable.example.com", # duplicate entry
module="sfp_shodan"
)
# 3️⃣ Load and execute rule
raw_rules = {
"duplicate_host": open("correlations/duplicate_host.yaml").read()
}
correlator = SpiderFootCorrelator(
dbh=db,
ruleset=raw_rules,
scanId="test-scan-001"
)
correlator.run_correlations()
# 4️⃣ Assert expected outcome
results = db.correlationSummary(scanId="test-scan-001")
assert len(results) == 1
assert "vulnerable.example.com" in results[0]["title"]
assert results[0]["risk"] == "MEDIUM"
Testing Best Practices
- Use
:memory:SQLite databases for fast, isolated tests - Include edge cases: zero matches, single matches, threshold boundaries
- Verify both positive matches (rule triggers) and negative matches (rule correctly skips)
- Test aggregation and analysis methods independently when logic grows complex
Advanced Correlation Rule Patterns
Aggregation with Field Grouping
Group events before analysis to detect patterns across data subsets:
collections:
- collect:
- method: exact
field: type
value: IP_ADDRESS
aggregation:
field: scanId # group by source scan
analysis:
- method: threshold
minimum: 5 # 5+ IPs per scanId triggers
field: data
Multiple Collection Blocks
Combine data from different event types in a single rule:
collections:
- collect:
- method: exact
field: type
value: VULNERABILITY_CVE
- collect:
- method: regex
field: type
value: "^VULNERABILITY_.*"
Summary
- Create correlation rules as YAML files in
correlations/withidmatching filename - Load rules via
SpiderFootHelpers.loadCorrelationRulesRaw()fromspiderfoot/helpers.py - Validate rule syntax automatically in
SpiderFootCorrelator.__init__()atspiderfoot/correlation.pylines 49-63 - Execute with
run_correlations()which processes each rule throughprocess_rule() - Verify results in
tbl_scan_correlation_resultsusingdb.correlationSummary()or direct queries - Test rules using in-memory databases with controlled mock events following the pattern in
test_spiderfootcorrelator.py
Frequently Asked Questions
What file naming convention must correlation rules follow?
The YAML filename must exactly match the rule's id field with .yaml appended. A rule with id: suspicious_cert must reside in correlations/suspicious_cert.yaml. The SpiderFootCorrelator constructor validates this match during initialization and logs warnings for discrepancies.
How do I debug a correlation rule that produces no results?
First, verify the rule passes validation by checking SpiderFoot logs for YAML syntax errors. Next, confirm your collections blocks match actual event types in your scan data—query the database directly with db.scanResultEvent() to see available types. Finally, test with lowered thresholds (e.g., minimum: 1) to ensure the collection logic works before tightening constraints.
Can correlation rules reference custom event types from private modules?
Yes. The correlator operates on the database layer and does not validate event types against SpiderFoot's built-in lists. Any string value in collections.collect[].value will match corresponding entries in tbl_scan_results. Document your custom types clearly to avoid maintenance issues.
What analysis methods are available beyond threshold?
SpiderFoot's spiderfoot/correlation.py implements multiple analysis methods including outlier for statistical anomaly detection, match for pattern-based filtering, and compare for cross-field relationships. Inspect the analysis_methods dictionary in the correlator source for the complete, version-specific list and their required parameters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →