How to Create Custom Correlation Rules in SpiderFoot Using YAML
SpiderFoot discovers relationships between data elements by applying correlation rules written in YAML files located in the correlations/ directory, which are automatically loaded at startup and executed through a four-stage pipeline.
SpiderFoot's correlation engine enables security analysts to detect meaningful patterns across reconnaissance data without writing Python code. By creating custom correlation rules in YAML, you can define domain-specific logic that groups, filters, and flags findings based on your threat model. This guide explains the rule structure and implementation process using the actual source code from the smicallef/spiderfoot repository.
Understanding the Correlation Rule Architecture
How Rules Are Loaded and Executed
When SpiderFoot initializes, the function SpiderFootHelpers.loadCorrelationRulesRaw() in [spiderfoot/helpers.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/helpers.py) traverses the correlations/ directory and reads every *.yaml file into an internal ruleset dictionary.
The execution engine in [spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) processes each rule through four sequential stages:
- collections – Gather matching records from the SpiderFoot database
- aggregation (optional) – Bucket records by a common field
- analysis (optional) – Apply conditional thresholds or counts
- headline – Generate the final correlation title with placeholders
Successful correlations are persisted via SpiderFootDB.correlationResultCreate() in [spiderfoot/db.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/db.py), storing results in tbl_scan_correlation_results and linked events in tbl_scan_correlation_results_events.
YAML Rule Structure Reference
Each correlation rule consists of five logical sections:
| Section | Required | Purpose |
|---|---|---|
| meta | Yes | Human-readable name, description, and risk level |
| collections | Yes | Data-gathering blocks with method chains |
| aggregation | No | Grouping field for pre-analysis bucketing |
| analysis | No | Conditional logic (thresholds, counts) |
| headline | Yes | Title template with field placeholders |
Available Collection Methods
The engine supports multiple matching strategies in collect blocks:
exact– Field matches value preciselyregex– Regular expression match against fieldexists– Field has any valuecontains– Field contains substring
Field references can target source.field, child.field, or entity.field depending on data context.
Step-by-Step: Creating Your First Custom Rule
Step 1: Copy the Template
Duplicate correlations/template.yaml and name your file to match the rule's id. The filename without extension must equal the id field inside the YAML.
cd /path/to/spiderfoot
cp correlations/template.yaml correlations/admin_interface.yaml
Step 2: Define Rule Metadata
Edit the meta section with clear, actionable descriptions:
id: admin_interface
version: 1
meta:
name: Exposed administrative interfaces
description: >
Detects hosts serving login panels or admin dashboards
on non-standard ports, indicating potential attack surface.
risk: HIGH
Risk levels: HIGH, MEDIUM, LOW, INFO.
Step 3: Configure Data Collections
Build one or more collect blocks with ordered method lists:
collections:
- collect:
- method: exact
field: type
value: TCP_PORT_OPEN
- method: regex
field: data
value:
- :80\b
- :443\b
- method: regex
field: module
value:
- ^sfp_spider$
Multiple methods in a collect block act as AND conditions. Multiple collect blocks act as OR conditions.
Step 4: Add Aggregation (Optional)
Group records before analysis to detect patterns across multiple findings:
aggregation:
field: source.data # bucket by parent entity
Common aggregation fields: data, source.data, module.
Step 5: Add Analysis Conditions (Optional)
Trigger correlations only when thresholds are met:
analysis:
- method: threshold
minimum: 5
field: data
Available analysis methods:
threshold– Trigger when bucket contains ≥minimumitems (or ≤maximum)count– Exact match to specified count
Step 6: Craft the Headline Template
Use placeholders matching your collected fields:
headline: "Host {source.data} has {count} open web ports"
Common placeholders: {data}, {module}, {type}, {source.data}, {count}.
Step 7: Validate and Restart
Validate YAML syntax to prevent startup failures:
yamllint correlations/admin_interface.yaml
Restart SpiderFoot to load the new rule. The rule appears in the Correlation Rules list via the scancorrelationrules API endpoint.
Complete YAML Examples
Example 1: Admin and Login Sub-domains
Detects hosts containing both "admin" and "login" sub-domains, indicating exposed administrative interfaces.
# File: correlations/admin_login.yaml
id: admin_login
version: 1
meta:
name: Hosts with admin and login sub-domains
description: >
Finds Internet names that have both an "admin" and a "login" sub-domain.
This often indicates a publicly exposed administrative interface.
risk: HIGH
collections:
- collect:
- method: exact
field: type
value: INTERNET_NAME
- method: regex
field: data
value:
- ^admin\..*$
- method: regex
field: data
value:
- ^login\..*$
aggregation:
field: data
analysis:
- method: threshold
minimum: 2
field: data
headline: "Host {data} exposes both admin and login sub-domains"
This rule uses two regex matches against the same field to require both patterns in the collected dataset, then aggregates by hostname and requires at least two matches.
Example 2: IP Addresses in Multiple WHOIS Records
Flags infrastructure appearing across multiple WHOIS registrations, suggesting shared or suspicious hosting.
id: ip_multiple_whois
version: 1
meta:
name: IP seen in multiple WHOIS records
description: >
An IP address that shows up in several WHOIS entries may belong to
a shared service or be part of a malicious infrastructure.
risk: MEDIUM
collections:
- collect:
- method: exact
field: type
value: IPV4_ADDRESS
- method: exists
field: module
value: sfp_whois
aggregation:
field: data
analysis:
- method: threshold
minimum: 3
field: data
headline: "IP {data} appears in ≥3 WHOIS records"
The exists method here ensures WHOIS data was actually retrieved by the sfp_whois module.
Key Source Files and Their Roles
| File | Function |
|---|---|
[spiderfoot/helpers.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/helpers.py) |
loadCorrelationRulesRaw() – loads all YAML files from correlations/ |
[spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) |
Parses, validates, and executes the rule pipeline; handles SyntaxError for malformed YAML |
[spiderfoot/db.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/db.py) |
correlationResultCreate() – persists results to database tables |
[correlations/template.yaml](https://github.com/smicallef/spiderfoot/blob/master/correlations/template.yaml) |
Canonical reference for all supported sections and keys |
[correlations/README.md](https://github.com/smicallef/spiderfoot/blob/master/correlations/README.md) |
Field naming conventions and method documentation |
Troubleshooting Custom Correlation Rules
Startup Errors
Malformed YAML triggers a clear SyntaxError from [spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py). Check:
- Consistent indentation (spaces, not tabs)
- Valid YAML list syntax for
value:fields with multiple items - Matching
idand filename
Rules Not Triggering
- Verify field names match actual SpiderFoot event types
- Check that aggregation and analysis fields exist in collected data
- Test regex patterns independently—anchors (
^,$) are often required
Performance Considerations
- Aggregation reduces memory pressure for high-volume matches
- Specific
exactmatches beforeregeximproves collection speed - Multiple collect blocks increase query load—consolidate where possible
Summary
- SpiderFoot loads correlation rules from YAML files in the
correlations/directory at startup viaSpiderFootHelpers.loadCorrelationRulesRaw() - Rules execute through four stages: collections → aggregation → analysis → headline
- The meta section defines UI-visible name, description, and risk level (
HIGHthroughINFO) - Collections use chained methods (
exact,regex,exists,contains) to select database records - Aggregation buckets results by field; analysis applies threshold or count conditions
- The headline template generates human-readable correlation titles with field placeholders
- Rules are validated at load time—malformed YAML causes startup errors with clear messages
- Reference implementations exist in [
correlations/template.yaml](https://github.com/smicallef/spiderfoot/blob/master/correlations/template.yaml) and [correlations/README.md](https://github.com/smicallef/spiderfoot/blob/master/correlations/README.md)
Frequently Asked Questions
What happens if my YAML file has syntax errors?
SpiderFoot will abort startup with a SyntaxError that identifies the problematic file and line. The error originates in [spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) during rule parsing. Validate your YAML with yamllint before deployment.
Can I reference fields from different event types in the same rule?
Yes. The collection engine allows field references from source.field (the parent entity), child.field (downstream elements), and entity.field (the originating target). These prefixes enable cross-event correlation logic.
How do I test a correlation rule without running a full scan?
SpiderFoot does not provide an isolated rule tester. Create a minimal scan with known data that should trigger your rule, then verify results appear under the Correlations tab. Check the scancorrelationrules API endpoint to confirm your rule loaded correctly.
Do correlation rules affect scan performance?
Complex rules with multiple collect blocks, unanchored regex patterns, or missing aggregation can increase database load. Optimize by using exact matches before regex, adding aggregation for high-cardinality fields, and limiting collect blocks to necessary conditions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →