How to Create Custom Correlation Rules in SpiderFoot Using YAML

SpiderFoot discovers relationships between data elements by applying correlation rules written in YAML files located in the correlations/ directory, which are automatically loaded at startup and executed through a four-stage pipeline.

SpiderFoot's correlation engine enables security analysts to detect meaningful patterns across reconnaissance data without writing Python code. By creating custom correlation rules in YAML, you can define domain-specific logic that groups, filters, and flags findings based on your threat model. This guide explains the rule structure and implementation process using the actual source code from the smicallef/spiderfoot repository.

Understanding the Correlation Rule Architecture

How Rules Are Loaded and Executed

When SpiderFoot initializes, the function SpiderFootHelpers.loadCorrelationRulesRaw() in [spiderfoot/helpers.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/helpers.py) traverses the correlations/ directory and reads every *.yaml file into an internal ruleset dictionary.

The execution engine in [spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) processes each rule through four sequential stages:

  1. collections – Gather matching records from the SpiderFoot database
  2. aggregation (optional) – Bucket records by a common field
  3. analysis (optional) – Apply conditional thresholds or counts
  4. headline – Generate the final correlation title with placeholders

Successful correlations are persisted via SpiderFootDB.correlationResultCreate() in [spiderfoot/db.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/db.py), storing results in tbl_scan_correlation_results and linked events in tbl_scan_correlation_results_events.

YAML Rule Structure Reference

Each correlation rule consists of five logical sections:

Section Required Purpose
meta Yes Human-readable name, description, and risk level
collections Yes Data-gathering blocks with method chains
aggregation No Grouping field for pre-analysis bucketing
analysis No Conditional logic (thresholds, counts)
headline Yes Title template with field placeholders

Available Collection Methods

The engine supports multiple matching strategies in collect blocks:

  • exact – Field matches value precisely
  • regex – Regular expression match against field
  • exists – Field has any value
  • contains – Field contains substring

Field references can target source.field, child.field, or entity.field depending on data context.

Step-by-Step: Creating Your First Custom Rule

Step 1: Copy the Template

Duplicate correlations/template.yaml and name your file to match the rule's id. The filename without extension must equal the id field inside the YAML.

cd /path/to/spiderfoot
cp correlations/template.yaml correlations/admin_interface.yaml

Step 2: Define Rule Metadata

Edit the meta section with clear, actionable descriptions:

id: admin_interface
version: 1
meta:
  name: Exposed administrative interfaces
  description: >
    Detects hosts serving login panels or admin dashboards
    on non-standard ports, indicating potential attack surface.
  risk: HIGH

Risk levels: HIGH, MEDIUM, LOW, INFO.

Step 3: Configure Data Collections

Build one or more collect blocks with ordered method lists:

collections:
  - collect:
      - method: exact
        field: type
        value: TCP_PORT_OPEN
      - method: regex
        field: data
        value:
          - :80\b
          - :443\b
      - method: regex
        field: module
        value:
          - ^sfp_spider$

Multiple methods in a collect block act as AND conditions. Multiple collect blocks act as OR conditions.

Step 4: Add Aggregation (Optional)

Group records before analysis to detect patterns across multiple findings:

aggregation:
  field: source.data   # bucket by parent entity

Common aggregation fields: data, source.data, module.

Step 5: Add Analysis Conditions (Optional)

Trigger correlations only when thresholds are met:

analysis:
  - method: threshold
    minimum: 5
    field: data

Available analysis methods:

  • threshold – Trigger when bucket contains ≥ minimum items (or ≤ maximum)
  • count – Exact match to specified count

Step 6: Craft the Headline Template

Use placeholders matching your collected fields:

headline: "Host {source.data} has {count} open web ports"

Common placeholders: {data}, {module}, {type}, {source.data}, {count}.

Step 7: Validate and Restart

Validate YAML syntax to prevent startup failures:

yamllint correlations/admin_interface.yaml

Restart SpiderFoot to load the new rule. The rule appears in the Correlation Rules list via the scancorrelationrules API endpoint.

Complete YAML Examples

Example 1: Admin and Login Sub-domains

Detects hosts containing both "admin" and "login" sub-domains, indicating exposed administrative interfaces.


# File: correlations/admin_login.yaml

id: admin_login
version: 1
meta:
  name: Hosts with admin and login sub-domains
  description: >
    Finds Internet names that have both an "admin" and a "login" sub-domain.
    This often indicates a publicly exposed administrative interface.
  risk: HIGH

collections:
  - collect:
      - method: exact
        field: type
        value: INTERNET_NAME
      - method: regex
        field: data
        value:
          - ^admin\..*$
      - method: regex
        field: data
        value:
          - ^login\..*$

aggregation:
  field: data

analysis:
  - method: threshold
    minimum: 2
    field: data

headline: "Host {data} exposes both admin and login sub-domains"

This rule uses two regex matches against the same field to require both patterns in the collected dataset, then aggregates by hostname and requires at least two matches.

Example 2: IP Addresses in Multiple WHOIS Records

Flags infrastructure appearing across multiple WHOIS registrations, suggesting shared or suspicious hosting.

id: ip_multiple_whois
version: 1
meta:
  name: IP seen in multiple WHOIS records
  description: >
    An IP address that shows up in several WHOIS entries may belong to
    a shared service or be part of a malicious infrastructure.
  risk: MEDIUM

collections:
  - collect:
      - method: exact
        field: type
        value: IPV4_ADDRESS
      - method: exists
        field: module
        value: sfp_whois

aggregation:
  field: data

analysis:
  - method: threshold
    minimum: 3
    field: data

headline: "IP {data} appears in ≥3 WHOIS records"

The exists method here ensures WHOIS data was actually retrieved by the sfp_whois module.

Key Source Files and Their Roles

File Function
[spiderfoot/helpers.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/helpers.py) loadCorrelationRulesRaw() – loads all YAML files from correlations/
[spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) Parses, validates, and executes the rule pipeline; handles SyntaxError for malformed YAML
[spiderfoot/db.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/db.py) correlationResultCreate() – persists results to database tables
[correlations/template.yaml](https://github.com/smicallef/spiderfoot/blob/master/correlations/template.yaml) Canonical reference for all supported sections and keys
[correlations/README.md](https://github.com/smicallef/spiderfoot/blob/master/correlations/README.md) Field naming conventions and method documentation

Troubleshooting Custom Correlation Rules

Startup Errors

Malformed YAML triggers a clear SyntaxError from [spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py). Check:

  • Consistent indentation (spaces, not tabs)
  • Valid YAML list syntax for value: fields with multiple items
  • Matching id and filename

Rules Not Triggering

  • Verify field names match actual SpiderFoot event types
  • Check that aggregation and analysis fields exist in collected data
  • Test regex patterns independently—anchors (^, $) are often required

Performance Considerations

  • Aggregation reduces memory pressure for high-volume matches
  • Specific exact matches before regex improves collection speed
  • Multiple collect blocks increase query load—consolidate where possible

Summary

Frequently Asked Questions

What happens if my YAML file has syntax errors?

SpiderFoot will abort startup with a SyntaxError that identifies the problematic file and line. The error originates in [spiderfoot/correlation.py](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) during rule parsing. Validate your YAML with yamllint before deployment.

Can I reference fields from different event types in the same rule?

Yes. The collection engine allows field references from source.field (the parent entity), child.field (downstream elements), and entity.field (the originating target). These prefixes enable cross-event correlation logic.

How do I test a correlation rule without running a full scan?

SpiderFoot does not provide an isolated rule tester. Create a minimal scan with known data that should trigger your rule, then verify results appear under the Correlations tab. Check the scancorrelationrules API endpoint to confirm your rule loaded correctly.

Do correlation rules affect scan performance?

Complex rules with multiple collect blocks, unanchored regex patterns, or missing aggregation can increase database load. Optimize by using exact matches before regex, adding aggregation for high-cardinality fields, and limiting collect blocks to necessary conditions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →