# How to Create Custom Correlation Rules in SpiderFoot Using YAML

> Learn to create custom correlation rules in SpiderFoot using YAML. Discover relationships between data elements with this straightforward guide.

- Repository: [Steve Micallef/spiderfoot](https://github.com/smicallef/spiderfoot)
- Tags: how-to-guide
- Published: 2026-08-15

---

**SpiderFoot discovers relationships between data elements by applying correlation rules written in YAML files located in the `correlations/` directory, which are automatically loaded at startup and executed through a four-stage pipeline.**

SpiderFoot's correlation engine enables security analysts to detect meaningful patterns across reconnaissance data without writing Python code. By creating custom correlation rules in YAML, you can define domain-specific logic that groups, filters, and flags findings based on your threat model. This guide explains the rule structure and implementation process using the actual source code from the [smicallef/spiderfoot](https://github.com/smicallef/spiderfoot) repository.

## Understanding the Correlation Rule Architecture

### How Rules Are Loaded and Executed

When SpiderFoot initializes, the function **`SpiderFootHelpers.loadCorrelationRulesRaw()`** in [[`spiderfoot/helpers.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/helpers.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/helpers.py) traverses the `correlations/` directory and reads every `*.yaml` file into an internal `ruleset` dictionary.

The execution engine in [[`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) processes each rule through four sequential stages:

1. **collections** – Gather matching records from the SpiderFoot database
2. **aggregation** *(optional)* – Bucket records by a common field
3. **analysis** *(optional)* – Apply conditional thresholds or counts
4. **headline** – Generate the final correlation title with placeholders

Successful correlations are persisted via **`SpiderFootDB.correlationResultCreate()`** in [[`spiderfoot/db.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/db.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/db.py), storing results in `tbl_scan_correlation_results` and linked events in `tbl_scan_correlation_results_events`.

## YAML Rule Structure Reference

Each correlation rule consists of five logical sections:

| Section | Required | Purpose |
|---------|----------|---------|
| **meta** | Yes | Human-readable name, description, and risk level |
| **collections** | Yes | Data-gathering blocks with method chains |
| **aggregation** | No | Grouping field for pre-analysis bucketing |
| **analysis** | No | Conditional logic (thresholds, counts) |
| **headline** | Yes | Title template with field placeholders |

### Available Collection Methods

The engine supports multiple matching strategies in collect blocks:

- **`exact`** – Field matches value precisely
- **`regex`** – Regular expression match against field
- **`exists`** – Field has any value
- **`contains`** – Field contains substring

Field references can target `source.field`, `child.field`, or `entity.field` depending on data context.

## Step-by-Step: Creating Your First Custom Rule

### Step 1: Copy the Template

Duplicate [`correlations/template.yaml`](https://github.com/smicallef/spiderfoot/blob/main/correlations/template.yaml) and name your file to match the rule's `id`. The filename without extension **must** equal the `id` field inside the YAML.

```bash
cd /path/to/spiderfoot
cp correlations/template.yaml correlations/admin_interface.yaml

```

### Step 2: Define Rule Metadata

Edit the **meta** section with clear, actionable descriptions:

```yaml
id: admin_interface
version: 1
meta:
  name: Exposed administrative interfaces
  description: >
    Detects hosts serving login panels or admin dashboards
    on non-standard ports, indicating potential attack surface.
  risk: HIGH

```

Risk levels: `HIGH`, `MEDIUM`, `LOW`, `INFO`.

### Step 3: Configure Data Collections

Build one or more `collect` blocks with ordered method lists:

```yaml
collections:
  - collect:
      - method: exact
        field: type
        value: TCP_PORT_OPEN
      - method: regex
        field: data
        value:
          - :80\b
          - :443\b
      - method: regex
        field: module
        value:
          - ^sfp_spider$

```

Multiple methods in a collect block act as **AND** conditions. Multiple collect blocks act as **OR** conditions.

### Step 4: Add Aggregation (Optional)

Group records before analysis to detect patterns across multiple findings:

```yaml
aggregation:
  field: source.data   # bucket by parent entity

```

Common aggregation fields: `data`, `source.data`, `module`.

### Step 5: Add Analysis Conditions (Optional)

Trigger correlations only when thresholds are met:

```yaml
analysis:
  - method: threshold
    minimum: 5
    field: data

```

Available analysis methods:

- **`threshold`** – Trigger when bucket contains ≥ `minimum` items (or ≤ `maximum`)
- **`count`** – Exact match to specified count

### Step 6: Craft the Headline Template

Use placeholders matching your collected fields:

```yaml
headline: "Host {source.data} has {count} open web ports"

```

Common placeholders: `{data}`, `{module}`, `{type}`, `{source.data}`, `{count}`.

### Step 7: Validate and Restart

Validate YAML syntax to prevent startup failures:

```bash
yamllint correlations/admin_interface.yaml

```

Restart SpiderFoot to load the new rule. The rule appears in the **Correlation Rules** list via the `scancorrelationrules` API endpoint.

## Complete YAML Examples

### Example 1: Admin and Login Sub-domains

Detects hosts containing both "admin" and "login" sub-domains, indicating exposed administrative interfaces.

```yaml

# File: correlations/admin_login.yaml

id: admin_login
version: 1
meta:
  name: Hosts with admin and login sub-domains
  description: >
    Finds Internet names that have both an "admin" and a "login" sub-domain.
    This often indicates a publicly exposed administrative interface.
  risk: HIGH

collections:
  - collect:
      - method: exact
        field: type
        value: INTERNET_NAME
      - method: regex
        field: data
        value:
          - ^admin\..*$
      - method: regex
        field: data
        value:
          - ^login\..*$

aggregation:
  field: data

analysis:
  - method: threshold
    minimum: 2
    field: data

headline: "Host {data} exposes both admin and login sub-domains"

```

This rule uses **two regex matches against the same field** to require both patterns in the collected dataset, then aggregates by hostname and requires at least two matches.

### Example 2: IP Addresses in Multiple WHOIS Records

Flags infrastructure appearing across multiple WHOIS registrations, suggesting shared or suspicious hosting.

```yaml
id: ip_multiple_whois
version: 1
meta:
  name: IP seen in multiple WHOIS records
  description: >
    An IP address that shows up in several WHOIS entries may belong to
    a shared service or be part of a malicious infrastructure.
  risk: MEDIUM

collections:
  - collect:
      - method: exact
        field: type
        value: IPV4_ADDRESS
      - method: exists
        field: module
        value: sfp_whois

aggregation:
  field: data

analysis:
  - method: threshold
    minimum: 3
    field: data

headline: "IP {data} appears in ≥3 WHOIS records"

```

The **`exists`** method here ensures WHOIS data was actually retrieved by the `sfp_whois` module.

## Key Source Files and Their Roles

| File | Function |
|------|----------|
| [[`spiderfoot/helpers.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/helpers.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/helpers.py) | **`loadCorrelationRulesRaw()`** – loads all YAML files from `correlations/` |
| [[`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) | Parses, validates, and executes the rule pipeline; handles `SyntaxError` for malformed YAML |
| [[`spiderfoot/db.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/db.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/db.py) | **`correlationResultCreate()`** – persists results to database tables |
| [[`correlations/template.yaml`](https://github.com/smicallef/spiderfoot/blob/main/correlations/template.yaml)](https://github.com/smicallef/spiderfoot/blob/master/correlations/template.yaml) | Canonical reference for all supported sections and keys |
| [[`correlations/README.md`](https://github.com/smicallef/spiderfoot/blob/main/correlations/README.md)](https://github.com/smicallef/spiderfoot/blob/master/correlations/README.md) | Field naming conventions and method documentation |

## Troubleshooting Custom Correlation Rules

### Startup Errors

Malformed YAML triggers a clear `SyntaxError` from [[`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py). Check:

- Consistent indentation (spaces, not tabs)
- Valid YAML list syntax for `value:` fields with multiple items
- Matching `id` and filename

### Rules Not Triggering

- Verify field names match actual SpiderFoot event types
- Check that aggregation and analysis fields exist in collected data
- Test regex patterns independently—anchors (`^`, `$`) are often required

### Performance Considerations

- Aggregation reduces memory pressure for high-volume matches
- Specific `exact` matches before `regex` improves collection speed
- Multiple collect blocks increase query load—consolidate where possible

## Summary

- SpiderFoot loads correlation rules from YAML files in the `correlations/` directory at startup via `SpiderFootHelpers.loadCorrelationRulesRaw()`
- Rules execute through four stages: **collections → aggregation → analysis → headline**
- The **meta** section defines UI-visible name, description, and risk level (`HIGH` through `INFO`)
- **Collections** use chained methods (`exact`, `regex`, `exists`, `contains`) to select database records
- **Aggregation** buckets results by field; **analysis** applies threshold or count conditions
- The **headline** template generates human-readable correlation titles with field placeholders
- Rules are validated at load time—malformed YAML causes startup errors with clear messages
- Reference implementations exist in [[`correlations/template.yaml`](https://github.com/smicallef/spiderfoot/blob/main/correlations/template.yaml)](https://github.com/smicallef/spiderfoot/blob/master/correlations/template.yaml) and [[`correlations/README.md`](https://github.com/smicallef/spiderfoot/blob/main/correlations/README.md)](https://github.com/smicallef/spiderfoot/blob/master/correlations/README.md)

## Frequently Asked Questions

### What happens if my YAML file has syntax errors?

SpiderFoot will abort startup with a `SyntaxError` that identifies the problematic file and line. The error originates in [[`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py)](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/correlation.py) during rule parsing. Validate your YAML with `yamllint` before deployment.

### Can I reference fields from different event types in the same rule?

Yes. The collection engine allows field references from `source.field` (the parent entity), `child.field` (downstream elements), and `entity.field` (the originating target). These prefixes enable cross-event correlation logic.

### How do I test a correlation rule without running a full scan?

SpiderFoot does not provide an isolated rule tester. Create a minimal scan with known data that should trigger your rule, then verify results appear under the **Correlations** tab. Check the `scancorrelationrules` API endpoint to confirm your rule loaded correctly.

### Do correlation rules affect scan performance?

Complex rules with multiple collect blocks, unanchored regex patterns, or missing aggregation can increase database load. Optimize by using `exact` matches before `regex`, adding aggregation for high-cardinality fields, and limiting collect blocks to necessary conditions.