# How to Create and Test Custom Correlation Rules in SpiderFoot

> Learn to create and test custom correlation rules in SpiderFoot. This guide explains YAML file creation, loading, execution, and verification for powerful threat intelligence.

- Repository: [Steve Micallef/spiderfoot](https://github.com/smicallef/spiderfoot)
- Tags: how-to-guide
- Published: 2026-08-15

---

**Create custom correlation rules in SpiderFoot by writing YAML files in the `correlations/` directory, loading them via `SpiderFootHelpers.loadCorrelationRulesRaw()`, and executing them through `SpiderFootCorrelator.run_correlations()`—then verify results in the database or via unit tests.**

SpiderFoot's correlation engine enables security analysts to define custom rules that combine data from multiple reconnaissance modules, aggregate findings, and surface meaningful security insights. This guide walks through the complete workflow for authoring, loading, executing, and testing custom correlation rules using the actual implementation from the smicallef/spiderfoot repository.

## How SpiderFoot's Correlation Engine Works

The correlation system consists of three integrated components working together to process rules against scan data.

| Component | File Location | Core Responsibility |
|-----------|-------------|---------------------|
| **Rule loader** | [`spiderfoot/helpers.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/helpers.py) lines 174-185 | Reads and parses raw YAML rule files from disk |
| **Correlator engine** | [`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py) lines 49-63, 108-131 | Validates rule syntax, executes collection queries, applies aggregation and analysis logic |
| **Result writer** | [`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py) lines 897-906, 929-447 | Generates human-readable titles and persists correlations to the database |

Understanding this pipeline helps you debug rules that fail validation or produce unexpected results.

## Step 1: Author a YAML Correlation Rule

SpiderFoot correlation rules use YAML syntax with a strict schema. Create your rule file in the repository's `correlations/` directory with a filename matching the rule's `id` field.

### Required YAML Structure

| Section | Purpose | Required Keys |
|---------|---------|---------------|
| `meta` | Human-readable metadata | `name`, `description`, `risk` |
| `collections` | Database queries to gather events | `collect` blocks with `method`, `field`, `value` |
| `aggregation` | Event grouping (optional) | `field` to bucket by |
| `analysis` | Thresholds and filters (optional) | `method`, `field`, plus method-specific parameters |
| `headline` | Result title template | Text with `{field}` placeholders |

### Example: Duplicate Hostname Detection

This minimal rule detects hostnames appearing multiple times across scan results—adapted from the template in [`correlations/template.yaml`](https://github.com/smicallef/spiderfoot/blob/main/correlations/template.yaml):

```yaml

# file: correlations/duplicate_host.yaml

id: duplicate_host
version: 1
meta:
  name: Duplicate Hostnames
  description: >
    Finds hostnames that appear more than once in the scan results,
    indicating potential infrastructure overlap or DNS inconsistencies.
  risk: MEDIUM
collections:
  - collect:
      - method: exact
        field: type
        value: INTERNET_NAME
analysis:
  - method: threshold
    minimum: 2        # trigger when count ≥ 2

    field: data
headline: "Duplicate hostname detected: {data}"

```

**Critical requirement**: The `id` field must exactly match the filename without the `.yaml` extension. Mismatches cause the correlator to skip the rule during validation in `SpiderFootCorrelator.__init__()`.

## Step 2: Load Rules into the Correlator

Use the helper utilities to load your custom rules before execution.

```python
from spiderfoot.helpers import SpiderFootHelpers
from spiderfoot.correlation import SpiderFootCorrelator
from spiderfoot.db import SpiderFootDb

# 1️⃣ Load all YAML files from the correlations directory

raw_rules = SpiderFootHelpers.loadCorrelationRulesRaw(
    path="./correlations/",          # directory containing *.yaml files

    ignore_files=["template.yaml"]   # optional: skip specific files

)

# 2️⃣ Initialize database connection

db = SpiderFootDb("./spiderfoot.db")  # or existing DB handle from a scan

# 3️⃣ Instantiate correlator with loaded rules

correlator = SpiderFootCorrelator(
    dbh=db,
    ruleset=raw_rules,
    scanId="your-scan-id-here"
)

```

The `loadCorrelationRulesRaw()` function at lines 174-185 in [`spiderfoot/helpers.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/helpers.py) returns a dictionary mapping rule IDs to raw YAML strings. The `SpiderFootCorrelator` constructor at lines 49-63 in [`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py) then parses each YAML document, validates schema compliance, and builds executable rule objects.

## Step 3: Execute Correlation Rules

Run all loaded rules against the scan data with a single method call:

```python

# Execute every validated rule against the scan database

correlator.run_correlations()

```

The `run_correlations()` method (lines 108-131 in [`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py)) iterates through `self.rules` and invokes `process_rule()` for each entry. Failed rules log warnings without halting execution—check logs for validation errors if results appear incomplete.

## Step 4: Verify Correlation Results

Correlations matching your criteria are stored in `tbl_scan_correlation_results`. Query them programmatically or inspect via SpiderFoot's web interface.

```python

# Retrieve all correlation titles for verification

summary = db.correlationSummary(scanId="your-scan-id-here")
for correlation in summary:
    print(f"Detected: {correlation['title']}")

# Access full correlation details

details = db.correlationResultGet(correlationId="specific-id")

```

The `build_correlation_title()` method (lines 897-906) substitutes `{field}` placeholders in your `headline` template with actual event data. Then `create_correlation()` (lines 929-447) persists the completed result to the database with proper risk scoring and metadata.

## Step 5: Unit Test Custom Correlation Rules

SpiderFoot includes a test suite at [`test/unit/spiderfoot/test_spiderfootcorrelator.py`](https://github.com/smicallef/spiderfoot/blob/main/test/unit/spiderfoot/test_spiderfootcorrelator.py). Follow this pattern to validate your rules programmatically.

### Complete Test Example

```python
import pytest
from spiderfoot.db import SpiderFootDb
from spiderfoot.correlation import SpiderFootCorrelator

def test_duplicate_host_correlation():
    """Verify duplicate hostname detection triggers correctly."""
    
    # 1️⃣ Create isolated in-memory database

    db = SpiderFootDb(":memory:")
    db.scanInstanceCreate(
        instanceId="test-scan-001",
        scanName="Unit Test Scan",
        scanner="pytest",
        moduleList="sfp_dnsdumpster"
    )
    
    # 2️⃣ Insert mock events matching collection criteria

    db.scanResultEventCreate(
        instanceId="test-scan-001",
        eventType="INTERNET_NAME",
        data="vulnerable.example.com",
        module="sfp_dnsdumpster"
    )
    db.scanResultEventCreate(
        instanceId="test-scan-001",
        eventType="INTERNET_NAME",
        data="vulnerable.example.com",  # duplicate entry

        module="sfp_shodan"
    )
    
    # 3️⃣ Load and execute rule

    raw_rules = {
        "duplicate_host": open("correlations/duplicate_host.yaml").read()
    }
    correlator = SpiderFootCorrelator(
        dbh=db,
        ruleset=raw_rules,
        scanId="test-scan-001"
    )
    correlator.run_correlations()
    
    # 4️⃣ Assert expected outcome

    results = db.correlationSummary(scanId="test-scan-001")
    assert len(results) == 1
    assert "vulnerable.example.com" in results[0]["title"]
    assert results[0]["risk"] == "MEDIUM"

```

### Testing Best Practices

- Use `:memory:` SQLite databases for fast, isolated tests
- Include edge cases: zero matches, single matches, threshold boundaries
- Verify both positive matches (rule triggers) and negative matches (rule correctly skips)
- Test aggregation and analysis methods independently when logic grows complex

## Advanced Correlation Rule Patterns

### Aggregation with Field Grouping

Group events before analysis to detect patterns across data subsets:

```yaml
collections:
  - collect:
      - method: exact
        field: type
        value: IP_ADDRESS
aggregation:
  field: scanId          # group by source scan

analysis:
  - method: threshold
    minimum: 5           # 5+ IPs per scanId triggers

    field: data

```

### Multiple Collection Blocks

Combine data from different event types in a single rule:

```yaml
collections:
  - collect:
      - method: exact
        field: type
        value: VULNERABILITY_CVE
  - collect:
      - method: regex
        field: type
        value: "^VULNERABILITY_.*"

```

## Summary

- **Create** correlation rules as YAML files in `correlations/` with `id` matching filename
- **Load** rules via `SpiderFootHelpers.loadCorrelationRulesRaw()` from [`spiderfoot/helpers.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/helpers.py)
- **Validate** rule syntax automatically in `SpiderFootCorrelator.__init__()` at [`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py) lines 49-63
- **Execute** with `run_correlations()` which processes each rule through `process_rule()`
- **Verify** results in `tbl_scan_correlation_results` using `db.correlationSummary()` or direct queries
- **Test** rules using in-memory databases with controlled mock events following the pattern in [`test_spiderfootcorrelator.py`](https://github.com/smicallef/spiderfoot/blob/main/test_spiderfootcorrelator.py)

## Frequently Asked Questions

### What file naming convention must correlation rules follow?

The YAML filename must exactly match the rule's `id` field with `.yaml` appended. A rule with `id: suspicious_cert` must reside in [`correlations/suspicious_cert.yaml`](https://github.com/smicallef/spiderfoot/blob/main/correlations/suspicious_cert.yaml). The `SpiderFootCorrelator` constructor validates this match during initialization and logs warnings for discrepancies.

### How do I debug a correlation rule that produces no results?

First, verify the rule passes validation by checking SpiderFoot logs for YAML syntax errors. Next, confirm your `collections` blocks match actual event types in your scan data—query the database directly with `db.scanResultEvent()` to see available types. Finally, test with lowered thresholds (e.g., `minimum: 1`) to ensure the collection logic works before tightening constraints.

### Can correlation rules reference custom event types from private modules?

Yes. The correlator operates on the database layer and does not validate event types against SpiderFoot's built-in lists. Any string value in `collections.collect[].value` will match corresponding entries in `tbl_scan_results`. Document your custom types clearly to avoid maintenance issues.

### What analysis methods are available beyond `threshold`?

SpiderFoot's [`spiderfoot/correlation.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/correlation.py) implements multiple analysis methods including `outlier` for statistical anomaly detection, `match` for pattern-based filtering, and `compare` for cross-field relationships. Inspect the `analysis_methods` dictionary in the correlator source for the complete, version-specific list and their required parameters.