# How to Write a Custom SpiderFoot Module in Python: A Step-by-Step Guide

> Learn to write a custom SpiderFoot module in Python with this step-by-step guide. Extend SpiderFoot's capabilities by creating your own modules and integrating new data sources.

- Repository: [Steve Micallef/spiderfoot](https://github.com/smicallef/spiderfoot)
- Tags: how-to-guide
- Published: 2026-08-15

---

**Create a custom SpiderFoot module by inheriting from `SpiderFootPlugin`, defining metadata, declaring watched and produced events, and implementing `handleEvent()` to react to incoming data and emit new findings.**

SpiderFoot is an open-source OSINT (Open Source Intelligence) automation platform that discovers data through an event-driven plugin architecture. Writing a **custom SpiderFoot module** allows you to integrate proprietary feeds, third-party APIs, or bespoke analysis logic into the scanning engine. According to the `smicallef/spiderfoot` source code, every module follows a consistent pattern built around the `SpiderFootPlugin` base class.

## Understanding the SpiderFoot Plugin Architecture

The core of any SpiderFoot module resides in **[`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py)**. This file defines `SpiderFootPlugin`, an abstract base class that handles logging, configuration, queue management, and the plugin lifecycle.

When SpiderFoot initiates a scan, it:

1. Loads every Python file under `modules/` containing a class that inherits from `SpiderFootPlugin`
2. Calls `setup()` on each plugin with the `SpiderFoot` instance and user-supplied options
3. Distributes root target events to all plugins that declared interest via `watchedEvents()`
4. Queues and routes new events produced by plugins to other interested plugins, forming a directed acyclic graph of discoveries

This **event-driven architecture** enables modules to chain together—one module's output becomes another's input without explicit dependencies.

## The Nine-Step Pattern for Custom SpiderFoot Modules

Every custom **SpiderFoot module in Python** follows this implementation pattern derived from the source code:

| Step | Method/Attribute | Purpose |
|------|------------------|---------|
| 1 | Import `SpiderFootPlugin`, `SpiderFootEvent` | Access base classes from the spiderfoot package |
| 2 | Class definition | Inherit from `SpiderFootPlugin` |
| 3 | `meta` dictionary | UI metadata: name, summary, categories, use cases |
| 4 | `opts` dictionary | Default configuration values |
| 5 | `optdescs` dictionary | Human-readable option descriptions for the UI |
| 6 | `watchedEvents()` | Return list of event types to consume |
| 7 | `producedEvents()` | Return list of event types to emit |
| 8 | `setup(self, sf, userOpts={})` | Initialize plugin state, store `self.sf` reference |
| 9 | `handleEvent(self, event)` | Core logic: process input, query sources, emit results |

## Complete Custom Module Example

Below is a production-ready **SpiderFoot module skeleton** based on patterns found in [`modules/sfp_customfeed.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_customfeed.py). Save this as [`modules/sfp_mycustom.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_mycustom.py):

```python

# -*- coding: utf-8 -*-

from spiderfoot import SpiderFootEvent, SpiderFootPlugin


class sfp_mycustom(SpiderFootPlugin):

    meta = {
        'name': "My Custom Module",
        'summary': "Demonstrates how to add a custom data source.",
        'flags': [],
        'useCases': ["Investigate"],
        'categories': ["Reputation Systems"]
    }

    # Default options – user can override them in the UI or via CLI.

    opts = {
        'url': '',
        'cacheperiod': 24          # hours

    }

    # Human-readable descriptions for the options.

    optdescs = {
        'url': "URL of the custom feed (one entry per line).",
        'cacheperiod': "Cache lifetime in hours; 0 forces a fresh download each run."
    }

    def setup(self, sf, userOpts=dict()):
        """Initialize the plugin before scan start."""
        self.sf = sf
        self.results = self.tempStorage()
        for opt in userOpts:
            self.opts[opt] = userOpts[opt]

    def watchedEvents(self):
        """Events we want to receive."""
        return ["INTERNET_NAME", "IP_ADDRESS"]

    def producedEvents(self):
        """Events we may emit."""
        return ["MALICIOUS_INTERNET_NAME", "MALICIOUS_IPADDR"]

    def handleEvent(self, event):
        """Core logic – lookup the target in the custom feed."""
        if self.opts['url'] == "":
            self.error("No feed URL configured")
            return

        # Avoid duplicate work using temporary storage.

        if event.data in self.results:
            return
        self.results[event.data] = True

        # Pull from cache or fetch fresh data.

        data = self.sf.cacheGet("sfmycustom", self.opts['cacheperiod'])
        if data is None:
            fetched = self.sf.fetchUrl(self.opts['url'])
            if fetched['content'] is None:
                self.error("Unable to download feed")
                return
            data = fetched['content']
            self.sf.cachePut("sfmycustom", data)

        # Simple exact-match check against feed entries.

        if event.data in data.splitlines():
            evttype = ("MALICIOUS_INTERNET_NAME"
                       if event.eventType.startswith("INTERNET")
                       else "MALICIOUS_IPADDR")
            newevt = SpiderFootEvent(
                evttype,
                f"Found in custom feed\n<SFURL>{self.opts['url']}</SFURL>",
                self.__name__,
                event
            )
            self.notifyListeners(newevt)

```

## Key SpiderFoot Module Components Explained

### The `SpiderFoot` Object (`self.sf`)

The `self.sf` reference, established in `setup()`, provides essential helper methods for **custom SpiderFoot modules**:

- `self.sf.fetchUrl(url)` – HTTP/HTTPS retrieval with automatic headers, timeout handling, and response parsing
- `self.sf.cacheGet(key, hours)` / `self.sf.cachePut(key, data)` – Persistent cross-scan caching
- `self.sf.hostDomain(hostname)` – Extract the domain portion from a hostname
- `self.sf.error(message)` – Log errors visible in the scan output

### Event Creation and Distribution

New findings propagate through the system via `SpiderFootEvent` and `self.notifyListeners()`:

```python
newevt = SpiderFootEvent(
    eventType="MALICIOUS_IPADDR",           # Must be in producedEvents()

    data=f"Evidence: {details}",            # The actual finding

    module=self.__name__,                   # Source module name

    sourceEvent=parent_event                # Traces lineage

)
self.notifyListeners(newevt)                # Routes to watching modules

```

The `sourceEvent` parameter maintains **event provenance**, enabling SpiderFoot to display complete chains of how data was discovered.

### Temporary Storage for State Management

Use `self.tempStorage()` to maintain per-scan state across multiple `handleEvent` invocations:

```python
def setup(self, sf, userOpts=dict()):
    self.sf = sf
    self.seen_targets = self.tempStorage()  # Dict-like storage

def handleEvent(self, event):
    if event.data in self.seen_targets:
        return  # Already processed

    self.seen_targets[event.data] = True

```

This avoids redundant API calls and prevents circular event generation.

## Advanced Pattern: CSV Feed Parsing with Regex

For modules parsing structured feeds, implement helper methods to keep `handleEvent` clean. This pattern appears in [`sfp_customfeed.py`](https://github.com/smicallef/spiderfoot/blob/main/sfp_customfeed.py):

```python
import re


def lookup_csv(self, target, regex):
    """Helper that loads a CSV feed and matches with a regex."""
    content = self.sf.cacheGet("sfmycsv", self.opts['cacheperiod'])
    if content is None:
        fetched = self.sf.fetchUrl(self.opts['url'])
        if fetched['content'] is None:
            self.error("Unable to download CSV feed")
            return None
        content = fetched['content']
        self.sf.cachePut("sfmycsv", content)

    pattern = re.compile(regex, re.IGNORECASE)
    for line in content.splitlines():
        if pattern.search(line):
            return line
    return None

```

## Critical Source Files for Module Development

| File | Role |
|------|------|
| [`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py) | `SpiderFootPlugin` base class—lifecycle, logging, queue handling |
| [`spiderfoot/event.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/event.py) | `SpiderFootEvent` definition and event type constants |
| [`modules/sfp_customfeed.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_customfeed.py) | Production example of a custom feed parser module |
| [`modules/__init__.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/__init__.py) | Package marker enabling automatic plugin discovery |
| [`sf.py`](https://github.com/smicallef/spiderfoot/blob/main/sf.py) | Entry point that loads modules and orchestrates scans |

## Testing Your Custom SpiderFoot Module

1. Place your file in [`modules/sfp_yourname.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_yourname.py)
2. Restart SpiderFoot to register the new module
3. Enable it in **Settings → Modules** and configure options
4. Run a scan against a target that triggers your `watchedEvents()`
5. Check **Browse** and **Scans** tabs for emitted `producedEvents()`

## Summary

- **Inherit from `SpiderFootPlugin`** in [`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py) to create any **custom SpiderFoot module**
- **Define `meta`, `opts`, `optdescs`** to control UI presentation and configuration
- **Implement `watchedEvents()` and `producedEvents()`** to declare your module's data contract
- **Use `setup()` for initialization** and `handleEvent()` for core logic
- **Leverage `self.sf` helpers** for HTTP, caching, and utility functions
- **Emit findings via `SpiderFootEvent`** and `self.notifyListeners()` to integrate with the event graph
- **Study [`modules/sfp_customfeed.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_customfeed.py)** for a complete, production-tested reference implementation

## Frequently Asked Questions

### What is the minimum code needed for a working SpiderFoot module?

You need a class inheriting from `SpiderFootPlugin` with `meta`, `watchedEvents()`, `producedEvents()`, `setup()`, and `handleEvent()` methods. The `setup()` method must store `self.sf = sf`, and `handleEvent()` must accept an event parameter. Even minimal modules can participate in the event-driven scan flow.

### How do I prevent my SpiderFoot module from processing the same data twice?

Use `self.tempStorage()` in `setup()` to create a per-scan dictionary, then check and populate it in `handleEvent()`. This pattern appears throughout `smicallef/spiderfoot` modules like [`sfp_customfeed.py`](https://github.com/smicallef/spiderfoot/blob/main/sfp_customfeed.py) to avoid redundant API calls and circular event chains.

### Where should I save my custom SpiderFoot module file?

Save it as [`modules/sfp_yourmodule.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_yourmodule.py) in the SpiderFoot root directory. The [`modules/__init__.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/__init__.py) package marker allows SpiderFoot to automatically discover any class inheriting from `SpiderFootPlugin` when the application starts.

### Can my custom module call external APIs or databases?

Yes. Use `self.sf.fetchUrl()` for HTTP/HTTPS APIs, or import standard Python libraries like `requests`, `sqlite3`, or `pymongo` directly in your module. Store credentials and endpoints in `opts` / `optdescs` so users can configure them through the SpiderFoot UI or CLI without modifying code.