How to Write a Custom SpiderFoot Module in Python: A Step-by-Step Guide

Create a custom SpiderFoot module by inheriting from SpiderFootPlugin, defining metadata, declaring watched and produced events, and implementing handleEvent() to react to incoming data and emit new findings.

SpiderFoot is an open-source OSINT (Open Source Intelligence) automation platform that discovers data through an event-driven plugin architecture. Writing a custom SpiderFoot module allows you to integrate proprietary feeds, third-party APIs, or bespoke analysis logic into the scanning engine. According to the smicallef/spiderfoot source code, every module follows a consistent pattern built around the SpiderFootPlugin base class.

Understanding the SpiderFoot Plugin Architecture

The core of any SpiderFoot module resides in spiderfoot/plugin.py. This file defines SpiderFootPlugin, an abstract base class that handles logging, configuration, queue management, and the plugin lifecycle.

When SpiderFoot initiates a scan, it:

  1. Loads every Python file under modules/ containing a class that inherits from SpiderFootPlugin
  2. Calls setup() on each plugin with the SpiderFoot instance and user-supplied options
  3. Distributes root target events to all plugins that declared interest via watchedEvents()
  4. Queues and routes new events produced by plugins to other interested plugins, forming a directed acyclic graph of discoveries

This event-driven architecture enables modules to chain together—one module's output becomes another's input without explicit dependencies.

The Nine-Step Pattern for Custom SpiderFoot Modules

Every custom SpiderFoot module in Python follows this implementation pattern derived from the source code:

Step Method/Attribute Purpose
1 Import SpiderFootPlugin, SpiderFootEvent Access base classes from the spiderfoot package
2 Class definition Inherit from SpiderFootPlugin
3 meta dictionary UI metadata: name, summary, categories, use cases
4 opts dictionary Default configuration values
5 optdescs dictionary Human-readable option descriptions for the UI
6 watchedEvents() Return list of event types to consume
7 producedEvents() Return list of event types to emit
8 setup(self, sf, userOpts={}) Initialize plugin state, store self.sf reference
9 handleEvent(self, event) Core logic: process input, query sources, emit results

Complete Custom Module Example

Below is a production-ready SpiderFoot module skeleton based on patterns found in modules/sfp_customfeed.py. Save this as modules/sfp_mycustom.py:


# -*- coding: utf-8 -*-

from spiderfoot import SpiderFootEvent, SpiderFootPlugin


class sfp_mycustom(SpiderFootPlugin):

    meta = {
        'name': "My Custom Module",
        'summary': "Demonstrates how to add a custom data source.",
        'flags': [],
        'useCases': ["Investigate"],
        'categories': ["Reputation Systems"]
    }

    # Default options – user can override them in the UI or via CLI.

    opts = {
        'url': '',
        'cacheperiod': 24          # hours

    }

    # Human-readable descriptions for the options.

    optdescs = {
        'url': "URL of the custom feed (one entry per line).",
        'cacheperiod': "Cache lifetime in hours; 0 forces a fresh download each run."
    }

    def setup(self, sf, userOpts=dict()):
        """Initialize the plugin before scan start."""
        self.sf = sf
        self.results = self.tempStorage()
        for opt in userOpts:
            self.opts[opt] = userOpts[opt]

    def watchedEvents(self):
        """Events we want to receive."""
        return ["INTERNET_NAME", "IP_ADDRESS"]

    def producedEvents(self):
        """Events we may emit."""
        return ["MALICIOUS_INTERNET_NAME", "MALICIOUS_IPADDR"]

    def handleEvent(self, event):
        """Core logic – lookup the target in the custom feed."""
        if self.opts['url'] == "":
            self.error("No feed URL configured")
            return

        # Avoid duplicate work using temporary storage.

        if event.data in self.results:
            return
        self.results[event.data] = True

        # Pull from cache or fetch fresh data.

        data = self.sf.cacheGet("sfmycustom", self.opts['cacheperiod'])
        if data is None:
            fetched = self.sf.fetchUrl(self.opts['url'])
            if fetched['content'] is None:
                self.error("Unable to download feed")
                return
            data = fetched['content']
            self.sf.cachePut("sfmycustom", data)

        # Simple exact-match check against feed entries.

        if event.data in data.splitlines():
            evttype = ("MALICIOUS_INTERNET_NAME"
                       if event.eventType.startswith("INTERNET")
                       else "MALICIOUS_IPADDR")
            newevt = SpiderFootEvent(
                evttype,
                f"Found in custom feed\n<SFURL>{self.opts['url']}</SFURL>",
                self.__name__,
                event
            )
            self.notifyListeners(newevt)

Key SpiderFoot Module Components Explained

The SpiderFoot Object (self.sf)

The self.sf reference, established in setup(), provides essential helper methods for custom SpiderFoot modules:

  • self.sf.fetchUrl(url) – HTTP/HTTPS retrieval with automatic headers, timeout handling, and response parsing
  • self.sf.cacheGet(key, hours) / self.sf.cachePut(key, data) – Persistent cross-scan caching
  • self.sf.hostDomain(hostname) – Extract the domain portion from a hostname
  • self.sf.error(message) – Log errors visible in the scan output

Event Creation and Distribution

New findings propagate through the system via SpiderFootEvent and self.notifyListeners():

newevt = SpiderFootEvent(
    eventType="MALICIOUS_IPADDR",           # Must be in producedEvents()

    data=f"Evidence: {details}",            # The actual finding

    module=self.__name__,                   # Source module name

    sourceEvent=parent_event                # Traces lineage

)
self.notifyListeners(newevt)                # Routes to watching modules

The sourceEvent parameter maintains event provenance, enabling SpiderFoot to display complete chains of how data was discovered.

Temporary Storage for State Management

Use self.tempStorage() to maintain per-scan state across multiple handleEvent invocations:

def setup(self, sf, userOpts=dict()):
    self.sf = sf
    self.seen_targets = self.tempStorage()  # Dict-like storage

def handleEvent(self, event):
    if event.data in self.seen_targets:
        return  # Already processed

    self.seen_targets[event.data] = True

This avoids redundant API calls and prevents circular event generation.

Advanced Pattern: CSV Feed Parsing with Regex

For modules parsing structured feeds, implement helper methods to keep handleEvent clean. This pattern appears in sfp_customfeed.py:

import re


def lookup_csv(self, target, regex):
    """Helper that loads a CSV feed and matches with a regex."""
    content = self.sf.cacheGet("sfmycsv", self.opts['cacheperiod'])
    if content is None:
        fetched = self.sf.fetchUrl(self.opts['url'])
        if fetched['content'] is None:
            self.error("Unable to download CSV feed")
            return None
        content = fetched['content']
        self.sf.cachePut("sfmycsv", content)

    pattern = re.compile(regex, re.IGNORECASE)
    for line in content.splitlines():
        if pattern.search(line):
            return line
    return None

Critical Source Files for Module Development

File Role
spiderfoot/plugin.py SpiderFootPlugin base class—lifecycle, logging, queue handling
spiderfoot/event.py SpiderFootEvent definition and event type constants
modules/sfp_customfeed.py Production example of a custom feed parser module
modules/__init__.py Package marker enabling automatic plugin discovery
sf.py Entry point that loads modules and orchestrates scans

Testing Your Custom SpiderFoot Module

  1. Place your file in modules/sfp_yourname.py
  2. Restart SpiderFoot to register the new module
  3. Enable it in Settings → Modules and configure options
  4. Run a scan against a target that triggers your watchedEvents()
  5. Check Browse and Scans tabs for emitted producedEvents()

Summary

  • Inherit from SpiderFootPlugin in spiderfoot/plugin.py to create any custom SpiderFoot module
  • Define meta, opts, optdescs to control UI presentation and configuration
  • Implement watchedEvents() and producedEvents() to declare your module's data contract
  • Use setup() for initialization and handleEvent() for core logic
  • Leverage self.sf helpers for HTTP, caching, and utility functions
  • Emit findings via SpiderFootEvent and self.notifyListeners() to integrate with the event graph
  • Study modules/sfp_customfeed.py for a complete, production-tested reference implementation

Frequently Asked Questions

What is the minimum code needed for a working SpiderFoot module?

You need a class inheriting from SpiderFootPlugin with meta, watchedEvents(), producedEvents(), setup(), and handleEvent() methods. The setup() method must store self.sf = sf, and handleEvent() must accept an event parameter. Even minimal modules can participate in the event-driven scan flow.

How do I prevent my SpiderFoot module from processing the same data twice?

Use self.tempStorage() in setup() to create a per-scan dictionary, then check and populate it in handleEvent(). This pattern appears throughout smicallef/spiderfoot modules like sfp_customfeed.py to avoid redundant API calls and circular event chains.

Where should I save my custom SpiderFoot module file?

Save it as modules/sfp_yourmodule.py in the SpiderFoot root directory. The modules/__init__.py package marker allows SpiderFoot to automatically discover any class inheriting from SpiderFootPlugin when the application starts.

Can my custom module call external APIs or databases?

Yes. Use self.sf.fetchUrl() for HTTP/HTTPS APIs, or import standard Python libraries like requests, sqlite3, or pymongo directly in your module. Store credentials and endpoints in opts / optdescs so users can configure them through the SpiderFoot UI or CLI without modifying code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →