How SpiderFoot's Module Plugin Architecture Works: A Deep Dive into the Plugin Framework

SpiderFoot uses a lightweight plugin framework centered on the abstract SpiderFootPlugin base class, where every data-source module acts as an interchangeable component that communicates through a shared event-driven system.

The smicallef/spiderfoot open-source OSINT tool is built entirely around extensible modular architecture. This design lets security researchers add new intelligence sources without modifying core scan logic. Understanding how SpiderFoot's module plugin architecture works reveals why the tool remains one of the most flexible reconnaissance platforms available.

The Core Plugin Contract: SpiderFootPlugin

All SpiderFoot modules inherit from a single abstract base class defined in spiderfoot/plugin.py. This class establishes the mandatory interface that every module must implement.

Lifecycle Hooks

The base class provides standardized hooks that the scan controller invokes:

  • setup(sf, userOpts) – Initializes the module with the central SpiderFoot object (sf) and user-supplied configuration options
  • start() – Signals the beginning of event processing
  • threadWorker() – Pulls events from incomingEventQueue when running in multi-threaded mode
  • finish() – Cleanup hook called when the scan terminates

Built-in Infrastructure

Modules receive ready-made functionality from SpiderFootPlugin:

  • Logging utilities – self.log, debug(), info(), and error() methods automatically embed the current scan ID
  • Thread-pool integration – Access to sharedThreadPool for concurrent execution; storage plugins run directly without threading
  • Event queues – incomingEventQueue (events to consume) and outgoingEventQueue (events to emit)
  • Output filtering – __outputFilter__ limits which event types propagate
  • Stop-signal handling – checkForStop() monitors abort requests and error states

Module Structure and Required Methods

Each module lives as a Python file under modules/ (for example, modules/sfp_dnsresolve.py). Every module must define several components:

Component Purpose
meta dictionary Human-readable metadata including name, summary, categories, and flags for UI display
opts and optdescs Configuration options exposed in the web interface
watchedEvents() Returns list of event types the module consumes (['DNS_NAME', 'IP_ADDRESS'] or ['*'] for all)
producedEvents() Returns event types the module generates
handleEvent(sfEvent) Core processing logic—most modules override this method
Optional threadWorker() Custom worker when maxThreads > 1

Here is a complete minimal module implementation:

from spiderfoot import SpiderFootPlugin

class sfp_example(SpiderFootPlugin):
    """Example module that demonstrates the plugin API."""
    meta = {
        'name': 'Example',
        'summary': 'Shows how a plugin works',
        'categories': ['Demo'],
        'flags': []
    }
    opts = {}
    optdescs = {}

    def watchedEvents(self):
        return ['DOMAIN_NAME']

    def producedEvents(self):
        return ['EXAMPLE_RESULT']

    def setup(self, sf, userOpts):
        self.sf = sf               # spiderfoot core object

        self.opts = userOpts

    def handleEvent(self, sfEvent):
        # Do something with sfEvent.data …

        result = f"processed {sfEvent.data}"
        self.sf.emitEvent(result, "EXAMPLE_RESULT", sfEvent)

Event-Driven Data Flow

SpiderFoot's module plugin architecture operates on a publish-subscribe event model. The controller orchestrates all module communication:

  1. Scan initialization – The controller creates a SpiderFoot instance and loads every module from modules/
  2. Module setup – Each module is instantiated and setup() receives the central SpiderFoot object plus configuration
  3. Root event seeding – A SpiderFootEvent('ROOT', target) enters every module's incomingEventQueue
  4. Event processing – The module's threadWorker() or direct handleEvent() processes events and calls self.sf.emitEvent() to generate new data
  5. Event propagation – notifyListeners() in SpiderFootPlugin routes events by checking each destination module's watchedEvents(), prevents duplicates via storeOnly logic, and either queues for threaded listeners or calls handleEvent() directly
  6. Scan completion – The FINISHED sentinel triggers finish() hooks across all modules

Dynamic Module Discovery

SpiderFoot discovers plugins automatically by scanning the modules/ package for SpiderFootPlugin subclasses. The main entry point in sf.py imports each Python file, instantiates the class, and collects meta dictionaries for UI rendering.

Because discovery relies on Python's import system, adding a new module requires only placing a properly structured .py file into modules/—no registry edits or core modifications needed.

Thread Pool Coordination

Non-storage plugins share a global SpiderFootThreadPool implemented in spiderfoot/threadpool.py. When a module submits work via self.poolExecute(callback, *args), the task receives a name derived from the module (f"{self.__name__}_threadWorker").

This architecture enables:

  • Concurrency limits – Enforced via maxThreads configuration
  • Work tracking – Monitor pending tasks with self.sharedThreadPool.countQueuedTasks()
  • Graceful aborts – Queue clearing and stop signals propagate to all running modules

Practical Examples

Creating a Complete New Plugin


# modules/sfp_mynewsource.py

from spiderfoot import SpiderFootPlugin

class sfp_mynewsource(SpiderFootPlugin):
    meta = {
        'name': 'MyNewSource',
        'summary': 'Fetches data from MyNewSource API',
        'categories': ['API'],
        'flags': ['requires_api_key']
    }

    def watchedEvents(self):
        return ['IP_ADDRESS']

    def producedEvents(self):
        return ['MYNEW_DATA']

    def setup(self, sf, userOpts):
        self.sf = sf
        self.opts = userOpts
        self.api_key = self.opts.get('api_key')

    def handleEvent(self, sfEvent):
        ip = sfEvent.data
        # API call implementation (omitted for brevity)

        data = self._query_api(ip)
        self.sf.emitEvent(data, 'MYNEW_DATA', sfEvent)

Save this as modules/sfp_mynewsource.py—SpiderFoot automatically loads it on next startup.

Emitting Events for Cross-Module Communication

def handleEvent(self, sfEvent):
    # Generate new intelligence

    domain = f"{sfEvent.data}.example.com"
    # Broadcast to interested modules (DNS resolution, WHOIS, etc.)

    self.sf.emitEvent(domain, "DOMAIN_NAME", sfEvent)

Key Source Files

File Role
spiderfoot/plugin.py Abstract base class SpiderFootPlugin defining plugin contracts and event routing
spiderfoot/threadpool.py Shared thread pool for concurrent module execution
spiderfoot/__init__.py Exports SpiderFootPlugin, SpiderFootTarget, and other core classes
modules/*.py Concrete plugin implementations (e.g., modules/sfp_dnsresolve.py)
sf.py CLI entry point handling dynamic discovery and plugin wiring

Summary

  • SpiderFoot's module plugin architecture centers on the SpiderFootPlugin abstract base class in spiderfoot/plugin.py
  • Modules implement watchedEvents(), producedEvents(), setup(), and handleEvent() to participate in scans
  • The event-driven publish-subscribe model decouples data sources, letting modules react to findings from any other module
  • Dynamic discovery via Python imports eliminates registration overhead—drop a file in modules/ and it works
  • Shared thread pooling in spiderfoot/threadpool.py controls concurrency and enables clean scan abortion

Frequently Asked Questions

How do I create a new SpiderFoot module from scratch?

Subclass SpiderFootPlugin from spiderfoot/plugin.py, implement watchedEvents(), producedEvents(), setup(), and handleEvent(), then save your file to modules/sfp_yourname.py. SpiderFoot automatically discovers and loads it on next startup. Include a meta dictionary with name, summary, and categories for UI integration.

What is the difference between watchedEvents() and producedEvents()?

watchedEvents() returns the event types your module wants to receive (your inputs), while producedEvents() declares what your module generates (your outputs). The controller uses these declarations to build the event routing graph—modules only receive events they explicitly watch.

Can SpiderFoot modules run in parallel?

Yes. When maxThreads exceeds 1, modules execute their threadWorker() method in the shared SpiderFootThreadPool from spiderfoot/threadpool.py. Storage modules typically run single-threaded. Use self.poolExecute() to submit additional background tasks with automatic module-name tracking.

How does SpiderFoot prevent infinite event loops between modules?

The framework implements duplicate detection through storeOnly logic in notifyListeners() and stop-signal checking via checkForStop(). Events carry source traceability, and the controller monitors queue depths and cycle patterns to halt runaway processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →