# How SpiderFoot's Module Plugin Architecture Works: A Deep Dive into the Plugin Framework

> Explore SpiderFoot's module plugin architecture. Learn how interchangeable data-source modules communicate via an event-driven system built on the SpiderFootPlugin base class.

- Repository: [Steve Micallef/spiderfoot](https://github.com/smicallef/spiderfoot)
- Tags: deep-dive
- Published: 2026-08-15

---

**SpiderFoot uses a lightweight plugin framework centered on the abstract `SpiderFootPlugin` base class, where every data-source module acts as an interchangeable component that communicates through a shared event-driven system.**

The `smicallef/spiderfoot` open-source OSINT tool is built entirely around extensible **modular architecture**. This design lets security researchers add new intelligence sources without modifying core scan logic. Understanding how SpiderFoot's module plugin architecture works reveals why the tool remains one of the most flexible reconnaissance platforms available.

## The Core Plugin Contract: `SpiderFootPlugin`

All SpiderFoot modules inherit from a single abstract base class defined in **[`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py)**. This class establishes the mandatory interface that every module must implement.

### Lifecycle Hooks

The base class provides standardized hooks that the scan controller invokes:

- **`setup(sf, userOpts)`** – Initializes the module with the central `SpiderFoot` object (`sf`) and user-supplied configuration options
- **`start()`** – Signals the beginning of event processing
- **`threadWorker()`** – Pulls events from `incomingEventQueue` when running in multi-threaded mode
- **`finish()`** – Cleanup hook called when the scan terminates

### Built-in Infrastructure

Modules receive ready-made functionality from `SpiderFootPlugin`:

- **Logging utilities** – `self.log`, `debug()`, `info()`, and `error()` methods automatically embed the current scan ID
- **Thread-pool integration** – Access to `sharedThreadPool` for concurrent execution; storage plugins run directly without threading
- **Event queues** – `incomingEventQueue` (events to consume) and `outgoingEventQueue` (events to emit)
- **Output filtering** – `__outputFilter__` limits which event types propagate
- **Stop-signal handling** – `checkForStop()` monitors abort requests and error states

## Module Structure and Required Methods

Each module lives as a Python file under **`modules/`** (for example, [`modules/sfp_dnsresolve.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_dnsresolve.py)). Every module must define several components:

| Component | Purpose |
|-----------|---------|
| `meta` dictionary | Human-readable metadata including `name`, `summary`, `categories`, and `flags` for UI display |
| `opts` and `optdescs` | Configuration options exposed in the web interface |
| `watchedEvents()` | Returns list of event types the module consumes (`['DNS_NAME', 'IP_ADDRESS']` or `['*']` for all) |
| `producedEvents()` | Returns event types the module generates |
| `handleEvent(sfEvent)` | Core processing logic—most modules override this method |
| Optional `threadWorker()` | Custom worker when `maxThreads > 1` |

Here is a complete minimal module implementation:

```python
from spiderfoot import SpiderFootPlugin

class sfp_example(SpiderFootPlugin):
    """Example module that demonstrates the plugin API."""
    meta = {
        'name': 'Example',
        'summary': 'Shows how a plugin works',
        'categories': ['Demo'],
        'flags': []
    }
    opts = {}
    optdescs = {}

    def watchedEvents(self):
        return ['DOMAIN_NAME']

    def producedEvents(self):
        return ['EXAMPLE_RESULT']

    def setup(self, sf, userOpts):
        self.sf = sf               # spiderfoot core object

        self.opts = userOpts

    def handleEvent(self, sfEvent):
        # Do something with sfEvent.data …

        result = f"processed {sfEvent.data}"
        self.sf.emitEvent(result, "EXAMPLE_RESULT", sfEvent)

```

## Event-Driven Data Flow

SpiderFoot's module plugin architecture operates on a **publish-subscribe event model**. The controller orchestrates all module communication:

1. **Scan initialization** – The controller creates a `SpiderFoot` instance and loads every module from `modules/`
2. **Module setup** – Each module is instantiated and `setup()` receives the central `SpiderFoot` object plus configuration
3. **Root event seeding** – A `SpiderFootEvent('ROOT', target)` enters every module's `incomingEventQueue`
4. **Event processing** – The module's `threadWorker()` or direct `handleEvent()` processes events and calls `self.sf.emitEvent()` to generate new data
5. **Event propagation** – `notifyListeners()` in `SpiderFootPlugin` routes events by checking each destination module's `watchedEvents()`, prevents duplicates via `storeOnly` logic, and either queues for threaded listeners or calls `handleEvent()` directly
6. **Scan completion** – The **FINISHED** sentinel triggers `finish()` hooks across all modules

## Dynamic Module Discovery

SpiderFoot discovers plugins automatically by scanning the **`modules/`** package for `SpiderFootPlugin` subclasses. The main entry point in **[`sf.py`](https://github.com/smicallef/spiderfoot/blob/main/sf.py)** imports each Python file, instantiates the class, and collects `meta` dictionaries for UI rendering.

Because discovery relies on **Python's import system**, adding a new module requires only placing a properly structured `.py` file into `modules/`—no registry edits or core modifications needed.

## Thread Pool Coordination

Non-storage plugins share a **global `SpiderFootThreadPool`** implemented in [`spiderfoot/threadpool.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/threadpool.py). When a module submits work via `self.poolExecute(callback, *args)`, the task receives a name derived from the module (`f"{self.__name__}_threadWorker"`).

This architecture enables:

- **Concurrency limits** – Enforced via `maxThreads` configuration
- **Work tracking** – Monitor pending tasks with `self.sharedThreadPool.countQueuedTasks()`
- **Graceful aborts** – Queue clearing and stop signals propagate to all running modules

## Practical Examples

### Creating a Complete New Plugin

```python

# modules/sfp_mynewsource.py

from spiderfoot import SpiderFootPlugin

class sfp_mynewsource(SpiderFootPlugin):
    meta = {
        'name': 'MyNewSource',
        'summary': 'Fetches data from MyNewSource API',
        'categories': ['API'],
        'flags': ['requires_api_key']
    }

    def watchedEvents(self):
        return ['IP_ADDRESS']

    def producedEvents(self):
        return ['MYNEW_DATA']

    def setup(self, sf, userOpts):
        self.sf = sf
        self.opts = userOpts
        self.api_key = self.opts.get('api_key')

    def handleEvent(self, sfEvent):
        ip = sfEvent.data
        # API call implementation (omitted for brevity)

        data = self._query_api(ip)
        self.sf.emitEvent(data, 'MYNEW_DATA', sfEvent)

```

Save this as **[`modules/sfp_mynewsource.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_mynewsource.py)**—SpiderFoot automatically loads it on next startup.

### Emitting Events for Cross-Module Communication

```python
def handleEvent(self, sfEvent):
    # Generate new intelligence

    domain = f"{sfEvent.data}.example.com"
    # Broadcast to interested modules (DNS resolution, WHOIS, etc.)

    self.sf.emitEvent(domain, "DOMAIN_NAME", sfEvent)

```

## Key Source Files

| File | Role |
|------|------|
| [`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py) | Abstract base class `SpiderFootPlugin` defining plugin contracts and event routing |
| [`spiderfoot/threadpool.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/threadpool.py) | Shared thread pool for concurrent module execution |
| [`spiderfoot/__init__.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/__init__.py) | Exports `SpiderFootPlugin`, `SpiderFootTarget`, and other core classes |
| `modules/*.py` | Concrete plugin implementations (e.g., [`modules/sfp_dnsresolve.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_dnsresolve.py)) |
| [`sf.py`](https://github.com/smicallef/spiderfoot/blob/main/sf.py) | CLI entry point handling dynamic discovery and plugin wiring |

## Summary

- **SpiderFoot's module plugin architecture** centers on the `SpiderFootPlugin` abstract base class in [`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py)
- Modules implement `watchedEvents()`, `producedEvents()`, `setup()`, and `handleEvent()` to participate in scans
- The **event-driven publish-subscribe model** decouples data sources, letting modules react to findings from any other module
- **Dynamic discovery** via Python imports eliminates registration overhead—drop a file in `modules/` and it works
- **Shared thread pooling** in [`spiderfoot/threadpool.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/threadpool.py) controls concurrency and enables clean scan abortion

## Frequently Asked Questions

### How do I create a new SpiderFoot module from scratch?

Subclass `SpiderFootPlugin` from [`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py), implement `watchedEvents()`, `producedEvents()`, `setup()`, and `handleEvent()`, then save your file to [`modules/sfp_yourname.py`](https://github.com/smicallef/spiderfoot/blob/main/modules/sfp_yourname.py). SpiderFoot automatically discovers and loads it on next startup. Include a `meta` dictionary with `name`, `summary`, and `categories` for UI integration.

### What is the difference between `watchedEvents()` and `producedEvents()`?

`watchedEvents()` returns the event types your module wants to receive (your inputs), while `producedEvents()` declares what your module generates (your outputs). The controller uses these declarations to build the event routing graph—modules only receive events they explicitly watch.

### Can SpiderFoot modules run in parallel?

Yes. When `maxThreads` exceeds 1, modules execute their `threadWorker()` method in the shared `SpiderFootThreadPool` from [`spiderfoot/threadpool.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/threadpool.py). Storage modules typically run single-threaded. Use `self.poolExecute()` to submit additional background tasks with automatic module-name tracking.

### How does SpiderFoot prevent infinite event loops between modules?

The framework implements duplicate detection through `storeOnly` logic in `notifyListeners()` and stop-signal checking via `checkForStop()`. Events carry source traceability, and the controller monitors queue depths and cycle patterns to halt runaway processing.