How SpiderFoot's Module Plugin Architecture Works: A Deep Dive into the Plugin Framework
SpiderFoot uses a lightweight plugin framework centered on the abstract SpiderFootPlugin base class, where every data-source module acts as an interchangeable component that communicates through a shared event-driven system.
The smicallef/spiderfoot open-source OSINT tool is built entirely around extensible modular architecture. This design lets security researchers add new intelligence sources without modifying core scan logic. Understanding how SpiderFoot's module plugin architecture works reveals why the tool remains one of the most flexible reconnaissance platforms available.
The Core Plugin Contract: SpiderFootPlugin
All SpiderFoot modules inherit from a single abstract base class defined in spiderfoot/plugin.py. This class establishes the mandatory interface that every module must implement.
Lifecycle Hooks
The base class provides standardized hooks that the scan controller invokes:
setup(sf, userOpts)– Initializes the module with the centralSpiderFootobject (sf) and user-supplied configuration optionsstart()– Signals the beginning of event processingthreadWorker()– Pulls events fromincomingEventQueuewhen running in multi-threaded modefinish()– Cleanup hook called when the scan terminates
Built-in Infrastructure
Modules receive ready-made functionality from SpiderFootPlugin:
- Logging utilities –
self.log,debug(),info(), anderror()methods automatically embed the current scan ID - Thread-pool integration – Access to
sharedThreadPoolfor concurrent execution; storage plugins run directly without threading - Event queues –
incomingEventQueue(events to consume) andoutgoingEventQueue(events to emit) - Output filtering –
__outputFilter__limits which event types propagate - Stop-signal handling –
checkForStop()monitors abort requests and error states
Module Structure and Required Methods
Each module lives as a Python file under modules/ (for example, modules/sfp_dnsresolve.py). Every module must define several components:
| Component | Purpose |
|---|---|
meta dictionary |
Human-readable metadata including name, summary, categories, and flags for UI display |
opts and optdescs |
Configuration options exposed in the web interface |
watchedEvents() |
Returns list of event types the module consumes (['DNS_NAME', 'IP_ADDRESS'] or ['*'] for all) |
producedEvents() |
Returns event types the module generates |
handleEvent(sfEvent) |
Core processing logic—most modules override this method |
Optional threadWorker() |
Custom worker when maxThreads > 1 |
Here is a complete minimal module implementation:
from spiderfoot import SpiderFootPlugin
class sfp_example(SpiderFootPlugin):
"""Example module that demonstrates the plugin API."""
meta = {
'name': 'Example',
'summary': 'Shows how a plugin works',
'categories': ['Demo'],
'flags': []
}
opts = {}
optdescs = {}
def watchedEvents(self):
return ['DOMAIN_NAME']
def producedEvents(self):
return ['EXAMPLE_RESULT']
def setup(self, sf, userOpts):
self.sf = sf # spiderfoot core object
self.opts = userOpts
def handleEvent(self, sfEvent):
# Do something with sfEvent.data …
result = f"processed {sfEvent.data}"
self.sf.emitEvent(result, "EXAMPLE_RESULT", sfEvent)
Event-Driven Data Flow
SpiderFoot's module plugin architecture operates on a publish-subscribe event model. The controller orchestrates all module communication:
- Scan initialization – The controller creates a
SpiderFootinstance and loads every module frommodules/ - Module setup – Each module is instantiated and
setup()receives the centralSpiderFootobject plus configuration - Root event seeding – A
SpiderFootEvent('ROOT', target)enters every module'sincomingEventQueue - Event processing – The module's
threadWorker()or directhandleEvent()processes events and callsself.sf.emitEvent()to generate new data - Event propagation –
notifyListeners()inSpiderFootPluginroutes events by checking each destination module'swatchedEvents(), prevents duplicates viastoreOnlylogic, and either queues for threaded listeners or callshandleEvent()directly - Scan completion – The FINISHED sentinel triggers
finish()hooks across all modules
Dynamic Module Discovery
SpiderFoot discovers plugins automatically by scanning the modules/ package for SpiderFootPlugin subclasses. The main entry point in sf.py imports each Python file, instantiates the class, and collects meta dictionaries for UI rendering.
Because discovery relies on Python's import system, adding a new module requires only placing a properly structured .py file into modules/—no registry edits or core modifications needed.
Thread Pool Coordination
Non-storage plugins share a global SpiderFootThreadPool implemented in spiderfoot/threadpool.py. When a module submits work via self.poolExecute(callback, *args), the task receives a name derived from the module (f"{self.__name__}_threadWorker").
This architecture enables:
- Concurrency limits – Enforced via
maxThreadsconfiguration - Work tracking – Monitor pending tasks with
self.sharedThreadPool.countQueuedTasks() - Graceful aborts – Queue clearing and stop signals propagate to all running modules
Practical Examples
Creating a Complete New Plugin
# modules/sfp_mynewsource.py
from spiderfoot import SpiderFootPlugin
class sfp_mynewsource(SpiderFootPlugin):
meta = {
'name': 'MyNewSource',
'summary': 'Fetches data from MyNewSource API',
'categories': ['API'],
'flags': ['requires_api_key']
}
def watchedEvents(self):
return ['IP_ADDRESS']
def producedEvents(self):
return ['MYNEW_DATA']
def setup(self, sf, userOpts):
self.sf = sf
self.opts = userOpts
self.api_key = self.opts.get('api_key')
def handleEvent(self, sfEvent):
ip = sfEvent.data
# API call implementation (omitted for brevity)
data = self._query_api(ip)
self.sf.emitEvent(data, 'MYNEW_DATA', sfEvent)
Save this as modules/sfp_mynewsource.py—SpiderFoot automatically loads it on next startup.
Emitting Events for Cross-Module Communication
def handleEvent(self, sfEvent):
# Generate new intelligence
domain = f"{sfEvent.data}.example.com"
# Broadcast to interested modules (DNS resolution, WHOIS, etc.)
self.sf.emitEvent(domain, "DOMAIN_NAME", sfEvent)
Key Source Files
| File | Role |
|---|---|
spiderfoot/plugin.py |
Abstract base class SpiderFootPlugin defining plugin contracts and event routing |
spiderfoot/threadpool.py |
Shared thread pool for concurrent module execution |
spiderfoot/__init__.py |
Exports SpiderFootPlugin, SpiderFootTarget, and other core classes |
modules/*.py |
Concrete plugin implementations (e.g., modules/sfp_dnsresolve.py) |
sf.py |
CLI entry point handling dynamic discovery and plugin wiring |
Summary
- SpiderFoot's module plugin architecture centers on the
SpiderFootPluginabstract base class inspiderfoot/plugin.py - Modules implement
watchedEvents(),producedEvents(),setup(), andhandleEvent()to participate in scans - The event-driven publish-subscribe model decouples data sources, letting modules react to findings from any other module
- Dynamic discovery via Python imports eliminates registration overhead—drop a file in
modules/and it works - Shared thread pooling in
spiderfoot/threadpool.pycontrols concurrency and enables clean scan abortion
Frequently Asked Questions
How do I create a new SpiderFoot module from scratch?
Subclass SpiderFootPlugin from spiderfoot/plugin.py, implement watchedEvents(), producedEvents(), setup(), and handleEvent(), then save your file to modules/sfp_yourname.py. SpiderFoot automatically discovers and loads it on next startup. Include a meta dictionary with name, summary, and categories for UI integration.
What is the difference between watchedEvents() and producedEvents()?
watchedEvents() returns the event types your module wants to receive (your inputs), while producedEvents() declares what your module generates (your outputs). The controller uses these declarations to build the event routing graph—modules only receive events they explicitly watch.
Can SpiderFoot modules run in parallel?
Yes. When maxThreads exceeds 1, modules execute their threadWorker() method in the shared SpiderFootThreadPool from spiderfoot/threadpool.py. Storage modules typically run single-threaded. Use self.poolExecute() to submit additional background tasks with automatic module-name tracking.
How does SpiderFoot prevent infinite event loops between modules?
The framework implements duplicate detection through storeOnly logic in notifyListeners() and stop-signal checking via checkForStop(). Events carry source traceability, and the controller monitors queue depths and cycle patterns to halt runaway processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →