How SpiderFoot's Thread Pool Manages Parallel Module Execution
SpiderFoot manages parallel module execution through the SpiderFootThreadPool class in spiderfoot/threadpool.py, which creates a configurable pool of worker threads that process tasks from named queues, enforces per-module concurrency limits, and provides submit() and map() APIs for module developers.
SpiderFoot is an open-source reconnaissance platform that automates the collection of intelligence about a target. A core challenge in its architecture is executing dozens of reconnaissance modules simultaneously without overwhelming system resources. The SpiderFootThreadPool solves this by implementing a sophisticated thread pool that isolates work by module while sharing a common worker pool across the entire scan.
Thread Pool Creation and Initialization
The SpiderFootThreadPool class (spiderfoot/threadpool.py, lines 34-51) initializes with a configurable maximum number of worker threads. The default is 100 threads, though this can be tuned based on system capabilities and scan requirements.
from spiderfoot import SpiderFootThreadPool
# Create a pool with 50 worker threads
pool = SpiderFootThreadPool(threads=50, name="reconPool")
pool.start() # Spawns ThreadPoolWorker threads (lines 53-59)
The constructor establishes critical internal structures:
inputQueues– dictionary mappingtaskNameto dedicated queueoutputQueues– dictionary for storing results per module- threading lock – for synchronization during queue operations
Calling start() spawns the requested number of ThreadPoolWorker threads. Each worker is a long-lived threading.Thread that continuously polls for work rather than being created per-task.
Task Submission: submit() and map() APIs
Modules enqueue work through two primary methods, both supporting per-module concurrency caps.
submit() for Fire-and-Forget Tasks
The submit() method (lines 17-34) places a tuple of (callback, args, kwargs) onto a module-specific input queue:
def submit(self, callback, *args, taskName="default", **kwargs):
# Enforces maxThreads limit before queuing (lines 28-31)
# Blocks until space available if limit reached
Before queuing, submit() checks the module's maxThreads limit and blocks until capacity is available (lines 28-31). This prevents any single module from monopolizing the worker pool.
map() for Iterable Processing
The map() method provides a higher-level interface for processing collections:
# Example: Scan multiple targets concurrently
targets = ["example.com", "test.org", "sample.net"]
with SpiderFootThreadPool(threads=10, name="scanPool") as pool:
for result in pool.map(
scan_target,
targets,
taskName="urlScanner",
saveResult=True, # Store results in output queue
):
print(result) # Yields results as completed
Per-Module Queue Isolation
Each taskName receives its own input and optional output queue, created lazily on first use (lines 55-66). This isolation provides several benefits:
- Fair scheduling – No module can starve others of queue space
- Ordered results – Output remains associated with the originating module
- Independent limits –
maxThreadsapplies per module, not globally
The countQueuedTasks() helper (lines 35-53) enables modules to monitor their own concurrency by returning queued items plus currently-running workers for a given taskName.
Worker Thread Execution Model
ThreadPoolWorker.run() implements a round-robin scheduler across all input queues (lines 41-47). For each iteration:
- Poll all input queues in rotation
- Mark worker as busy when job acquired
- Extract
taskNameandsaveResultflag - Execute the callback with provided arguments
- Optionally store result in corresponding output queue
- Handle exceptions, log errors, and continue processing
This design ensures that a single slow module cannot block workers indefinitely—remaining workers continue processing other queues.
Result Retrieval and Context Manager Support
Results flow through dedicated output queues per taskName. Two retrieval patterns exist:
Streaming with results():
for item in pool.results(taskName="dnsLookup", wait=True):
process(item) # Blocks until all work complete if wait=True
Bulk collection via context manager (lines 72-84, 94-115):
with SpiderFootThreadPool(threads=20) as pool:
pool.submit(dns_query, "example.com", taskName="dns")
# shutdown() called automatically on exit
# Returns dict of all collected results per taskName
The context manager automatically triggers shutdown(), which drains queues, stops all workers, and returns accumulated results.
Shared Thread Pool Architecture
The main scan driver in sfscan.py (lines 213-218) creates a single shared pool for the entire scan:
# sfscan.py lines 213-218
self.__sharedThreadPool = SpiderFootThreadPool(
threads=self.__config['__threads'],
name='sharedThreadPool'
)
Each loaded module receives this pool via mod.setSharedThreadPool() (lines 554-560 in spiderfoot/plugin.py). This architecture ensures:
- Resource efficiency – One pool serves all modules instead of N separate pools
- Global limits – Total threads capped at scan configuration level
- Module autonomy – Each module still controls its own
maxThreadsconstraint
Complete Working Example
from spiderfoot import SpiderFootThreadPool
def check_ip(ip):
"""Perform a network lookup with potential blocking."""
import time
time.sleep(0.1) # Simulated I/O
return f"{ip} is responsive"
# Process 100 IPs with controlled concurrency
ip_list = [f"192.168.1.{i}" for i in range(1, 101)]
with SpiderFootThreadPool(threads=20, name="ipRecon") as pool:
for result in pool.map(
check_ip,
ip_list,
taskName="ipChecker",
saveResult=True,
maxThreads=5 # Cap this module at 5 concurrent checks
):
print(result)
# Context exit automatically collects all results
Summary
SpiderFootThreadPoolinspiderfoot/threadpool.pyprovides the core concurrency mechanism- Named queues isolate work per module while sharing a common worker pool
maxThreadsenforcement prevents individual modules from exhausting resources- Round-robin scheduling ensures fair distribution of worker time across modules
- Context manager support simplifies lifecycle management and result collection
sfscan.pycreates one shared pool distributed to all loaded modules viaspiderfoot/plugin.py
Frequently Asked Questions
How does SpiderFoot prevent one module from blocking all workers?
The ThreadPoolWorker.run() method polls all input queues in round-robin order (lines 41-47). If one module's queue is full or its callbacks run slowly, workers continue checking other queues. Additionally, the maxThreads parameter caps concurrent execution per module, and submit() blocks when that limit is reached rather than flooding the queue.
What is the default thread pool size and can it be changed?
The default is 100 worker threads, set in the SpiderFootThreadPool constructor. This is configurable via the threads= parameter. The main scan driver in sfscan.py passes self.__config['__threads'] to respect user-defined scan settings.
Why use named queues instead of a single global queue?
Named queues (indexed by taskName) provide module isolation and result ordering. Each module gets distinct input and output queues created lazily on first use (lines 55-66). This ensures that results remain associated with their originating module and that countQueuedTasks() can accurately track per-module load.
What happens to pending tasks when a scan completes?
When used as a context manager, exiting the with block triggers shutdown() (lines 94-115). This method drains both input and output queues, stops all worker threads, and returns a dictionary of collected results per taskName. For non-context usage, shutdown() must be called explicitly to ensure clean termination.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →