# How SpiderFoot's Thread Pool Manages Parallel Module Execution

> Discover how SpiderFoot's thread pool efficiently manages parallel module execution using worker threads, named queues, and configurable concurrency for faster threat intelligence gathering.

- Repository: [Steve Micallef/spiderfoot](https://github.com/smicallef/spiderfoot)
- Tags: internals
- Published: 2026-08-15

---

**SpiderFoot manages parallel module execution through the `SpiderFootThreadPool` class in [`spiderfoot/threadpool.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/threadpool.py), which creates a configurable pool of worker threads that process tasks from named queues, enforces per-module concurrency limits, and provides `submit()` and `map()` APIs for module developers.**

SpiderFoot is an open-source reconnaissance platform that automates the collection of intelligence about a target. A core challenge in its architecture is executing dozens of reconnaissance modules simultaneously without overwhelming system resources. The `SpiderFootThreadPool` solves this by implementing a sophisticated thread pool that isolates work by module while sharing a common worker pool across the entire scan.

## Thread Pool Creation and Initialization

The `SpiderFootThreadPool` class ([`spiderfoot/threadpool.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/threadpool.py), lines 34-51) initializes with a configurable maximum number of worker threads. The default is **100 threads**, though this can be tuned based on system capabilities and scan requirements.

```python
from spiderfoot import SpiderFootThreadPool

# Create a pool with 50 worker threads

pool = SpiderFootThreadPool(threads=50, name="reconPool")
pool.start()  # Spawns ThreadPoolWorker threads (lines 53-59)

```

The constructor establishes critical internal structures:
- **`inputQueues`** – dictionary mapping `taskName` to dedicated queue
- **`outputQueues`** – dictionary for storing results per module
- **threading lock** – for synchronization during queue operations

Calling `start()` spawns the requested number of `ThreadPoolWorker` threads. Each worker is a long-lived `threading.Thread` that continuously polls for work rather than being created per-task.

## Task Submission: submit() and map() APIs

Modules enqueue work through two primary methods, both supporting per-module concurrency caps.

### submit() for Fire-and-Forget Tasks

The `submit()` method (lines 17-34) places a tuple of `(callback, args, kwargs)` onto a module-specific input queue:

```python
def submit(self, callback, *args, taskName="default", **kwargs):
    # Enforces maxThreads limit before queuing (lines 28-31)

    # Blocks until space available if limit reached

```

Before queuing, `submit()` checks the module's `maxThreads` limit and **blocks until capacity is available** (lines 28-31). This prevents any single module from monopolizing the worker pool.

### map() for Iterable Processing

The `map()` method provides a higher-level interface for processing collections:

```python

# Example: Scan multiple targets concurrently

targets = ["example.com", "test.org", "sample.net"]

with SpiderFootThreadPool(threads=10, name="scanPool") as pool:
    for result in pool.map(
        scan_target,
        targets,
        taskName="urlScanner",
        saveResult=True,  # Store results in output queue

    ):
        print(result)  # Yields results as completed

```

## Per-Module Queue Isolation

Each `taskName` receives its **own input and optional output queue**, created lazily on first use (lines 55-66). This isolation provides several benefits:

- **Fair scheduling** – No module can starve others of queue space
- **Ordered results** – Output remains associated with the originating module
- **Independent limits** – `maxThreads` applies per module, not globally

The `countQueuedTasks()` helper (lines 35-53) enables modules to monitor their own concurrency by returning queued items plus currently-running workers for a given `taskName`.

## Worker Thread Execution Model

`ThreadPoolWorker.run()` implements a **round-robin scheduler** across all input queues (lines 41-47). For each iteration:

1. Poll all input queues in rotation
2. Mark worker as busy when job acquired
3. Extract `taskName` and `saveResult` flag
4. Execute the callback with provided arguments
5. Optionally store result in corresponding output queue
6. Handle exceptions, log errors, and continue processing

This design ensures that a single slow module cannot block workers indefinitely—remaining workers continue processing other queues.

## Result Retrieval and Context Manager Support

Results flow through dedicated output queues per `taskName`. Two retrieval patterns exist:

**Streaming with results():**

```python
for item in pool.results(taskName="dnsLookup", wait=True):
    process(item)  # Blocks until all work complete if wait=True

```

**Bulk collection via context manager (lines 72-84, 94-115):**

```python
with SpiderFootThreadPool(threads=20) as pool:
    pool.submit(dns_query, "example.com", taskName="dns")
    # shutdown() called automatically on exit

    # Returns dict of all collected results per taskName

```

The context manager automatically triggers `shutdown()`, which drains queues, stops all workers, and returns accumulated results.

## Shared Thread Pool Architecture

The main scan driver in [`sfscan.py`](https://github.com/smicallef/spiderfoot/blob/main/sfscan.py) (lines 213-218) creates a **single shared pool** for the entire scan:

```python

# sfscan.py lines 213-218

self.__sharedThreadPool = SpiderFootThreadPool(
    threads=self.__config['__threads'],
    name='sharedThreadPool'
)

```

Each loaded module receives this pool via `mod.setSharedThreadPool()` (lines 554-560 in [`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py)). This architecture ensures:

- **Resource efficiency** – One pool serves all modules instead of N separate pools
- **Global limits** – Total threads capped at scan configuration level
- **Module autonomy** – Each module still controls its own `maxThreads` constraint

## Complete Working Example

```python
from spiderfoot import SpiderFootThreadPool

def check_ip(ip):
    """Perform a network lookup with potential blocking."""
    import time
    time.sleep(0.1)  # Simulated I/O

    return f"{ip} is responsive"

# Process 100 IPs with controlled concurrency

ip_list = [f"192.168.1.{i}" for i in range(1, 101)]

with SpiderFootThreadPool(threads=20, name="ipRecon") as pool:
    for result in pool.map(
        check_ip,
        ip_list,
        taskName="ipChecker",
        saveResult=True,
        maxThreads=5  # Cap this module at 5 concurrent checks

    ):
        print(result)

# Context exit automatically collects all results

```

## Summary

- **`SpiderFootThreadPool`** in [`spiderfoot/threadpool.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/threadpool.py) provides the core concurrency mechanism
- **Named queues** isolate work per module while sharing a common worker pool
- **`maxThreads` enforcement** prevents individual modules from exhausting resources
- **Round-robin scheduling** ensures fair distribution of worker time across modules
- **Context manager support** simplifies lifecycle management and result collection
- **[`sfscan.py`](https://github.com/smicallef/spiderfoot/blob/main/sfscan.py)** creates one shared pool distributed to all loaded modules via [`spiderfoot/plugin.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/plugin.py)

## Frequently Asked Questions

### How does SpiderFoot prevent one module from blocking all workers?

The `ThreadPoolWorker.run()` method polls **all input queues in round-robin order** (lines 41-47). If one module's queue is full or its callbacks run slowly, workers continue checking other queues. Additionally, the `maxThreads` parameter caps concurrent execution per module, and `submit()` blocks when that limit is reached rather than flooding the queue.

### What is the default thread pool size and can it be changed?

The default is **100 worker threads**, set in the `SpiderFootThreadPool` constructor. This is configurable via the `threads=` parameter. The main scan driver in [`sfscan.py`](https://github.com/smicallef/spiderfoot/blob/main/sfscan.py) passes `self.__config['__threads']` to respect user-defined scan settings.

### Why use named queues instead of a single global queue?

Named queues (indexed by `taskName`) provide **module isolation** and **result ordering**. Each module gets distinct input and output queues created lazily on first use (lines 55-66). This ensures that results remain associated with their originating module and that `countQueuedTasks()` can accurately track per-module load.

### What happens to pending tasks when a scan completes?

When used as a context manager, exiting the `with` block triggers `shutdown()` (lines 94-115). This method drains both input and output queues, stops all worker threads, and returns a dictionary of collected results per `taskName`. For non-context usage, `shutdown()` must be called explicitly to ensure clean termination.