How to Configure the Maximum Concurrent Threads for SpiderFoot Module Execution

Set the _maxthreads parameter via CLI flag (-max-threads), configuration file, or the default config dictionary in sf.py to control how many SpiderFoot modules run simultaneously.

SpiderFoot executes its scanning modules in parallel using a shared thread pool, and the maximum number of concurrent threads directly impacts scan performance and system resource consumption. This guide explains three methods to configure this setting based on the smicallef/spiderfoot source code implementation.

Method 1: Use the Command-Line Flag

The simplest way to override the default thread limit for a single scan is the -max-threads CLI argument defined in sf.py at line 115.

python sf.py -max-threads 20 -s example.com

This runs the scan with up to 20 modules executing simultaneously, regardless of the default configuration.

Method 2: Modify the Default Configuration in sf.py

The built-in default is stored in the sfConfig dictionary inside sf.py at line 56.


# sf.py - default configuration

sfConfig = {
    # ... other options ...

    '_maxthreads': 3,   # Default: 3 concurrent threads

    # ...

}

Change this value to set a new default for all future scans run from this installation.

Method 3: Set _maxthreads in a Custom Configuration File

For reusable, environment-specific settings, create or edit a configuration file:


# sf.cfg

_maxthreads = 15

Then reference it when launching:

python sf.py -c sf.cfg -s example.com

SpiderFoot loads this file and applies the _maxthreads value throughout the scan lifecycle.

How the Thread Limit Is Enforced Internally

The actual thread pool creation occurs in sfscan.py at line 213, where the _maxthreads value from your selected configuration method is passed to SpiderFootThreadPool:


# sfscan.py – thread pool instantiation

self.__sharedThreadPool = SpiderFootThreadPool(
    threads=self.__config.get("_maxthreads", 3),
    name='sharedThreadPool')

All modules that support threading submit work to this shared instance, meaning the _maxthreads value caps total concurrent activity across the entire scan, not per-module behavior.

Programmatic Configuration Example

When embedding SpiderFoot in another Python script, adjust the thread count directly:

import sf

# Load default configuration and override

cfg = sf.sfConfig.copy()
cfg['_maxthreads'] = 25

# Initialize and start scan with modified config

scanner = sf.sfScan(target='example.com', sfConfig=cfg)
scanner.start()

Key Source Files

File Purpose
sf.py Defines _maxthreads default, CLI flag (-max-threads), and configuration schema
sfscan.py Creates the shared SpiderFootThreadPool using the configured thread count
threadpool.py Implements SpiderFootThreadPool class that manages worker threads and enforces the limit

Summary

  • Primary control: The _maxthreads configuration option governs SpiderFoot's module concurrency
  • Three configuration methods: CLI flag (-max-threads), sf.py default dictionary, or external config file
  • Implementation location: sfscan.py line 213 initializes the thread pool with your specified value
  • Scope: The limit applies globally to all threaded modules within a single scan

Frequently Asked Questions

What is the default maximum concurrent threads in SpiderFoot?

The default value is 3 threads, defined in the sfConfig dictionary at sf.py line 56. This conservative default minimizes resource impact on standard hardware.

Does increasing _maxthreads always improve scan speed?

Not necessarily. Higher values increase CPU and memory utilization and may trigger rate limiting from target services. The optimal setting depends on network bandwidth, target sensitivity, and local system capacity.

Can different scans use different thread limits simultaneously?

Yes. Each sfScan instance creates its own SpiderFootThreadPool with the thread count specified in its configuration. Concurrent scans with different _maxthreads values operate independently.

Where is the thread pool actually implemented?

The SpiderFootThreadPool class resides in threadpool.py, which provides the underlying thread management and task queueing mechanisms used by sfscan.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →