How to Optimize SpiderFoot Scan Performance for Large Targets: Thread Pool Tuning and DNS Optimization Guide
Increase the global _maxthreads limit above the default of 3 and configure per-module concurrency caps to eliminate I/O bottlenecks when scanning large targets with SpiderFoot.
SpiderFoot implements a shared thread pool architecture that controls all concurrent operations during reconnaissance. For large targets—domains with extensive subdomains, disparate IP ranges, or heavy API utilization—the default configuration creates significant I/O wait times that throttle scan throughput. Understanding the concurrency controls in sfscan.py and spiderfoot/threadpool.py enables precise performance tuning without destabilizing the scanner.
Understanding SpiderFoot's Concurrency Architecture
SpiderFoot distributes work through three interconnected layers:
| Component | Default Value | Location |
|---|---|---|
Global thread pool (_maxthreads) |
3 threads | sf.py line 56; sfscan.py line 213 |
| Per-module thread limits | Varies (often 100) | Module option dictionaries (e.g., sfp_dnsbrute.py lines 40-51) |
Queue size (qsize) |
10 | SpiderFootThreadPool.__init__ lines 34-45 |
The SpiderFootThreadPool class in spiderfoot/threadpool.py maintains persistent worker threads throughout a scan lifetime, avoiding the overhead of thread creation per task. Workers pull from module-specific inputQueue instances and return results via outputQueue.
The event distribution loop in sfscan.py (lines 475-540) coordinates completion through waitForThreads, which monitors threadsFinished status with 0.1-second polling intervals to prevent CPU-intensive busy-waiting.
Why Default Settings Fail on Large Targets
Large targets generate bursty, high-volume I/O workloads: DNS brute-force enumeration, concurrent WHOIS queries, and parallel API requests to services like Shodan or VirusTotal. With only 3 global threads, the scanner serializes these operations, leaving network interfaces underutilized while workers block on responses.
The SpiderFootThreadPool instantiation in sfscan.py demonstrates this constraint explicitly:
# sfscan.py lines 213-214
self.threadpool = SpiderFootThreadPool(
threads=self.__config.get("_maxthreads", 3),
...
)
Raising this value parallelizes the I/O pipeline, though excessive concurrency risks API rate limiting and memory pressure.
Step-by-Step Performance Optimization
1. Increase the Global Thread Pool
The global _maxthreads parameter governs all module execution. Adjust via CLI flag or programmatic configuration.
CLI approach:
python sf.py -t large-target.com \
--max-threads 25 \
--modules sfp_dnsbrute,sfp_whois,sfp_shodan
Programmatic approach:
from sfscan import SpiderFootScanner
cfg = {
'_maxthreads': 25, # Override default of 3
# ... other options
}
scanner = SpiderFootScanner(
scanName='optimized-scan',
scanId='scan-001',
targetValue='large-target.com',
targetType='DOMAIN_NAME',
moduleList=['sfp_dnsbrute', 'sfp_whois', 'sfp_shodan'],
globalOpts=cfg,
start=True
)
The scanner passes this value directly to SpiderFootThreadPool.__init__ as the threads parameter.
2. Tune Module-Specific Concurrency Limits
Many I/O-intensive modules enforce internal _maxthreads caps independent of the global pool. The DNS brute-force module exemplifies this pattern:
# modules/sfp_dnsbrute.py lines 40-52 (simplified)
opts = {
'_maxthreads': 100, # Module-specific default
# ...
}
def query(self, qry):
# Respects self.opts['_maxthreads'] for concurrent lookups
...
Override these limits when the global pool increase exceeds module defaults:
python sf.py -t large-target.com \
--max-threads 30 \
--module-opts "sfp_dnsbrute._maxthreads=250,sfp_tldsearch._maxthreads=200"
Or via configuration dictionary:
cfg['__modules__'] = {
'sfp_dnsbrute': {'opts': {'_maxthreads': 250}},
'sfp_tldsearch': {'opts': {'_maxthreads': 200}},
}
3. Optimize DNS Resolution Performance
DNS latency dominates domain-heavy scans. SpiderFoot provides two optimization mechanisms in sfscan.py lines 87-94:
- Resolver override: Replace the system resolver with a specified server
- Persistent resolver instance: Reuse connections across lookups
Configuration example:
cfg = {
'_maxthreads': 30,
'_dnsserver': '1.1.1.1', # Cloudflare DNS
}
This invokes dns.resolver.override_system_resolver() during scanner initialization, directing all DNS queries through the high-performance resolver.
4. Warm the TLD Cache
SpiderFoot downloads the Internet TLD list once per scan unless cached. In sfscan.py lines 99-107, the scanner checks _internettlds_cache for a valid local copy before network retrieval.
Pre-populate this cache to eliminate startup latency:
- Ensure
_internettlds_cachepoints to a writable, persistent path - Verify the cached file isn't expired (SpiderFoot validates freshness)
- For air-gapped environments, manually populate the cache file
5. Adjust Queue Sizes for Bursty Workloads
The SpiderFootThreadPool constructor accepts a qsize parameter (default 10) that bounds per-module queues. For modules generating rapid event bursts—particularly sfp_dnsbrute with large wordlists—queue saturation causes blocking.
Modify queue depth by editing the pool instantiation in sfscan.py (advanced use):
# Requires source modification; no CLI exposure
self.threadpool = SpiderFootThreadPool(
threads=self.__config.get("_maxthreads", 3),
qsize=50, # Increased from default 10
...
)
Monitor for "Queue full" warnings in scan output to identify this bottleneck.
Balancing Performance Against Constraints
API Rate Limiting
External modules interacting with rate-limited services (VirusTotal, Shodan, Censys) require asymmetric tuning: maintain high global thread counts for local operations while restricting specific modules. Use module-specific _maxthreads as a throttle:
| Module | Typical Limit | Rationale |
|---|---|---|
sfp_shodan |
1-3 threads | API key request quotas |
sfp_virustotal |
4 threads | Daily lookup limits |
sfp_dnsbrute |
200+ threads | Local DNS, no external limits |
Memory and CPU Boundaries
Each thread consumes approximately 8MB stack space (Python default). A 100-thread configuration requires ~800MB resident memory before accounting for SpiderFoot's data structures. Recommended ceilings:
- Small VPS (2GB RAM): Maximum 20-30 global threads
- Medium server (8GB RAM): 50-75 threads viable
- Large infrastructure (32GB+ RAM): 100+ threads with monitoring
Complete Optimized Configuration Example
#!/bin/bash
# High-performance scan for enterprise-scale target
python sf.py \
-t corp-target.com \
--max-threads 40 \
--modules sfp_dnsbrute,sfp_tldsearch,sfp_whois,sfp_shodan,sfp_censys \
--module-opts "\
sfp_dnsbrute._maxthreads=300,\
sfp_tldsearch._maxthreads=250,\
sfp_shodan._maxthreads=2,\
sfp_censys._maxthreads=2" \
-o json \
> scan-results.json
Programmatic equivalent with full configuration:
from sfscan import SpiderFootScanner
cfg = {
'_maxthreads': 40,
'_dnsserver': '9.9.9.9', # Quad9 resolver
'_internettlds_cache': '/var/cache/spiderfoot/tlds.cache',
'__modules__': {
'sfp_dnsbrute': {'opts': {'_maxthreads': 300}},
'sfp_tldsearch': {'opts': {'_maxthreads': 250}},
'sfp_shodan': {'opts': {'_maxthreads': 2, '_apikey': '...'}},
'sfp_censys': {'opts': {'_maxthreads': 2, '_apikey': '...'}},
'sfp_whois': {'opts': {}},
},
}
scanner = SpiderFootScanner(
scanName='enterprise-recon',
scanId='ent-2024-001',
targetValue='corp-target.com',
targetType='DOMAIN_NAME',
moduleList=list(cfg['__modules__'].keys()),
globalOpts=cfg,
start=True
)
scanner.waitForThreads() # Blocks until completion
Key Source Files for Deep Customization
| File | Critical Function |
|---|---|
spiderfoot/threadpool.py |
SpiderFootThreadPool class; worker lifecycle and queue management |
sfscan.py |
SpiderFootScanner orchestration; pool instantiation, DNS setup, event loop |
sf.py |
Default configuration values; CLI argument parsing |
modules/sfp_dnsbrute.py |
Reference implementation of per-module _maxthreads |
modules/sfp_tldsearch.py |
Additional DNS-heavy module with concurrency controls |
Summary
- Raise global
_maxthreadsfrom 3 to 20-50+ based on target size and infrastructure capacity - Override module-specific limits for I/O-intensive modules (
sfp_dnsbrute,sfp_tldsearch) while throttling API-dependent modules - Configure fast DNS resolver via
_dnsserverto eliminate resolution latency - Ensure TLD cache availability to prevent redundant network fetches
- Monitor queue saturation and memory utilization when scaling thread counts
Parallelism tuning in SpiderFoot operates at two distinct levels: the global SpiderFootThreadPool governing cross-module execution and per-module _maxthreads controlling internal concurrency. Effective optimization coordinates both layers while respecting external API constraints and system resource limits.
Frequently Asked Questions
What is the default thread pool size in SpiderFoot and why is it so low?
The default _maxthreads value is 3, defined in sf.py line 56. This conservative setting prioritizes stability and API safety over raw performance, ensuring new users don't immediately encounter rate limits or memory exhaustion. Production deployments scanning large targets should increase this value substantially based on available infrastructure and target characteristics.
How do I know if my scan is thread-bound or network-bound?
Thread-bound scans exhibit low network utilization (monitored via iftop, nload, or cloud provider metrics) with sustained CPU activity in the SpiderFoot process. Network-bound scans show high interface throughput with threads frequently idle in I/O wait states. Increase _maxthreads until network saturation or API errors appear, then back off slightly.
Can module-specific thread limits exceed the global pool size?
Yes. Per-module _maxthreads controls how many concurrent operations a module attempts to queue, while the global pool determines how many execute simultaneously. A module with _maxthreads=500 and a global pool of 40 will queue 500 tasks but execute only 40 concurrently. This configuration benefits bursty, fast operations like DNS brute-forcing that complete quickly once scheduled.
Does increasing threads affect result accuracy or cause missed findings?
No. SpiderFoot's event-driven architecture in sfscan.py ensures thread-safe event distribution regardless of concurrency level. However, excessive parallelism against rate-limited APIs causes transient failures that may require re-scanning. Implement module-specific throttling for external services to preserve result completeness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →