How SpiderFoot Implements Rate Limiting Across Its Modules: A Deep Dive into the Thread-Pool Architecture
SpiderFoot uses a two-layer defense: a shared thread-pool that caps concurrent requests per module, combined with explicit HTTP 429 detection and error handling in individual modules.
SpiderFoot's open-source reconnaissance framework implements rate limiting through a sophisticated combination of cooperative throttling and reactive safeguards. Understanding how SpiderFoot handles rate limiting is essential for module developers and security researchers who need to scan aggressively without triggering API bans or exhausting service quotas. This article examines the complete rate-limiting architecture as implemented in the smicallef/spiderfoot repository.
The Core Mechanism: Per-Module Thread-Pool Throttling
The foundation of SpiderFoot's rate limiting lives in spiderfoot/threadpool.py. The SpiderFootThreadPool class provides a shared execution context where every module submits its network-bound tasks.
How the Thread Pool Enforces Concurrency Limits
When a module calls self.sharedThreadPool.submit(), the pool applies a hard cap on simultaneous operations. Here is the critical logic from the submit() method:
def submit(self, callback, *args, **kwargs):
taskName = kwargs.get('taskName', 'default')
maxThreads = kwargs.pop('maxThreads', 100)
while self.countQueuedTasks(taskName) >= maxThreads:
sleep(.01) # block until a worker slot is free
self.inputQueue(taskName).put((callback, args, kwargs))
The taskName parameter identifies which module owns the task. The maxThreads value—defaulting to 100 but overridden per module—determines how many concurrent slots that module may occupy. When the limit is reached, the submission thread sleeps for 10 milliseconds and retries.
Configuring Module-Level Concurrency
Each module declares its preferred parallelism in setup(). The maxThreads attribute, inherited from SpiderFootPlugin in spiderfoot/plugin.py, controls this behavior:
def setup(self, sf, userOpts):
self.maxThreads = 5 # Allow at most 5 parallel tasks for this module
self.sharedThreadPool = sf.sharedThreadPool
Most modules default maxThreads to 1, ensuring conservative, sequential requests. Modules interacting with high-capacity APIs may raise this value, but the global _maxthreads option in sf.py still bounds total workers across the entire scan.
Global Coordination via the Scan Controller
The main scan controller in sf.py instantiates a single SpiderFootThreadPool shared by all modules:
SpiderFootThreadPool(self.opts["_maxthreads"])
This design guarantees uniform enforcement. No matter how individual modules are configured, the total active worker count never exceeds the global limit. The controller passes this shared pool reference to every plugin during initialization, ensuring consistent rate-limiting semantics across heterogeneous data sources.
Reactive Safeguard: Detecting and Handling HTTP 429 Responses
When preventive throttling is insufficient, SpiderFoot modules react to explicit rate-limit signals from remote services.
The 429 Detection Pattern
Modules inspect HTTP response codes returned by SpiderFoot's helper functions. The pattern appears consistently across the codebase:
# Excerpt from modules/sfp_xforce.py
res = self.sf.fetch(url, headers=headers)
if res['code'] == '429':
self.error("Rate limit exceeded")
return # Abort further processing for this request
Similar logic exists in modules/sfp_zonefiles.py and other network-dependent modules:
# Excerpt from modules/sfp_zonefiles.py
if res['code'] == '429':
self.error("Told to go away by zonefiles.io")
return None
Error State Propagation and Scan-Level Effects
Calling self.error() triggers behavior defined in spiderfoot/plugin.py. The base class sets errorState, which signals the controller to skip further invocations of that module for the current scan. This prevents a rate-limited module from continuing to hammer a service that has already rejected it.
Some modules implement additional back-off. For example, sfp_binaryedge.py checks for both 429 and 500 responses and may insert deliberate delays to allow remote services recovery time.
Two-Layer Defense: Preventive and Reactive Working Together
SpiderFoot's rate limiting succeeds because these mechanisms complement each other:
| Layer | Purpose | Implementation |
|---|---|---|
| Preventive | Stop bursts before they happen | Thread-pool slot limits in spiderfoot/threadpool.py |
| Reactive | Respond to service feedback | 429 detection and error state handling in individual modules |
The preventive layer spreads load over time, preventing accidental request floods that would trigger remote rate limits. The reactive layer provides graceful degradation when limits are encountered anyway, preserving scan integrity and protecting API relationships.
Key Files and Their Roles
spiderfoot/threadpool.py— Shared thread pool withsubmit()throttling logicspiderfoot/plugin.py— BaseSpiderFootPluginclass providingmaxThreads,error(), anderrorStatesf.py— Scan controller that instantiates the global thread poolmodules/sfp_xforce.py— Representative module demonstrating 429 handlingmodules/sfp_zonefiles.py— Additional 429 detection examplemodules/sfp_binaryedge.py— Extended error handling with recovery pauses
Summary
- Thread-pool throttling in
spiderfoot/threadpool.pycaps concurrent requests per module using a blocking counter that sleeps whenmaxThreadsis reached - Global coordination through
sf.pyensures uniform enforcement via a single shared pool passed to all plugins - Per-module configuration via
maxThreadsinsetup()allows API-appropriate parallelism while respecting global limits - 429 detection in individual modules triggers error logging and module deactivation for the current scan
- Error state propagation through
SpiderFootPluginprevents continued requests to rate-limited services
Frequently Asked Questions
How do I increase the rate limit for a specific SpiderFoot module?
Set self.maxThreads to a higher value in the module's setup() method. The default is typically 1 for conservative sequential requests. The global _maxthreads option in your SpiderFoot configuration still bounds total workers across all modules.
What happens when a remote API returns HTTP 429 to a SpiderFoot module?
The module detects the 429 status code, calls self.error() with a descriptive message, and returns early. The base class sets errorState, causing the scan controller to skip further calls to that module for the remainder of the current scan.
Where is the global request concurrency limit configured?
The global limit is set via the _maxthreads option in your SpiderFoot configuration, used when instantiating SpiderFootThreadPool in sf.py. This value caps total active workers regardless of individual module maxThreads settings.
Can a completely rate-limited module resume operation mid-scan?
No. Once errorState is set via self.error(), the module is effectively disabled for that scan. This design prioritizes scan completion and API relationship preservation over exhaustive data collection from uncooperative services.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →