How Different SpiderFoot Target Types (IP, Domain, Email, ASN) Affect Scan Behavior
SpiderFoot uses target types to determine which modules can run, how relationships are validated, and how discovered data maps back to your original seed.
Every scan in SpiderFoot begins with a SpiderFootTarget object defined in spiderfoot/target.py. This object pairs a raw value (like 192.0.2.1 or example.com) with a target type that gates module execution and drives the matching logic used throughout the scan. Understanding how SpiderFoot target types affect scan behavior is essential for configuring effective reconnaissance workflows.
Core Target Types in SpiderFoot
SpiderFoot recognizes seven primary target types, each defined in SpiderFootTarget._validTypes at lines 23-25 of spiderfoot/target.py. The type is automatically detected via SpiderFootHelpers.targetTypeFromString() in spiderfoot/helpers.py (lines 25-37) using regex patterns.
IP_ADDRESS and IPV6_ADDRESS Targets
IP_ADDRESS and IPV6_ADDRESS targets trigger network-centric modules: port scanners, reverse DNS lookups, GeoIP resolution, and SSL/TLS certificate analyzers.
When you seed a scan with an IP, the engine:
- Enables modules that list IP types in their
__inputs__attribute - Stores the IP as the base target value
- Makes it available via
SpiderFootTarget.getAddresses()
from spiderfoot.target import SpiderFootTarget
from spiderfoot.helpers import SpiderFootHelpers
# Auto-detect and create IP target
target_type = SpiderFootHelpers.targetTypeFromString("198.51.100.42")
# Returns: 'IP_ADDRESS'
target = SpiderFootTarget("198.51.100.42", target_type)
print(target.getAddresses())
# → ['198.51.100.42']
The matches() method validates whether discovered IPs belong to the target scope. For plain IP targets, this is an exact match check.
NETBLOCK_OWNER and NETBLOCKV6_OWNER Targets
NETBLOCK_OWNER targets represent CIDR blocks (e.g., 192.0.2.0/24). The detection regex in spiderfoot/helpers.py lines 26-27 matches IPv4 netblocks with the pattern r'\d+\.\d+\.\d+\.\d+/\d+$' and IPv6 netblocks with r'[0-9a-f:]+::/[0-9]+$'.
Netblock targets unlock:
- ASN and routing table modules
- Netblock WHOIS lookups
- Bulk IP enumeration within the defined range
The critical difference is in SpiderFootTarget.matches() (lines 101-107 of spiderfoot/target.py), which uses netaddr.IPNetwork to test membership:
from spiderfoot.target import SpiderFootTarget
netblock = SpiderFootTarget("192.0.2.0/24", "NETBLOCK_OWNER")
# Test if discovered IP falls within target scope
if netblock.matches("192.0.2.45"):
print("IP is within target netblock") # ← prints
if netblock.matches("203.0.113.5"):
print("IP is within target netblock") # ← does not print
INTERNET_NAME Targets
INTERNET_NAME is SpiderFoot's designation for domain names. Detection uses the regex at spiderfoot/helpers.py lines 35-36: ^(([a-z0-9]([a-z0-9-]*[a-z0-9])?)\.)+([a-z0-9]([a-z0-9-]*[a-z0-9])?)$.
Domain targets activate the broadest module set:
- DNS resolution and enumeration
- Subdomain discovery
- Certificate transparency log monitoring
- Web framework and technology detection
The SpiderFootTarget.getNames() method (lines 28-33 of spiderfoot/target.py) aggregates the base domain plus any aliases. The matches() method supports flexible relationship testing:
from spiderfoot.target import SpiderFootTarget
domain = SpiderFootTarget("example.com", "INTERNET_NAME")
# Default: includeChildren=True treats subdomains as matches
print(domain.matches("sub.example.com")) # → True
# Can also match parent relationships
print(domain.matches("example.com", includeParents=True)) # → True
EMAILADDR Targets
EMAILADDR targets enable breach data modules, Gravatar lookups, and MX record analysis. Detected by the simple regex .*@.* at spiderfoot/helpers.py lines 28-29.
Email targets are stored directly in targetValue, and getNames() adds the address to the name list. This allows modules to pivot from email addresses to associated domains and infrastructure.
from spiderfoot.helpers import SpiderFootHelpers
# Auto-detect email target
email_type = SpiderFootHelpers.targetTypeFromString("admin@example.com")
print(email_type) # → 'EMAILADDR'
BGP_AS_OWNER Targets
BGP_AS_OWNER targets are autonomous system numbers, detected by the numeric-only regex ^[0-9]+$ at spiderfoot/helpers.py lines 32-33.
ASN targets run:
- ASN registration and organization lookups
- Associated netblock discovery
- Routing history analysis
The matches() method short-circuits for ASN targets (lines 90-94 of spiderfoot/target.py) because AS numbers lack direct network representation—discovered data relationships are validated through other means.
Loose Target Types: HUMAN_NAME, PHONE_NUMBER, USERNAME, BITCOIN_ADDRESS
HUMAN_NAME, PHONE_NUMBER, USERNAME, and BITCOIN_ADDRESS are treated as "loose" types. The matches() method returns True for any value when the target uses these types, since they don't map to concrete network boundaries.
These types are detected by specific regexes in targetTypeFromString():
HUMAN_NAME: quoted names ("John Doe")PHONE_NUMBER: international phone patternsUSERNAME: platform-specific handle formatsBITCOIN_ADDRESS: Base58Check-encoded addresses
How Target Types Drive Module Selection
The scan engine filters modules through a three-stage pipeline:
Stage 1: Target Instantiation
In sfscan.py lines 52-63, SfScan.__init__ wraps your input in a SpiderFootTarget:
# From sfscan.py (simplified)
from spiderfoot.target import SpiderFootTarget
self.target = SpiderFootTarget(targetValue, targetType)
Stage 2: Module Filtering
Each SpiderFoot plugin declares compatible inputs via self.__inputs__. In spiderfoot/plugin.py lines 334-336, the scanner builds the runnable set by testing:
# From spiderfoot/plugin.py (conceptual)
if module.__inputs__ and self.target.targetType not in module.__inputs__:
# Skip this module — incompatible target type
continue
Example module declarations:
# modules/sfp_dnsresolve.py — domain-only
class sfp_dnsresolve(SpiderFootPlugin):
__inputs__ = ["INTERNET_NAME"]
__outputs__ = ["IP_ADDRESS", "INTERNET_NAME"]
# modules/sfp_whois.py — dual compatibility
class sfp_whois(SpiderFootPlugin):
__inputs__ = ["IP_ADDRESS", "IPV6_ADDRESS", "INTERNET_NAME", "NETBLOCK_OWNER"]
Stage 3: Relationship Validation
As modules execute, they produce events. The SpiderFootTarget.matches() method (lines 57-84) determines if discovered data relates to your original target:
| Target Type | Match Logic |
|---|---|
IP_ADDRESS |
Exact IP match or alias match |
NETBLOCK_OWNER |
CIDR containment via netaddr |
INTERNET_NAME |
Exact domain, child subdomain, or parent domain (configurable) |
EMAILADDR |
Exact match or domain relationship |
BGP_AS_OWNER |
Always returns True (handled separately) |
Loose types (HUMAN_NAME, etc.) |
Always returns True |
Alias Expansion and Scan Scope
Modules can broaden the scan scope without changing the original target type by calling SpiderFootTarget.setAlias():
from spiderfoot.target import SpiderFootTarget
netblock = SpiderFootTarget("192.0.2.0/24", "NETBLOCK_OWNER")
# Module discovers related hostname
netblock.setAlias("www.example.com", "INTERNET_NAME")
# Now getNames() includes the alias
print(netblock.getNames()) # → ['www.example.com']
# Matches() considers both original and alias values
print(netblock.matches("192.0.2.10")) # → True (CIDR match)
Aliases feed back into getAddresses() and getNames(), allowing cross-type pivots while maintaining type-specific matching rules.
Practical Configuration Examples
Scanning a Corporate Netblock
from spiderfoot.target import SpiderFootTarget
# Target an owned /22
target = SpiderFootTarget("203.0.113.0/22", "NETBLOCK_OWNER")
# Matches any IP in 203.0.113.0 - 203.0.116.255
for ip in discovered_ips:
if target.matches(ip):
process_host(ip) # Within scope
Multi-Phase Domain Investigation
from spiderfoot.helpers import SpiderFootHelpers
from spiderfoot.target import SpiderFootTarget
# Phase 1: Domain reconnaissance
domain = "example.com"
target_type = SpiderFootHelpers.targetTypeFromString(domain) # 'INTERNET_NAME'
target = SpiderFootTarget(domain, target_type)
# Phase 2: After DNS resolution, add IP aliases
target.setAlias("93.184.216.34", "IP_ADDRESS")
# Phase 3: IP-specific modules now consider this target relevant
# via getAddresses() and matches() expansion
Summary
- SpiderFoot target types are defined in
SpiderFootTarget._validTypesand detected viatargetTypeFromString()regex patterns - Module gatekeeping happens through
__inputs__matching inspiderfoot/plugin.pylines 334-336 - Relationship validation uses type-specific logic in
SpiderFootTarget.matches(): CIDR containment for netblocks, domain hierarchy for names, exact matches for IPs and emails - Alias expansion via
setAlias()allows cross-type pivots without changing the base target type - Loose types (
HUMAN_NAME,PHONE_NUMBER, etc.) bypass strict matching and accept any discovered value
Frequently Asked Questions
How does SpiderFoot automatically detect target types?
SpiderFoot uses SpiderFootHelpers.targetTypeFromString() in spiderfoot/helpers.py (lines 25-37). This function applies ordered regex patterns to match IPs, netblocks, domains, emails, AS numbers, and other formats. The first matching pattern determines the type.
Can I force a specific target type instead of auto-detection?
Yes. While the CLI and web interface auto-detect, you can instantiate SpiderFootTarget directly with any valid type from SpiderFootTarget._validTypes. This overrides detection when working programmatically or building custom scan orchestration.
Why doesn't my module run against certain target types?
Each module declares compatible inputs in self.__inputs__. If your target type isn't listed, the scanner excludes the module during initialization at spiderfoot/plugin.py lines 334-336. Check the module source or use sf.py -M (module information) to view __inputs__ requirements.
How do netblock targets handle IPv6 address discovery?
NETBLOCKV6_OWNER uses the same matching logic as IPv4 via netaddr.IPNetwork, supporting CIDR notation like 2001:db8::/32. The matches() method at spiderfoot/target.py lines 101-107 automatically handles both address families when the target type specifies IPv6.
What happens when a module discovers data outside the original target scope?
Modules can still report external data as events. However, SpiderFootTarget.matches() returns False for out-of-scope values, affecting how the engine correlates findings and whether subsequent modules process those events as primary targets versus supplementary intelligence.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →