How to Monitor SpiderFoot Scan Progress Programmatically via Its REST API
Poll the /scanopts endpoint with a scan ID to retrieve real-time status, timestamps, and metadata as JSON without direct database access.
SpiderFoot's built-in web interface doubles as a REST API server that exposes scan state through lightweight HTTP endpoints. This article explains how to monitor SpiderFoot scan progress programmatically using the /scanopts endpoint, which queries the underlying SQLite database and returns structured JSON suitable for automation scripts and external monitoring systems.
The Core Endpoint: /scanopts
The primary interface for scan monitoring is implemented in sfwebui.py as the scanopts method. This endpoint accepts a single query parameter id containing the scan identifier and returns a JSON object with three top-level keys:
meta– array containing scan name, target, timestamps, and current statusconfig– the scan's module configurationconfigdesc– human-readable configuration descriptions
The method is decorated with @cherrypy.tools.json_out(), which ensures responses are automatically serialized to JSON with Content-Type: application/json headers. Internally, the endpoint delegates to SpiderFootDb.scanInstanceGet() in spiderfoot/db.py, which executes:
def scanInstanceGet(self, instanceId):
# Returns: (name, target, created, started, ended, status)
qry = "SELECT name, seed_target, created, started, ended, status \
FROM tbl_scan_instance WHERE guid = ?"
...
The status field in position 5 of the returned tuple contains one of: STARTING, RUNNING, FINISHED, ABORTED, or ERROR-FAILED.
Step-by-Step Monitoring Workflow
1. Obtain a Scan ID
Start a scan via /startscan (POST with scanname, scantarget, and optional modulelist) or launch through the CLI. The response contains the generated scan UUID. Example with curl:
curl -X POST "http://localhost:5001/startscan" \
-d "scanname=ExampleScan" \
-d "scantarget=example.com" \
-d "modulelist=sfp_dnsresolve,sfp_portscan_tcp"
Parse the JSON response to extract id.
2. Poll for Status Changes
Repeatedly GET /scanopts?id=<scan_id> until the scan reaches a terminal state. The meta array uses fixed indices:
| Index | Field | Type |
|---|---|---|
| 0 | Scan name | string |
| 1 | Target | string |
| 2 | Created timestamp | integer (epoch) |
| 3 | Started datetime | string |
| 4 | Ended datetime | string or "Not yet" |
| 5 | Status | string |
3. Parse Timestamps for Progress Metrics
The started and ended fields allow elapsed time calculation. When ended equals "Not yet", the scan is still active.
Complete Python Implementation
This production-ready script demonstrates polling with exponential backoff and proper error handling:
import time
import sys
import requests
from datetime import datetime
BASE_URL = "http://localhost:5001" # Adjust to your SpiderFoot server
POLL_INTERVAL = 5 # Seconds between checks
MAX_RETRIES = 3
class ScanMonitor:
def __init__(self, base_url: str):
self.base_url = base_url.rstrip("/")
self.session = requests.Session()
def get_status(self, scan_id: str) -> dict:
"""Fetch scan metadata from /scanopts endpoint."""
url = f"{self.base_url}/scanopts"
for attempt in range(MAX_RETRIES):
try:
resp = self.session.get(url, params={"id": scan_id}, timeout=30)
resp.raise_for_status()
return resp.json()
except requests.RequestException as e:
if attempt == MAX_RETRIES - 1:
raise RuntimeError(f"Failed to query scan status: {e}")
time.sleep(2 ** attempt) # Exponential backoff
return {} # Unreachable
def parse_meta(self, data: dict) -> dict:
"""Extract structured fields from meta array."""
meta = data.get("meta", [])
if len(meta) < 6:
raise ValueError(f"Unexpected meta format: {meta}")
return {
"name": meta[0],
"target": meta[1],
"created": meta[2],
"started": meta[3],
"ended": meta[4],
"status": meta[5],
}
def monitor(self, scan_id: str) -> dict:
"""Block until scan completes, yielding progress updates."""
print(f"Monitoring scan {scan_id}...")
while True:
data = self.get_status(scan_id)
parsed = self.parse_meta(data)
status = parsed["status"]
started = parsed["started"]
target = parsed["target"]
# Calculate runtime
if started != "Not started":
start_dt = datetime.strptime(started, "%Y-%m-%d %H:%M:%S")
elapsed = datetime.now() - start_dt
elapsed_str = str(elapsed).split(".")[0] # Trim microseconds
else:
elapsed_str = "N/A"
print(f"[{elapsed_str}] {target}: {status}")
if status in ("FINISHED", "ABORTED", "ERROR-FAILED"):
return parsed
time.sleep(POLL_INTERVAL)
if __name__ == "__main__":
if len(sys.argv) != 2:
print(f"Usage: {sys.argv[0]} <scan_id>")
sys.exit(1)
monitor = ScanMonitor(BASE_URL)
final = monitor.monitor(sys.argv[1])
print(f"\nFinal status: {final['status']}")
sys.exit(0 if final["status"] == "FINISHED" else 1)
Run with: python monitor.py a1b2c3d4-e5f6-7890-abcd-ef1234567890
Alternative: Query Endpoint for Raw Data
For advanced use cases, /query accepts raw SQL against the SQLite database (requires appropriate permissions). To fetch result counts during a scan:
curl "http://localhost:5001/query?query=\
SELECT type, COUNT(*) FROM tbl_scan_results \
WHERE scan_instance_id='YOUR_SCAN_ID' GROUP BY type"
Security note: The /query endpoint exposes the full database schema. Restrict access in production deployments.
Key Implementation Files
| File | Purpose | Critical Function/Line |
|---|---|---|
sfwebui.py |
CherryPy web server and REST routing | scanopts method at line 747 [source] |
spiderfoot/db.py |
SQLite abstraction and schema management | scanInstanceGet() at line 719 [source] |
sf.py |
CLI entry point and server initialization | Instantiates SpiderFootWebUi class |
sflib.py |
Core scanning engine | startSpiderFootScanner creates scan records |
Status State Machine
Understanding state transitions helps design robust monitoring logic:
STARTING → RUNNING → FINISHED
↓ ↓
ABORTED ERROR-FAILED
- STARTING: Scan record created, modules loading
- RUNNING: Active data collection in progress
- FINISHED: Normal completion, results available
- ABORTED: User-initiated cancellation
- ERROR-FAILED: Unrecoverable module or system error
Summary
- Use
/scanopts?id=<scan_id>as the canonical endpoint for SpiderFoot scan progress monitoring via REST API - Parse the
metaarray at fixed indices—status is always position 5 - Poll every 5-10 seconds with exponential backoff for network resilience
- Terminate polling when status reaches
FINISHED,ABORTED, orERROR-FAILED - Query
sfwebui.pyandspiderfoot/db.pysource files to understand data flow and extend functionality
Frequently Asked Questions
Does SpiderFoot support WebSocket push notifications for scan status?
No. According to the sfwebui.py implementation, SpiderFoot uses a polling-based REST architecture. The CherryPy server does not implement WebSocket endpoints for real-time status streaming. For low-latency updates, reduce polling intervals to 1-2 seconds or implement a local event listener using SpiderFoot's Python API directly.
Can I monitor multiple scans simultaneously?
Yes. The /scanopts endpoint is stateless—issue parallel requests for different scan IDs. Each call executes an independent SQLite query via scanInstanceGet(). No session binding exists between requests, though the underlying database connection pool may limit concurrent throughput.
What authentication does the REST API require?
SpiderFoot's web interface uses HTTP Basic Auth or session cookies configured during startup. Pass credentials in your HTTP client:
requests.get(url, auth=("username", "password"))
Without authentication, endpoints return 401 Unauthorized. The --no-auth CLI flag disables this requirement (not recommended for production).
How do I calculate scan completion percentage?
SpiderFoot does not expose percentage-complete metrics natively. Approximate progress by comparing elapsed time against typical scan duration for your target and module set, or count results returned via /query against expected thresholds. The meta array contains no total-work estimator.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →