How the SpiderFoot Web UI Communicates with the Backend Scan Engine: Architecture Deep Dive
SpiderFoot's web interface uses CherryPy HTTP endpoints, multiprocessing workers, and a shared SQLite database to coordinate between the frontend and the scan engine.
The open-source OSINT platform SpiderFoot separates its web interface from the heavy lifting of data collection. Understanding how these components communicate helps developers extend the tool or troubleshoot scan issues. This article breaks down the exact mechanisms used in the smicallef/spiderfoot repository, with references to specific files and methods.
Core Communication Architecture
SpiderFoot employs a three-tier architecture for UI-to-engine communication:
- HTTP layer: CherryPy serves AJAX endpoints that the JavaScript frontend consumes
- Process layer: Each scan spawns a separate
multiprocessing.Processworker - Data layer: SQLite acts as the single source of truth for scan state, results, and logs
The web UI never talks directly to a running scan. Instead, it creates records in the database and spawns processes that independently read from and write to that same database.
Frontend: JavaScript AJAX Calls to CherryPy Endpoints
The frontend code in spiderfoot/static/js/spiderfoot.js provides wrapper functions that execute HTTP requests. The sf.fetchData() utility handles GET and POST calls, while higher-level functions like sf.startscan() encapsulate specific API operations.
Key frontend functions include:
sf.fetchData()— generic AJAX wrapper at lines 24-33sf.startscan()— initiates a new scan via POST to/startscansf.scanDelete(),sf.rerunscan()— control existing scans
When a user clicks "Start Scan", the following JavaScript executes:
// spiderfoot/static/js/spiderfoot.js
sf.startscan = function(name, target, modules, types, usecase, callback) {
sf.fetchData(
docroot + "/startscan",
{
scanname: name,
scantarget: target,
modulelist: modules,
typelist: types,
usecase: usecase
},
callback
);
};
The docroot variable ensures the correct base URL, and responses are handled via callbacks that update the UI state.
Backend: CherryPy Handlers and Process Spawning
The SpiderFootWebUi class in sfwebui.py exposes HTTP endpoints using the @cherrypy.expose decorator. Each endpoint maps to a class method that validates input, interacts with the database, and returns JSON responses.
The /startscan Endpoint Implementation
The startscan() method at lines 374-418 of sfwebui.py demonstrates the full spawning workflow:
# sfwebui.py
@cherrypy.expose
def startscan(self, scanname, scantarget, modulelist, typelist, usecase):
# Validation and setup...
targetType = SpiderFootHelpers.targetTypeFromString(scantarget)
cfg = deepcopy(self.config) # isolate scan configuration
sf = SpiderFoot(cfg) # create engine instance
modlist = modulelist.split(',')
scanId = SpiderFootHelpers.genScanInstanceId()
# Spawn the scanner in its own process
p = mp.Process(
target=startSpiderFootScanner,
args=(self.loggingQueue, scanname, scanId,
scantarget, targetType, modlist, cfg)
)
p.daemon = True
p.start()
raise cherrypy.HTTPRedirect(f"{self.docroot}/scaninfo?id={scanId}")
Critical details in this implementation:
deepcopy(self.config)creates an isolated configuration snapshotmp.ProcesslaunchesstartSpiderFootScannerwith themultiprocessingmoduleself.loggingQueueis passed to enable log streaming back to the UI- The method redirects to
/scaninforather than returning JSON directly
Scan Engine: The Multiprocessing Worker
Once spawned, the worker process runs SpiderFootScanner from sfscan.py. This class manages a single scan lifecycle independently of the web server process.
Scanner Initialization and Database Setup
The SpiderFootScanner.__init__ method at lines 30-53 establishes database connectivity:
# sfscan.py
class SpiderFootScanner:
def __init__(self, scanName, scanId, targetValue,
targetType, moduleList, globalOpts, start=True):
self.__config = deepcopy(globalOpts)
self.__dbh = SpiderFootDb(self.__config) # database handle
self.__sf = SpiderFoot(self.__config) # core engine
self.__sf.dbh = self.__dbh # share DB connection
self.__scanId = scanId
self.__dbh.scanInstanceCreate(
self.__scanId,
self.__scanName,
self.__targetValue
)
# Module loading, DNS setup, proxy configuration...
The scanner registers itself via scanInstanceCreate() before beginning data collection. All discovered information flows through SpiderFootDb methods like:
scanResultEvent()— stores harvested data pointsscanLog()— records operational eventsscanConfigSet()— persists module configuration
The SQLite Database as Communication Hub
spiderfoot/db.py implements SpiderFootDb, the abstraction layer used by both UI handlers and scanner processes. This design pattern—database-as-message-bus—eliminates direct socket communication between processes.
Key methods in the database wrapper:
| Method | Used By | Purpose |
|---|---|---|
scanInstanceCreate() |
Scanner | Register new scan with metadata |
scanResultEvent() |
Scanner | Store discovered data events |
scanInstanceSet() |
Both | Update scan status (RUNNING, FINISHED, etc.) |
scanLog() |
Scanner | Append log entries |
scanConfigGet/Set() |
Both | Read/write module configuration |
The UI polls these tables via endpoints like /scaninfo, /search, and /eventtypes to render current scan state without blocking on active workers.
Real-Time Log Streaming via Multiprocessing Queue
Beyond database polling, SpiderFoot implements immediate log delivery through a multiprocessing.Queue. This enables the "Live Log" view in the web interface.
Queue setup occurs in SpiderFootWebUi.__init__ at lines 84-92:
# sfwebui.py
class SpiderFootWebUi:
def __init__(self, config):
# ...
self.loggingQueue = mp.Queue()
self.scanManager = None
The queue is passed to every spawned scanner. When the scan engine generates log events, they enter this queue. The CherryPy process consumes from the same queue to stream updates to connected browsers—typically via WebSocket-style long polling or periodic AJAX fetches.
API Endpoints for Scan Control
The SpiderFootWebUi class exposes additional endpoints for lifecycle management:
- GET
/scaninfo?id={scanId}— retrieves scan metadata and current status - POST
/search— queries accumulated results with filtering - POST
/scandelete?id={scanId}— terminates and removes a scan - POST
/stopscan?id={scanId}— signals graceful shutdown
These endpoints operate on database state rather than direct process manipulation. To stop a scan, the UI updates a status flag that the worker process checks periodically.
Why This Architecture Matters
The separation of concerns in SpiderFoot's communication design provides several operational benefits:
- Resilience: Web server restarts don't kill running scans
- Scalability: Multiple scans execute in parallel without GIL contention
- Isolation: Module crashes in one scan don't affect others or the UI
- Auditability: The database provides complete historical record of all activity
For developers extending SpiderFoot, this means that:
- New endpoints are added to
SpiderFootWebUiinsfwebui.py - New data types are stored via
SpiderFootDbmethods inspiderfoot/db.py - New scan logic runs within
SpiderFootScannerinsfscan.py
Summary
- The SpiderFoot web UI communicates with the backend through CherryPy HTTP endpoints, not direct function calls
- AJAX requests from
spiderfoot.jstrigger process spawning viamultiprocessing.Processinsfwebui.py - Each scan runs as an isolated worker process executing
SpiderFootScannerfromsfscan.py - SQLite serves as the shared state store, with
SpiderFootDbproviding the abstraction layer - Log streaming uses a
multiprocessing.Queuefor real-time updates without database polling - The database-as-message-bus pattern decouples the UI from long-running scan operations
Frequently Asked Questions
How does SpiderFoot handle multiple concurrent scans?
Each scan spawns a separate multiprocessing.Process with its own SpiderFootScanner instance. The processes share no memory except through the SQLite database and optional logging queue. The SpiderFootDb class manages connection pooling and transaction isolation.
Can the web UI restart without losing active scans?
Yes. Because scan state persists in SQLite and workers run as independent processes, restarting the CherryPy server does not terminate running scans. The UI simply reconnects to existing scan records on restart.
What protocol does SpiderFoot use for frontend-to-backend communication?
Standard HTTP/HTTPS with AJAX. The frontend uses vanilla JavaScript XMLHttpRequest wrappers in sf.fetchData(), not WebSockets. Real-time updates rely on polling endpoints or the multiprocessing logging queue feeding into the HTTP response stream.
Where is scan configuration stored during a scan?
Configuration is snapshotted via deepcopy(self.config) in SpiderFootWebUi.startscan() and passed to the worker process. The worker may additionally persist settings via SpiderFootDb.scanConfigSet(). This ensures scans use consistent settings even if global configuration changes mid-scan.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →