How the SpiderFoot Web UI Communicates with the Backend Scan Engine: Architecture Deep Dive

SpiderFoot's web interface uses CherryPy HTTP endpoints, multiprocessing workers, and a shared SQLite database to coordinate between the frontend and the scan engine.

The open-source OSINT platform SpiderFoot separates its web interface from the heavy lifting of data collection. Understanding how these components communicate helps developers extend the tool or troubleshoot scan issues. This article breaks down the exact mechanisms used in the smicallef/spiderfoot repository, with references to specific files and methods.

Core Communication Architecture

SpiderFoot employs a three-tier architecture for UI-to-engine communication:

  • HTTP layer: CherryPy serves AJAX endpoints that the JavaScript frontend consumes
  • Process layer: Each scan spawns a separate multiprocessing.Process worker
  • Data layer: SQLite acts as the single source of truth for scan state, results, and logs

The web UI never talks directly to a running scan. Instead, it creates records in the database and spawns processes that independently read from and write to that same database.

Frontend: JavaScript AJAX Calls to CherryPy Endpoints

The frontend code in spiderfoot/static/js/spiderfoot.js provides wrapper functions that execute HTTP requests. The sf.fetchData() utility handles GET and POST calls, while higher-level functions like sf.startscan() encapsulate specific API operations.

Key frontend functions include:

  • sf.fetchData() — generic AJAX wrapper at lines 24-33
  • sf.startscan() — initiates a new scan via POST to /startscan
  • sf.scanDelete(), sf.rerunscan() — control existing scans

When a user clicks "Start Scan", the following JavaScript executes:

// spiderfoot/static/js/spiderfoot.js
sf.startscan = function(name, target, modules, types, usecase, callback) {
    sf.fetchData(
        docroot + "/startscan",
        {
            scanname: name,
            scantarget: target,
            modulelist: modules,
            typelist: types,
            usecase: usecase
        },
        callback
    );
};

The docroot variable ensures the correct base URL, and responses are handled via callbacks that update the UI state.

Backend: CherryPy Handlers and Process Spawning

The SpiderFootWebUi class in sfwebui.py exposes HTTP endpoints using the @cherrypy.expose decorator. Each endpoint maps to a class method that validates input, interacts with the database, and returns JSON responses.

The /startscan Endpoint Implementation

The startscan() method at lines 374-418 of sfwebui.py demonstrates the full spawning workflow:


# sfwebui.py

@cherrypy.expose
def startscan(self, scanname, scantarget, modulelist, typelist, usecase):
    # Validation and setup...

    targetType = SpiderFootHelpers.targetTypeFromString(scantarget)
    cfg = deepcopy(self.config)               # isolate scan configuration

    sf = SpiderFoot(cfg)                       # create engine instance

    modlist = modulelist.split(',')
    scanId = SpiderFootHelpers.genScanInstanceId()

    # Spawn the scanner in its own process

    p = mp.Process(
        target=startSpiderFootScanner,
        args=(self.loggingQueue, scanname, scanId,
              scantarget, targetType, modlist, cfg)
    )
    p.daemon = True
    p.start()
    raise cherrypy.HTTPRedirect(f"{self.docroot}/scaninfo?id={scanId}")

Critical details in this implementation:

  • deepcopy(self.config) creates an isolated configuration snapshot
  • mp.Process launches startSpiderFootScanner with the multiprocessing module
  • self.loggingQueue is passed to enable log streaming back to the UI
  • The method redirects to /scaninfo rather than returning JSON directly

Scan Engine: The Multiprocessing Worker

Once spawned, the worker process runs SpiderFootScanner from sfscan.py. This class manages a single scan lifecycle independently of the web server process.

Scanner Initialization and Database Setup

The SpiderFootScanner.__init__ method at lines 30-53 establishes database connectivity:


# sfscan.py

class SpiderFootScanner:
    def __init__(self, scanName, scanId, targetValue,
                 targetType, moduleList, globalOpts, start=True):
        self.__config = deepcopy(globalOpts)
        self.__dbh = SpiderFootDb(self.__config)   # database handle

        self.__sf = SpiderFoot(self.__config)      # core engine

        self.__sf.dbh = self.__dbh                  # share DB connection

        self.__scanId = scanId
        self.__dbh.scanInstanceCreate(
            self.__scanId,
            self.__scanName,
            self.__targetValue
        )
        # Module loading, DNS setup, proxy configuration...

The scanner registers itself via scanInstanceCreate() before beginning data collection. All discovered information flows through SpiderFootDb methods like:

  • scanResultEvent() — stores harvested data points
  • scanLog() — records operational events
  • scanConfigSet() — persists module configuration

The SQLite Database as Communication Hub

spiderfoot/db.py implements SpiderFootDb, the abstraction layer used by both UI handlers and scanner processes. This design pattern—database-as-message-bus—eliminates direct socket communication between processes.

Key methods in the database wrapper:

Method Used By Purpose
scanInstanceCreate() Scanner Register new scan with metadata
scanResultEvent() Scanner Store discovered data events
scanInstanceSet() Both Update scan status (RUNNING, FINISHED, etc.)
scanLog() Scanner Append log entries
scanConfigGet/Set() Both Read/write module configuration

The UI polls these tables via endpoints like /scaninfo, /search, and /eventtypes to render current scan state without blocking on active workers.

Real-Time Log Streaming via Multiprocessing Queue

Beyond database polling, SpiderFoot implements immediate log delivery through a multiprocessing.Queue. This enables the "Live Log" view in the web interface.

Queue setup occurs in SpiderFootWebUi.__init__ at lines 84-92:


# sfwebui.py

class SpiderFootWebUi:
    def __init__(self, config):
        # ...

        self.loggingQueue = mp.Queue()
        self.scanManager = None

The queue is passed to every spawned scanner. When the scan engine generates log events, they enter this queue. The CherryPy process consumes from the same queue to stream updates to connected browsers—typically via WebSocket-style long polling or periodic AJAX fetches.

API Endpoints for Scan Control

The SpiderFootWebUi class exposes additional endpoints for lifecycle management:

  • GET /scaninfo?id={scanId} — retrieves scan metadata and current status
  • POST /search — queries accumulated results with filtering
  • POST /scandelete?id={scanId} — terminates and removes a scan
  • POST /stopscan?id={scanId} — signals graceful shutdown

These endpoints operate on database state rather than direct process manipulation. To stop a scan, the UI updates a status flag that the worker process checks periodically.

Why This Architecture Matters

The separation of concerns in SpiderFoot's communication design provides several operational benefits:

  • Resilience: Web server restarts don't kill running scans
  • Scalability: Multiple scans execute in parallel without GIL contention
  • Isolation: Module crashes in one scan don't affect others or the UI
  • Auditability: The database provides complete historical record of all activity

For developers extending SpiderFoot, this means that:

  1. New endpoints are added to SpiderFootWebUi in sfwebui.py
  2. New data types are stored via SpiderFootDb methods in spiderfoot/db.py
  3. New scan logic runs within SpiderFootScanner in sfscan.py

Summary

  • The SpiderFoot web UI communicates with the backend through CherryPy HTTP endpoints, not direct function calls
  • AJAX requests from spiderfoot.js trigger process spawning via multiprocessing.Process in sfwebui.py
  • Each scan runs as an isolated worker process executing SpiderFootScanner from sfscan.py
  • SQLite serves as the shared state store, with SpiderFootDb providing the abstraction layer
  • Log streaming uses a multiprocessing.Queue for real-time updates without database polling
  • The database-as-message-bus pattern decouples the UI from long-running scan operations

Frequently Asked Questions

How does SpiderFoot handle multiple concurrent scans?

Each scan spawns a separate multiprocessing.Process with its own SpiderFootScanner instance. The processes share no memory except through the SQLite database and optional logging queue. The SpiderFootDb class manages connection pooling and transaction isolation.

Can the web UI restart without losing active scans?

Yes. Because scan state persists in SQLite and workers run as independent processes, restarting the CherryPy server does not terminate running scans. The UI simply reconnects to existing scan records on restart.

What protocol does SpiderFoot use for frontend-to-backend communication?

Standard HTTP/HTTPS with AJAX. The frontend uses vanilla JavaScript XMLHttpRequest wrappers in sf.fetchData(), not WebSockets. Real-time updates rely on polling endpoints or the multiprocessing logging queue feeding into the HTTP response stream.

Where is scan configuration stored during a scan?

Configuration is snapshotted via deepcopy(self.config) in SpiderFootWebUi.startscan() and passed to the worker process. The worker may additionally persist settings via SpiderFootDb.scanConfigSet(). This ensures scans use consistent settings even if global configuration changes mid-scan.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →