# How the SpiderFoot Web UI Communicates with the Backend Scan Engine: Architecture Deep Dive

> Discover how SpiderFoot's web UI interacts with its backend scan engine. Explore the architecture using CherryPy HTTP endpoints, multiprocessing, and SQLite for seamless coordination.

- Repository: [Steve Micallef/spiderfoot](https://github.com/smicallef/spiderfoot)
- Tags: architecture
- Published: 2026-08-15

---

**SpiderFoot's web interface uses CherryPy HTTP endpoints, multiprocessing workers, and a shared SQLite database to coordinate between the frontend and the scan engine.**

The open-source OSINT platform SpiderFoot separates its web interface from the heavy lifting of data collection. Understanding how these components communicate helps developers extend the tool or troubleshoot scan issues. This article breaks down the exact mechanisms used in the smicallef/spiderfoot repository, with references to specific files and methods.

## Core Communication Architecture

SpiderFoot employs a **three-tier architecture** for UI-to-engine communication:

- **HTTP layer**: CherryPy serves AJAX endpoints that the JavaScript frontend consumes
- **Process layer**: Each scan spawns a separate `multiprocessing.Process` worker
- **Data layer**: SQLite acts as the single source of truth for scan state, results, and logs

The web UI never talks directly to a running scan. Instead, it creates records in the database and spawns processes that independently read from and write to that same database.

## Frontend: JavaScript AJAX Calls to CherryPy Endpoints

The frontend code in [`spiderfoot/static/js/spiderfoot.js`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/static/js/spiderfoot.js) provides wrapper functions that execute HTTP requests. The `sf.fetchData()` utility handles GET and POST calls, while higher-level functions like `sf.startscan()` encapsulate specific API operations.

Key frontend functions include:

- `sf.fetchData()` — generic AJAX wrapper at lines 24-33
- `sf.startscan()` — initiates a new scan via POST to `/startscan`
- `sf.scanDelete()`, `sf.rerunscan()` — control existing scans

When a user clicks **"Start Scan"**, the following JavaScript executes:

```javascript
// spiderfoot/static/js/spiderfoot.js
sf.startscan = function(name, target, modules, types, usecase, callback) {
    sf.fetchData(
        docroot + "/startscan",
        {
            scanname: name,
            scantarget: target,
            modulelist: modules,
            typelist: types,
            usecase: usecase
        },
        callback
    );
};

```

The `docroot` variable ensures the correct base URL, and responses are handled via callbacks that update the UI state.

## Backend: CherryPy Handlers and Process Spawning

The `SpiderFootWebUi` class in [`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py) exposes HTTP endpoints using the `@cherrypy.expose` decorator. Each endpoint maps to a class method that validates input, interacts with the database, and returns JSON responses.

### The `/startscan` Endpoint Implementation

The `startscan()` method at lines 374-418 of [`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py) demonstrates the full spawning workflow:

```python

# sfwebui.py

@cherrypy.expose
def startscan(self, scanname, scantarget, modulelist, typelist, usecase):
    # Validation and setup...

    targetType = SpiderFootHelpers.targetTypeFromString(scantarget)
    cfg = deepcopy(self.config)               # isolate scan configuration

    sf = SpiderFoot(cfg)                       # create engine instance

    modlist = modulelist.split(',')
    scanId = SpiderFootHelpers.genScanInstanceId()

    # Spawn the scanner in its own process

    p = mp.Process(
        target=startSpiderFootScanner,
        args=(self.loggingQueue, scanname, scanId,
              scantarget, targetType, modlist, cfg)
    )
    p.daemon = True
    p.start()
    raise cherrypy.HTTPRedirect(f"{self.docroot}/scaninfo?id={scanId}")

```

Critical details in this implementation:

- `deepcopy(self.config)` creates an isolated configuration snapshot
- `mp.Process` launches `startSpiderFootScanner` with the `multiprocessing` module
- `self.loggingQueue` is passed to enable log streaming back to the UI
- The method redirects to `/scaninfo` rather than returning JSON directly

## Scan Engine: The Multiprocessing Worker

Once spawned, the worker process runs `SpiderFootScanner` from [`sfscan.py`](https://github.com/smicallef/spiderfoot/blob/main/sfscan.py). This class manages a single scan lifecycle independently of the web server process.

### Scanner Initialization and Database Setup

The `SpiderFootScanner.__init__` method at lines 30-53 establishes database connectivity:

```python

# sfscan.py

class SpiderFootScanner:
    def __init__(self, scanName, scanId, targetValue,
                 targetType, moduleList, globalOpts, start=True):
        self.__config = deepcopy(globalOpts)
        self.__dbh = SpiderFootDb(self.__config)   # database handle

        self.__sf = SpiderFoot(self.__config)      # core engine

        self.__sf.dbh = self.__dbh                  # share DB connection

        self.__scanId = scanId
        self.__dbh.scanInstanceCreate(
            self.__scanId,
            self.__scanName,
            self.__targetValue
        )
        # Module loading, DNS setup, proxy configuration...

```

The scanner registers itself via `scanInstanceCreate()` before beginning data collection. All discovered information flows through `SpiderFootDb` methods like:

- `scanResultEvent()` — stores harvested data points
- `scanLog()` — records operational events
- `scanConfigSet()` — persists module configuration

## The SQLite Database as Communication Hub

[`spiderfoot/db.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/db.py) implements `SpiderFootDb`, the abstraction layer used by both UI handlers and scanner processes. This design pattern—**database-as-message-bus**—eliminates direct socket communication between processes.

Key methods in the database wrapper:

| Method | Used By | Purpose |
|--------|---------|---------|
| `scanInstanceCreate()` | Scanner | Register new scan with metadata |
| `scanResultEvent()` | Scanner | Store discovered data events |
| `scanInstanceSet()` | Both | Update scan status (RUNNING, FINISHED, etc.) |
| `scanLog()` | Scanner | Append log entries |
| `scanConfigGet/Set()` | Both | Read/write module configuration |

The UI polls these tables via endpoints like `/scaninfo`, `/search`, and `/eventtypes` to render current scan state without blocking on active workers.

## Real-Time Log Streaming via Multiprocessing Queue

Beyond database polling, SpiderFoot implements **immediate log delivery** through a `multiprocessing.Queue`. This enables the "Live Log" view in the web interface.

Queue setup occurs in `SpiderFootWebUi.__init__` at lines 84-92:

```python

# sfwebui.py

class SpiderFootWebUi:
    def __init__(self, config):
        # ...

        self.loggingQueue = mp.Queue()
        self.scanManager = None

```

The queue is passed to every spawned scanner. When the scan engine generates log events, they enter this queue. The CherryPy process consumes from the same queue to stream updates to connected browsers—typically via WebSocket-style long polling or periodic AJAX fetches.

## API Endpoints for Scan Control

The `SpiderFootWebUi` class exposes additional endpoints for lifecycle management:

- **GET `/scaninfo?id={scanId}`** — retrieves scan metadata and current status
- **POST `/search`** — queries accumulated results with filtering
- **POST `/scandelete?id={scanId}`** — terminates and removes a scan
- **POST `/stopscan?id={scanId}`** — signals graceful shutdown

These endpoints operate on database state rather than direct process manipulation. To stop a scan, the UI updates a status flag that the worker process checks periodically.

## Why This Architecture Matters

The separation of concerns in SpiderFoot's communication design provides several operational benefits:

- **Resilience**: Web server restarts don't kill running scans
- **Scalability**: Multiple scans execute in parallel without GIL contention
- **Isolation**: Module crashes in one scan don't affect others or the UI
- **Auditability**: The database provides complete historical record of all activity

For developers extending SpiderFoot, this means that:

1. New endpoints are added to `SpiderFootWebUi` in [`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py)
2. New data types are stored via `SpiderFootDb` methods in [`spiderfoot/db.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/db.py)
3. New scan logic runs within `SpiderFootScanner` in [`sfscan.py`](https://github.com/smicallef/spiderfoot/blob/main/sfscan.py)

## Summary

- The **SpiderFoot web UI** communicates with the backend through **CherryPy HTTP endpoints**, not direct function calls
- **AJAX requests** from [`spiderfoot.js`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot.js) trigger **process spawning** via `multiprocessing.Process` in [`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py)
- Each scan runs as an **isolated worker process** executing `SpiderFootScanner` from [`sfscan.py`](https://github.com/smicallef/spiderfoot/blob/main/sfscan.py)
- **SQLite** serves as the shared state store, with `SpiderFootDb` providing the abstraction layer
- **Log streaming** uses a `multiprocessing.Queue` for real-time updates without database polling
- The **database-as-message-bus** pattern decouples the UI from long-running scan operations

## Frequently Asked Questions

### How does SpiderFoot handle multiple concurrent scans?

Each scan spawns a separate `multiprocessing.Process` with its own `SpiderFootScanner` instance. The processes share no memory except through the SQLite database and optional logging queue. The `SpiderFootDb` class manages connection pooling and transaction isolation.

### Can the web UI restart without losing active scans?

Yes. Because scan state persists in SQLite and workers run as independent processes, restarting the CherryPy server does not terminate running scans. The UI simply reconnects to existing scan records on restart.

### What protocol does SpiderFoot use for frontend-to-backend communication?

Standard HTTP/HTTPS with AJAX. The frontend uses vanilla JavaScript `XMLHttpRequest` wrappers in `sf.fetchData()`, not WebSockets. Real-time updates rely on polling endpoints or the multiprocessing logging queue feeding into the HTTP response stream.

### Where is scan configuration stored during a scan?

Configuration is snapshotted via `deepcopy(self.config)` in `SpiderFootWebUi.startscan()` and passed to the worker process. The worker may additionally persist settings via `SpiderFootDb.scanConfigSet()`. This ensures scans use consistent settings even if global configuration changes mid-scan.