# How to Monitor SpiderFoot Scan Progress Programmatically via Its REST API

> Monitor SpiderFoot scan progress programmatically using the REST API. Poll the scanopts endpoint for real-time status and metadata without database access.

- Repository: [Steve Micallef/spiderfoot](https://github.com/smicallef/spiderfoot)
- Tags: how-to-guide
- Published: 2026-08-15

---

**Poll the `/scanopts` endpoint with a scan ID to retrieve real-time status, timestamps, and metadata as JSON without direct database access.**

SpiderFoot's built-in web interface doubles as a **REST API server** that exposes scan state through lightweight HTTP endpoints. This article explains how to monitor SpiderFoot scan progress programmatically using the `/scanopts` endpoint, which queries the underlying SQLite database and returns structured JSON suitable for automation scripts and external monitoring systems.

## The Core Endpoint: `/scanopts`

The primary interface for scan monitoring is implemented in **[`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py)** as the `scanopts` method. This endpoint accepts a single query parameter `id` containing the scan identifier and returns a JSON object with three top-level keys:

- **`meta`** – array containing scan name, target, timestamps, and current status
- **`config`** – the scan's module configuration
- **`configdesc`** – human-readable configuration descriptions

The method is decorated with `@cherrypy.tools.json_out()`, which ensures responses are automatically serialized to JSON with `Content-Type: application/json` headers. Internally, the endpoint delegates to `SpiderFootDb.scanInstanceGet()` in **[`spiderfoot/db.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/db.py)**, which executes:

```python
def scanInstanceGet(self, instanceId):
    # Returns: (name, target, created, started, ended, status)

    qry = "SELECT name, seed_target, created, started, ended, status \
           FROM tbl_scan_instance WHERE guid = ?"
    ...

```

The `status` field in position 5 of the returned tuple contains one of: `STARTING`, `RUNNING`, `FINISHED`, `ABORTED`, or `ERROR-FAILED`.

## Step-by-Step Monitoring Workflow

### 1. Obtain a Scan ID

Start a scan via `/startscan` (POST with `scanname`, `scantarget`, and optional `modulelist`) or launch through the CLI. The response contains the generated scan UUID. Example with curl:

```bash
curl -X POST "http://localhost:5001/startscan" \
     -d "scanname=ExampleScan" \
     -d "scantarget=example.com" \
     -d "modulelist=sfp_dnsresolve,sfp_portscan_tcp"

```

Parse the JSON response to extract `id`.

### 2. Poll for Status Changes

Repeatedly GET `/scanopts?id=<scan_id>` until the scan reaches a terminal state. The `meta` array uses fixed indices:

| Index | Field | Type |
|-------|-------|------|
| 0 | Scan name | string |
| 1 | Target | string |
| 2 | Created timestamp | integer (epoch) |
| 3 | Started datetime | string |
| 4 | Ended datetime | string or "Not yet" |
| 5 | **Status** | string |

### 3. Parse Timestamps for Progress Metrics

The `started` and `ended` fields allow elapsed time calculation. When `ended` equals "Not yet", the scan is still active.

## Complete Python Implementation

This production-ready script demonstrates polling with exponential backoff and proper error handling:

```python
import time
import sys
import requests
from datetime import datetime

BASE_URL = "http://localhost:5001"  # Adjust to your SpiderFoot server

POLL_INTERVAL = 5                   # Seconds between checks

MAX_RETRIES = 3

class ScanMonitor:
    def __init__(self, base_url: str):
        self.base_url = base_url.rstrip("/")
        self.session = requests.Session()

    def get_status(self, scan_id: str) -> dict:
        """Fetch scan metadata from /scanopts endpoint."""
        url = f"{self.base_url}/scanopts"
        for attempt in range(MAX_RETRIES):
            try:
                resp = self.session.get(url, params={"id": scan_id}, timeout=30)
                resp.raise_for_status()
                return resp.json()
            except requests.RequestException as e:
                if attempt == MAX_RETRIES - 1:
                    raise RuntimeError(f"Failed to query scan status: {e}")
                time.sleep(2 ** attempt)  # Exponential backoff

        return {}  # Unreachable

    def parse_meta(self, data: dict) -> dict:
        """Extract structured fields from meta array."""
        meta = data.get("meta", [])
        if len(meta) < 6:
            raise ValueError(f"Unexpected meta format: {meta}")
        return {
            "name": meta[0],
            "target": meta[1],
            "created": meta[2],
            "started": meta[3],
            "ended": meta[4],
            "status": meta[5],
        }

    def monitor(self, scan_id: str) -> dict:
        """Block until scan completes, yielding progress updates."""
        print(f"Monitoring scan {scan_id}...")
        while True:
            data = self.get_status(scan_id)
            parsed = self.parse_meta(data)

            status = parsed["status"]
            started = parsed["started"]
            target = parsed["target"]

            # Calculate runtime

            if started != "Not started":
                start_dt = datetime.strptime(started, "%Y-%m-%d %H:%M:%S")
                elapsed = datetime.now() - start_dt
                elapsed_str = str(elapsed).split(".")[0]  # Trim microseconds

            else:
                elapsed_str = "N/A"

            print(f"[{elapsed_str}] {target}: {status}")

            if status in ("FINISHED", "ABORTED", "ERROR-FAILED"):
                return parsed

            time.sleep(POLL_INTERVAL)


if __name__ == "__main__":
    if len(sys.argv) != 2:
        print(f"Usage: {sys.argv[0]} <scan_id>")
        sys.exit(1)

    monitor = ScanMonitor(BASE_URL)
    final = monitor.monitor(sys.argv[1])
    print(f"\nFinal status: {final['status']}")
    sys.exit(0 if final["status"] == "FINISHED" else 1)

```

Run with: `python monitor.py a1b2c3d4-e5f6-7890-abcd-ef1234567890`

## Alternative: Query Endpoint for Raw Data

For advanced use cases, `/query` accepts raw SQL against the SQLite database (requires appropriate permissions). To fetch result counts during a scan:

```bash
curl "http://localhost:5001/query?query=\
SELECT type, COUNT(*) FROM tbl_scan_results \
WHERE scan_instance_id='YOUR_SCAN_ID' GROUP BY type"

```

**Security note:** The `/query` endpoint exposes the full database schema. Restrict access in production deployments.

## Key Implementation Files

| File | Purpose | Critical Function/Line |
|------|---------|------------------------|
| [`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py) | CherryPy web server and REST routing | `scanopts` method at line 747 [[source](https://github.com/smicallef/spiderfoot/blob/master/sfwebui.py#L747)] |
| [`spiderfoot/db.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/db.py) | SQLite abstraction and schema management | `scanInstanceGet()` at line 719 [[source](https://github.com/smicallef/spiderfoot/blob/master/spiderfoot/db.py#L719)] |
| [`sf.py`](https://github.com/smicallef/spiderfoot/blob/main/sf.py) | CLI entry point and server initialization | Instantiates `SpiderFootWebUi` class |
| [`sflib.py`](https://github.com/smicallef/spiderfoot/blob/main/sflib.py) | Core scanning engine | `startSpiderFootScanner` creates scan records |

## Status State Machine

Understanding state transitions helps design robust monitoring logic:

```

STARTING → RUNNING → FINISHED
    ↓         ↓
ABORTED   ERROR-FAILED

```

- **STARTING**: Scan record created, modules loading
- **RUNNING**: Active data collection in progress
- **FINISHED**: Normal completion, results available
- **ABORTED**: User-initiated cancellation
- **ERROR-FAILED**: Unrecoverable module or system error

## Summary

- **Use `/scanopts?id=<scan_id>`** as the canonical endpoint for SpiderFoot scan progress monitoring via REST API
- **Parse the `meta` array** at fixed indices—status is always position 5
- **Poll every 5-10 seconds** with exponential backoff for network resilience
- **Terminate polling** when status reaches `FINISHED`, `ABORTED`, or `ERROR-FAILED`
- **Query [`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py) and [`spiderfoot/db.py`](https://github.com/smicallef/spiderfoot/blob/main/spiderfoot/db.py)** source files to understand data flow and extend functionality

## Frequently Asked Questions

### Does SpiderFoot support WebSocket push notifications for scan status?

No. According to the [`sfwebui.py`](https://github.com/smicallef/spiderfoot/blob/main/sfwebui.py) implementation, SpiderFoot uses a **polling-based REST architecture**. The CherryPy server does not implement WebSocket endpoints for real-time status streaming. For low-latency updates, reduce polling intervals to 1-2 seconds or implement a local event listener using SpiderFoot's Python API directly.

### Can I monitor multiple scans simultaneously?

Yes. The `/scanopts` endpoint is stateless—issue parallel requests for different scan IDs. Each call executes an independent SQLite query via `scanInstanceGet()`. No session binding exists between requests, though the underlying database connection pool may limit concurrent throughput.

### What authentication does the REST API require?

SpiderFoot's web interface uses **HTTP Basic Auth or session cookies** configured during startup. Pass credentials in your HTTP client:

```python
requests.get(url, auth=("username", "password"))

```

Without authentication, endpoints return `401 Unauthorized`. The `--no-auth` CLI flag disables this requirement (not recommended for production).

### How do I calculate scan completion percentage?

SpiderFoot does not expose percentage-complete metrics natively. Approximate progress by **comparing elapsed time against typical scan duration** for your target and module set, or count results returned via `/query` against expected thresholds. The `meta` array contains no total-work estimator.