How to Export SpiderFoot Scan Results to CSV, JSON, or GEXF Formats
SpiderFoot exports scan results through three interfaces: web UI endpoints, CLI commands, and direct HTTP requests—all powered by a shared SQLite-backed pipeline.
You can export SpiderFoot reconnaissance data in multiple structured formats using either the built-in command-line tool, the web interface, or programmatic HTTP calls. This guide breaks down the exact methods, source code locations, and practical commands needed to extract your scan results.
Web UI Export Endpoints
The SpiderFootWebUi class in sfwebui.py implements three HTTP endpoints that handle all export functionality. These methods query the SQLite database via SpiderFootDb, transform the data, and stream formatted responses to clients.
JSON Export
The scanexportjsonmulti endpoint (lines 611–662) retrieves event rows for one or more scans, builds a list of dictionaries, and serializes with json.dumps. It forces a download using the Content-Disposition: attachment header.
curl -X POST -d "ids=42,43" http://localhost:5001/scanexportjsonmulti -o scans.json
Response structure: array of event dictionaries containing fields like generated, event_type, data, and module.
CSV and Excel Export
The scaneventresultexportmulti endpoint (lines 441–486) iterates over event rows and outputs either:
- CSV format: Uses Python's
csv.writerwithtext/csvcontent type - Excel format: Builds an in-memory workbook via
buildExcel()withapplication/vnd.openxmlformats-officedocument.spreadsheetml.sheet
# CSV output
curl -X POST -d "ids=42" http://localhost:5001/scaneventresultexportmulti -o scan42.csv
# Excel output (when requested via UI)
GEXF Graph Export
The scanviz endpoint (lines 665–700) generates graph visualizations. When the gexf query parameter is non-zero, it invokes SpiderFootHelpers.buildGraphGexf and returns XML with Content-Type: application/gexf.
curl -X POST -d "ids=42" "http://localhost:5001/scanviz?gexf=1" -o scan42.gexf
Command-Line Interface (sfcli)
The CLI wrapper in sfcli.py (do_export, lines 791–820) translates user commands into POST requests against the same web endpoints. This provides a scriptable interface without manual HTTP construction.
# Export scan 42 as JSON (default) and write to file
sfcli export 42 -f scan42.json
# Export as CSV
sfcli export 42 -t csv -f scan42.csv
# Export as GEXF for Gephi or other graph tools
sfcli export 42 -t gexf -f scan42.gexf
The CLI validates format options (json, csv, gexf), posts scan ID(s), and either prints raw output to STDOUT or writes to a file when -f is specified.
Programmatic Python Export
For automation pipelines, use requests to call the export endpoints directly:
import requests
import json
BASE_URL = "http://localhost:5001"
# JSON export: parseable Python structures
resp = requests.post(
f"{BASE_URL}/scanexportjsonmulti",
data={"ids": "42"}
)
events = resp.json() # List of event dictionaries
print(json.dumps(events, indent=2))
# CSV export: binary content for file writing
resp = requests.post(
f"{BASE_URL}/scaneventresultexportmulti",
data={"ids": "42"}
)
with open("scan42.csv", "wb") as f:
f.write(resp.content)
# GEXF export: XML for graph visualization tools
resp = requests.post(
f"{BASE_URL}/scanviz?gexf=1",
data={"ids": "42"}
)
with open("scan42.gexf", "wb") as f:
f.write(resp.content)
Helper Utilities and Data Flow
Two helper functions in spiderfoot/helpers.py power graph-based exports:
buildGraphJson: Creates JSON-compatible node/edge structures for custom visualizationbuildGraphGexf: Generates GEXF XML documents compatible with Gephi, Cytoscape, and similar tools
Both are invoked by scanviz when building graph representations.
The complete export pipeline follows this sequence:
- Web request hits endpoint (
scanexportjsonmulti,scaneventresultexportmulti,scanviz) SpiderFootDbquery retrieves matching event rows from SQLite- Data transformation converts to JSON, CSV/Excel, or GEXF format
- HTTP headers set appropriate content types and disposition
- Client download streams the formatted result
Key Source Files
| File | Purpose |
|---|---|
sfwebui.py |
HTTP export endpoints (scanexportjsonmulti, scaneventresultexportmulti, scanviz) |
sfcli.py |
CLI command do_export that wraps web endpoints |
spiderfoot/helpers.py |
Graph builders buildGraphJson and buildGraphGexf |
sfdb.py (via SpiderFootDb import) |
Low-level SQLite access for event retrieval |
Export Format Selection Guide
| Format | Use Case | Endpoint | CLI Type |
|---|---|---|---|
| JSON | API integration, scripting, data processing | scanexportjsonmulti |
json (default) |
| CSV | Spreadsheet analysis, quick review | scaneventresultexportmulti |
csv |
| Excel | Formal reporting, styled tables | scaneventresultexportmulti |
N/A (UI only) |
| GEXF | Graph visualization (Gephi, etc.) | scanviz?gexf=1 |
gexf |
Summary
- SpiderFoot scan results export uses three HTTP endpoints in
sfwebui.pythat all query the same SQLite database. - JSON export (
scanexportjsonmulti) returns parseable event arrays ideal for automation. - CSV/Excel export (
scaneventresultexportmulti) produces spreadsheet-compatible outputs using standard Python csv or openpyxl. - GEXF export (
scanviz) generates graph XML viaSpiderFootHelpers.buildGraphGexffor network visualization. - The CLI (
sfcli export) provides the simplest interface, forwarding to these endpoints with automatic file handling. - Direct HTTP access enables custom integrations when the full SpiderFoot stack isn't required client-side.
Frequently Asked Questions
What formats does SpiderFoot support for scan export?
SpiderFoot supports JSON, CSV, Excel (XLSX), and GEXF formats. JSON and CSV/Excel handle raw event data. GEXF produces graph structures for visualization tools like Gephi. Excel output is UI-only; JSON, CSV, and GEXF work via both web and CLI interfaces.
Why does the CLI use HTTP endpoints instead of direct database access?
The sfcli.py design wraps web endpoints to ensure consistent security boundaries and single source of truth for export logic. This avoids duplicating database query and formatting code, and lets the CLI work remotely against any SpiderFoot instance without filesystem access.
How do I export multiple scans at once?
Pass comma-separated IDs to the ids parameter. Both scanexportjsonmulti and CLI commands accept multiple scans: sfcli export 42,43,44 -f combined.json or curl -d "ids=42,43" .../scanexportjsonmulti.
Can I automate SpiderFoot exports in a CI/CD pipeline?
Yes. Use the CLI with file output (sfcli export ID -t json -f output.json) for shell-based automation, or the Python requests example above for native integration. Both methods require only network access to the SpiderFoot web port (default 5001).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →