How the Cross-Scan Engine Works in user-scanner: A Deep Dive into Multi-Pass OSINT Discovery

The cross-scan engine in user-scanner is a second-pass mechanism that expands initial email or username scans by extracting new identifiers (usernames and emails) from profile metadata and recursively feeding them back into scan modules.

The cross-scan engine enables recursive OSINT discovery by treating every found profile as a potential source of additional targets. This article explains exactly how this multi-pass architecture works in the kaifcodec/user-scanner repository, based on direct analysis of the source code.

Cross-Scan Engine Architecture Overview

The cross-scan engine operates as a pipeline processor that transforms first-pass results into new scan targets. Unlike a simple single-query tool, it treats metadata as a generative resource—every username link, verified handle, and discovered email becomes fuel for deeper investigation.

The entry point is run_cross_scan() in [user_scanner/core/cross_scan.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/cross_scan.py#L82), which coordinates scope filtering, pivot extraction, budget allocation, and recursive execution.

How the Cross-Scan Engine Processes Results

Entry Point and Configuration

The run_cross_scan() function accepts three critical inputs:

  1. First-pass results — the Results object containing accounts, emails, and metadata
  2. ScanConfig — generic CLI configuration (modules, timeouts, output format)
  3. CrossScanConfig — cross-scan specific controls: links class, emails class, sweep budget, and depth
from user_scanner.core.cross_scan import run_cross_scan, CrossScanConfig
from user_scanner.core.helpers import ScanConfig

cross_cfg = CrossScanConfig(
    links="verified",    # "all", "verified", or "none"

    emails="verified",   # filter extracted emails

    sweep=3,             # maximum new targets per round

    depth=2              # recursive rounds

)
second_results = run_cross_scan(first_results, ScanConfig(), cross_cfg)

Scope Determination: Which Modules Run

Before extracting targets, the engine determines which modules are available for the second pass.

Username scope — _scope() (L87-L107) filters modules based on -m / -c CLI restrictions.

Email scope — _email_scope() (L110-L150) additionally excludes "loud" modules (those with high API cost or rate limits) unless --allow-loud is explicitly set.

Pivot Extraction: Mining Usernames from Metadata

The Pivot Pipeline

Pivots are usernames discovered in profile metadata. The cross-scan engine classifies and ranks them through a three-stage pipeline in user_scanner/core/pivots.py:

Stage Function Purpose
Extraction extract_pivots() Pulls handles from links, verified, and text fields
Selection select_pivots() Applies the links filter (all/verified/none)
Ranking rank_usernames() Scores by confidence class: handle > verified > link

The _fresh_pivots() function (L65-L80) implements deduplication—discarding usernames already swept or checked—to prevent redundant work.

Pivot Confidence Classes

The engine assigns confidence tiers based on source reliability:

  • handle — explicit username field from a confirmed account (highest confidence)
  • verified — marked as verified by the platform
  • link — found in URL or bio text (lowest confidence, may be coincidental)

Email Extraction and Anchor Building

Email Target Selection

The _fresh_emails() function (L92-L103) extracts email addresses from first-pass results, applying the emails filter.

Email confidence ranking uses rank_emails() from core/confidence.py, which distinguishes:

  • field — email from dedicated email field (higher confidence)
  • text — email scraped from body text (lower confidence)

Anchor-Based Scoring

The build_anchors() function creates a ground-truth set from:

  • Confirmed accounts from initial scan
  • Email-derived verification points
  • Pivot source URLs

These anchors feed _apply_confidence() (L84-L108), which recomputes confidence ratings after each round.

Budget Allocation and Target Prioritization

Splitting the Sweep Budget

The _split_budget() function (L18-L30) implements fair-share allocation:


# Pseudocode from source analysis

username_slots = max(1, budget // 2) if usernames_present else 0
email_slots = max(1, budget - username_slots) if emails_present else 0

This guarantees both identifier types receive at least one slot when the budget permits.

Target Selection Functions

  • _sweep_targets() (L33-L60) — selects top-ranked usernames by pivot confidence
  • _email_targets() (L62-L89) — truncates ranked email list to budget
  • _named_targets() (L13-L52) — handles pivots with pre-attached sites (direct module calls without sweeping)

Cross-Scan Engine Execution Flow

Username Module Invocation

For each swept username, the engine routes to core/orchestrator.py:

  • run_user_full() — executes all modules in scope
  • run_user_module() — executes restricted module set (when -m specified)

Email Module Invocation

Email targets route through core/email_orchestrator.py:

  • run_email_full_batch() — full module suite
  • run_email_module_batch() — restricted execution

Result Tagging and Confidence Recalculation

The _tag() and _tag_emails() functions attach pivot_source metadata to each hit, enabling audit trails for how targets were discovered.

After execution, _apply_confidence() updates ratings using the expanded anchor set, producing final classifications:

  • confirmed — multiple independent verifications
  • likely — strong anchor support
  • candidate — single source, unverified
  • conflicting — contradictory evidence

Recursive Depth and Chain Discovery

When --cross-depth N > 1, the engine loops:

  1. New results from round n become source for round n+1
  2. _followable() filters out conflicting or failed results
  3. Process repeats until depth exhausted or no new pivots found

This enables multi-hop discovery—finding a GitHub profile from an email, then LinkedIn from GitHub links, then Twitter from LinkedIn, etc.

Command-Line Examples for Cross-Scan Engine Use


# Basic cross-scan from email with defaults (depth 1, sweep 3)

user-scanner -e alice@example.com --cross-scan

# Verified-only pivots to reduce noise

user-scanner -u alice --cross-scan --cross-links verified

# Two-hop deep investigation

user-scanner -e alice@example.com --cross-scan --cross-depth 2

# Disable sweeping—named checks only (faster, narrower)

user-scanner -u alice --cross-scan --cross-sweep 0

# Full aggressive scan: loud modules, all links, deep recursion

user-scanner -e alice@example.com --cross-scan \
    --allow-loud --cross-links all --cross-emails all \
    --cross-sweep 10 --cross-depth 3

Key Source Files for Cross-Scan Engine Implementation

File Purpose
[user_scanner/core/cross_scan.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/cross_scan.py) Main orchestration: run_cross_scan(), budgeting, depth looping
[user_scanner/core/pivots.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/pivots.py) Pivot dataclass, extraction, classification, ranking
[user_scanner/core/confidence.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/confidence.py) Anchor building, email ranking, confidence scoring
[user_scanner/core/orchestrator.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/orchestrator.py) Username module execution
[user_scanner/core/email_orchestrator.py](https://github.com/kaifcodec/user-scanner/blob/main/user_scanner/core/email_orchestrator.py) Email module execution
[docs/CROSS_SCAN.md](https://github.com/kaifcodec/user-scanner/blob/main/docs/CROSS_SCAN.md) User documentation

Summary

  • The cross-scan engine transforms single-query OSINT into recursive discovery by mining metadata from found profiles
  • Pivots (extracted usernames) and emails are ranked by confidence, then fed back into appropriate modules
  • _split_budget() ensures balanced resource allocation between username and email targets
  • Depth looping enables multi-hop investigations when --cross-depth > 1
  • Anchor-based confidence scoring in core/confidence.py provides reliable result triage
  • All core logic resides in user_scanner/core/cross_scan.py with supporting modules for pivots, confidence, and orchestration

Frequently Asked Questions

What is the difference between a pivot and a regular search target?

A pivot is a username discovered inside profile metadata during a scan, while a regular target is supplied by the user. Pivots carry provenance metadata (pivot_source) and confidence classifications based on where they were found—handle fields, verified badges, or link text—enabling the cross-scan engine to prioritize higher-quality leads.

How does the sweep budget prevent runaway scans?

The --cross-sweep parameter caps total new targets per round, and _split_budget() enforces minimum allocation slots. When combined with deduplication in _fresh_pivots() and depth limits via --cross-depth, these controls bound execution time and API usage regardless of how many potential pivots exist in metadata.

Can I use cross-scan mode with module restrictions?

Yes. The _scope() and _email_scope() functions honor -m (module selection) and -c (category selection) flags from the base ScanConfig. Additionally, _email_scope() filters "loud" modules unless --allow-loud is set, protecting rate-limited APIs during automated recursive scans.

Why do results get re-tagged with pivot_source?

The _tag() and _tag_emails() functions attach provenance labels to every cross-scan hit, creating an audit trail that shows which original result yielded each discovered account. This enables investigators to trace chains of discovery and assess the reliability of multi-hop findings based on their source path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →