HNSW Vector Search Architecture for Room Fingerprinting in RuView

RuView implements a Hierarchical Navigable Small World (HNSW) graph to convert raw Channel State Information (CSI) into 329-dimensional vectors, enabling sub-millisecond room identification through cosine-similarity search against a product-quantized index of up to one million fingerprints.

The RuView sensing platform uses Wi-Fi Channel State Information (CSI) to recognize physical environments without cameras. By implementing an HNSW vector search architecture for room fingerprinting, the system transforms raw radio signals into searchable embeddings, achieving millisecond-scale inference while maintaining a compact memory footprint through 8-bit product quantization.

Core Architecture Components

Feature Extraction and Vector Encoding

Raw CSI frames undergo feature extraction to produce CsiFeatures, which are encoded into a 329-dimensional vector suitable for HNSW indexing. This dimensionality corresponds to the processed subcarrier information from Wi-Fi sensing hardware. The encoding logic resides in the Rust-based sensing server, specifically within the vector pipeline that prepares data for the env_fingerprint index.

HNSW Index Configuration

The env_fingerprint index is tuned for 329-dimensional CSI embeddings with parameters that balance recall against memory consumption. According to docs/adr/ADR-004-hnsw-vector-search-fingerprinting.md, the configuration uses:

  • m: 16 (maximum connections per node)
  • ef_construction: 200 (ensures >99% recall during bulk building)
  • ef_search: 64 (provides >95% recall with <1ms latency at 100k vectors)
  • metric: Cosine (optimal for normalized CSI embeddings)
  • max_elements: 1,000,000 (upper bound for large deployments)
  • quantization: PQ8 (8-bit product quantization yielding 8× compression with <5% recall loss)

The configuration appears in the ADR as Rust-style pseudo-code and is used when constructing the env_fingerprint index.

/// HNSW configuration tuned for CSI vector characteristics
pub struct CsiHnswConfig {
    dim: usize,          // 329 for 64 subcarriers
    m: usize,            // 16
    ef_construction: usize, // 200
    ef_search: usize,    // 64
    metric: DistanceMetric, // Cosine
    max_elements: usize, // 1_000_000
    simd: bool,          // true
    quantization: Quantization, // PQ8
}

Core Data Structures

rvf_pipeline.rs defines lightweight, serializable structures that enable the HNSW index to be embedded within RVF containers for edge deployment.

/// A single node in an HNSW layer.
#[derive(Debug, Clone)]
pub struct HnswNode {
    pub id: usize,
    pub neighbors: Vec<usize>,
    pub vector: Vec<f32>,
}

/// One layer of the HNSW graph.
#[derive(Debug, Clone)]
pub struct HnswLayer {
    pub nodes: Vec<HnswNode>,
}

/// Serializable HNSW index used for sparse inference neuron routing.
#[derive(Debug, Clone)]
pub struct HnswIndex {
    pub layers: Vec<HnswLayer>,
    pub entry_point: usize,
    pub ef_construction: usize,
    pub m: usize,
}

Reference: rust-port/wifi-densepose-rs/crates/wifi-densepose-sensing-server/src/rvf_pipeline.rs, lines 31-48.

The to_bytes and from_bytes methods (lines 50-66) facilitate packing the index into RVF containers for transmission to ESP32 or WASM runtimes.

Similarity-Based Detection Pipeline

At runtime, the SimilarityDetector struct performs room identification by comparing incoming CSI vectors against both human-present and empty-room indices.

The detection process follows these steps:

  1. Encode the CSI window into a query vector using features.to_rvf_vector().
  2. Search both the human_patterns and empty_patterns indices with k=5 nearest neighbors.
  3. Calculate similarity confidence as the ratio of distances: avg_empty_dist / (avg_human_dist + avg_empty_dist).
  4. Fuse with legacy threshold detection using a weighted average (fusion_alpha = 0.7).
  5. Decide presence if the fused confidence exceeds 0.5.
pub struct SimilarityDetector {
    human_patterns: HnswIndex,
    empty_patterns: HnswIndex,
    fusion_alpha: f64,
}

impl SimilarityDetector {
    pub fn detect(&self, features: &CsiFeatures) -> DetectionResult {
        let query_vec = features.to_rvf_vector();

        let human_neighbors = self.human_patterns.search(&query_vec, k=5);
        let empty_neighbors = self.empty_patterns.search(&query_vec, k=5);

        let avg_human_dist = human_neighbors.mean_distance();
        let avg_empty_dist = empty_neighbors.mean_distance();

        let similarity_confidence = avg_empty_dist / (avg_human_dist + avg_empty_dist);
        let threshold_confidence = self.traditional_threshold_detect(features);

        let fused_confidence = self.fusion_alpha * similarity_confidence
                             + (1.0 - self.fusion_alpha) * threshold_confidence;

        DetectionResult {
            human_detected: fused_confidence > 0.5,
            confidence: fused_confidence,
            similarity_confidence,
            threshold_confidence,
            nearest_human_pattern: human_neighbors[0].metadata.clone(),
            nearest_empty_pattern: empty_neighbors[0].metadata.clone(),
        }
    }
}

Reference: docs/adr/ADR-004-hnsw-vector-search-fingerprinting.md, lines 33-67.

Incremental Learning and Index Maintenance

The architecture supports continuous improvement through online learning. When the system confirms a detection, it inserts the corresponding vector back into the appropriate index (human_patterns or empty_patterns).

The incremental learning loop operates as follows:

  1. Capture CSI and extract features.
  2. Search existing indices to establish baseline distances.
  3. Confirm the detection through similarity and threshold fusion.
  4. Insert the new vector into the index, updating neighbor links via the GNN layer (as described in ADR-006).
  5. Adapt the fusion weight dynamically (per ADR-005).
  6. Maintain the index through periodic pruning and layer re-balancing to prevent degradation.

This approach allows the system to adapt to environmental changes without requiring full retraining or index rebuilding.

Reference: docs/adr/ADR-004-hnsw-vector-search-fingerprinting.md, lines 79-92.

Performance Characteristics

The HNSW implementation delivers substantial speedups over brute-force search while maintaining manageable memory footprints through product quantization.

Storage requirements for the env_fingerprint index:

| # Vectors | Raw size | PQ8 compressed | HNSW overhead | Total (≈) |

|-----------|----------|----------------|---------------|-----------| | 10 k | 12.9 MB | 1.6 MB | 2.5 MB | 4.1 MB | | 100 k | 129 MB | 16 MB | 25 MB | 41 MB | | 1 M | 1.29 GB | 160 MB | 250 MB | 410 MB |

Search latency comparisons (329-dimensional vectors, ef_search = 64):

Vectors Brute-Force HNSW Speed-up
10 k 3.2 ms 0.08 ms 40×
100 k 32 ms 0.3 ms 107×
1 M 320 ms 0.9 ms 356×

These metrics demonstrate that the system can query indices containing up to one million room fingerprints in under one millisecond, making it suitable for real-time Wi-Fi sensing applications.

Reference: docs/adr/ADR-004-hnsw-vector-search-fingerprinting.md, lines 94-111.

Implementation Examples

Creating a Room-Fingerprint Index in Rust

The following example demonstrates how to initialize the env_fingerprint index using the configuration parameters defined in ADR-004:

use wifi_densepose_sensing_server::rvf_pipeline::{HnswIndex, HnswLayer, HnswNode};

/// HNSW configuration tuned for CSI vector characteristics
pub struct CsiHnswConfig {
    dim: usize,          // 329 for 64 subcarriers
    m: usize,            // 16
    ef_construction: usize, // 200
    ef_search: usize,    // 64
    metric: DistanceMetric, // Cosine
    max_elements: usize, // 1_000_000
    simd: bool,          // true
    quantization: Quantization, // PQ8
}

fn make_env_fingerprint_index(cfg: &CsiHnswConfig) -> HnswIndex {
    // Allocate empty layers (level 0 … log_M(n))
    let mut layers = Vec::new();
    for _ in 0..cfg.max_elements.ilog2() as usize + 1 {
        layers.push(HnswLayer { nodes: Vec::new() });
    }

    HnswIndex {
        layers,
        entry_point: 0,          // will be set after first insertion
        ef_construction: cfg.ef_construction,
        m: cfg.m,
    }
}

Adding Vectors to the Index

Vector insertion updates the graph structure and establishes neighbor connections:

fn add_vector(idx: &mut HnswIndex, id: usize, vec: Vec<f32>) {
    // Simple insertion – real implementation uses the HNSW algorithm
    let node = HnswNode { id, neighbors: Vec::new(), vector: vec };
    idx.layers[0].nodes.push(node);
    // Update entry point if this is the first node
    if idx.layers[0].nodes.len() == 1 {
        idx.entry_point = id;
    }
}

Querying from Python

The Rust implementation exposes bindings for the Python pipeline:

from ruvector import VectorIndex  # Python bindings to RuVector HNSW

# Load the index shipped inside an RVF container

index = VectorIndex.from_rvf_file("room_fingerprint.rvf")

# `query` is a 329‑dim numpy array (normalized CSI embedding)

neighbors = index.search(query, k=5)   # returns (ids, distances, metadata)

confidence = 1.0 - neighbors.distances.mean()   # higher = more similar

print("Room similarity confidence:", confidence)

Key Source Files

Component File Role
Architecture description [docs/adr/ADR-004-hnsw-vector-search-fingerprinting.md](https://github.com/ruvnet/RuView/blob/main/docs/adr/ADR-004-hnsw-vector-search-fingerprinting.md) Design rationale, index types, performance numbers
Room‑fingerprint index config (pseudo‑code in ADR‑004) Shows CsiHnswConfig values
HNSW data structures & (de)serialization [rust-port/wifi-densepose-rs/crates/wifi-densepose-sensing-server/src/rvf_pipeline.rs](https://github.com/ruvnet/RuView/blob/main/rust-port/wifi-densepose-rs/crates/wifi-densepose-sensing-server/src/rvf_pipeline.rs) HnswNode, HnswLayer, HnswIndex, to_bytes, from_bytes
Similarity‑based detector Same file (rvf_pipeline.rs, lines 33‑67) SimilarityDetector implementation
Embedding generation (AETHER) [docs/adr/ADR-024-contrastive-csi-embedding-model.md](https://github.com/ruvnet/RuView/blob/main/docs/adr/ADR-024-contrastive-csi-embedding-model.md) Produces the 128‑dim z_csi that is later reduced to 329‑dim room vectors
Python bindings wifi-densepose-nn crate (exposed as ruvector Python package) Allows Rust HNSW indices to be queried from the RuView Python stack

Summary

  • RuView converts Channel State Information (CSI) into 329-dimensional vectors to create searchable room fingerprints.
  • The HNSW index (env_fingerprint) uses cosine similarity with m=16, ef_construction=200, and PQ8 quantization to balance speed and memory.
  • Sub-millisecond query latency is achieved even with 1 million vectors, delivering 356× speedup over brute-force search.
  • The SimilarityDetector fuses HNSW distances with legacy threshold detection using a configurable alpha weight.
  • Incremental learning allows confirmed detections to be inserted back into the index without full rebuilds.

Frequently Asked Questions

What makes HNSW suitable for room fingerprinting in RuView?

HNSW graphs provide approximate nearest neighbor search with logarithmic complexity, making them ideal for matching high-dimensional CSI vectors in real-time. The architecture supports incremental updates, allowing the system to adapt to environmental changes without retraining, while product quantization keeps memory usage low enough for edge deployment on ESP32 or WASM runtimes.

How does RuView handle the high dimensionality of CSI data?

RuView encodes raw CSI frames into 329-dimensional vectors through the CsiFeatures pipeline, which captures the essential characteristics of the radio environment. The HNSW index uses cosine distance to measure angular similarity between these normalized embeddings, and PQ8 quantization compresses the storage by 8× while preserving search accuracy above 95%.

Can the room fingerprint index be updated without stopping the system?

Yes. The architecture supports incremental learning where confirmed detections are inserted directly into the human_patterns or empty_patterns indices. The system updates neighbor links through the GNN layer (as described in ADR-006) and periodically re-balances the graph layers to maintain search quality, all without requiring a full index rebuild or service restart.

What performance can be expected with large-scale deployments?

With ef_search=64, the system queries 100,000 vectors in 0.3ms and 1,000,000 vectors in 0.9ms, achieving 107× to 356× speedup over brute-force search. Memory consumption remains manageable through PQ8 compression: a 1-million-vector index requires approximately 410MB total storage compared to 1.29GB uncompressed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →