WiFi DensePose Data Flow: From Raw CSI to Human Pose Estimation

WiFi DensePose transforms raw Channel State Information (CSI) into 2D human poses through a four-stage pipeline: hardware acquisition, signal preprocessing, feature extraction with human detection, and neural modality translation to DensePose outputs.

The WiFi DensePose architecture enables pose estimation using standard WiFi signals instead of cameras. This article examines how data flows through the ruvnet/wifi-densepose repository, tracing the journey from raw CSI packets to dense human body part segmentation.

Overview of the WiFi DensePose Pipeline

The WiFi DensePose data flow consists of four logical stages, each implemented by dedicated modules in the codebase:

  1. Hardware acquisition – Captures raw CSI packets from ESP-32 sensors or routers
  2. Signal-level preprocessing – Cleans and normalizes amplitude and phase matrices
  3. Feature extraction and human detection – Derives compact descriptors and validates human presence
  4. Pose estimation – Translates CSI features into visual-domain tensors for DensePose inference

Stage 1: Hardware Acquisition and CSI Extraction

The pipeline begins in src/hardware/csi_extractor.py, which handles communication with WiFi hardware and raw byte parsing.

Connecting to WiFi Hardware

The CSIExtractor class establishes hardware interfaces through the connect() method. This supports both ESP-32-based sensors and standard routers, automatically selecting the appropriate protocol handler.

Parsing Raw CSI Bytes

Once connected, extract_csi() reads raw byte frames from the hardware interface. The system uses hardware-specific parsers—ESP32CSIParser or RouterCSIParser—to convert these bytes into structured CSIData objects. Each dataclass contains:

  • Timestamp
  • Amplitude matrix
  • Phase measurements
  • Metadata (frequency, bandwidth, antenna configuration)

Stage 2: Signal-Level Preprocessing

Raw CSI contains environmental noise and hardware artifacts. The src/core/csi_processor.py module handles normalization through three private methods: _remove_noise, _apply_windowing, and _normalize_amplitude.

Noise Removal and Windowing

The processor applies a noise mask based on configurable dB thresholds to eliminate low-power interference. It then applies a Hamming window to reduce spectral leakage before further processing.

Amplitude Normalization

Finally, the amplitude matrix is scaled to unit variance using _normalize_amplitude. This standardization ensures consistent feature magnitudes regardless of transmitter power or distance variations.

Stage 3: Feature Extraction and Human Detection

The same csi_processor.py module extracts descriptive features and determines whether a human subject is present in the sensing area.

Computing CSI Features

The _extract_* methods compute multiple signal descriptors:

  • Amplitude mean and variance across subcarriers
  • Phase-difference matrices
  • Cross-correlation between antenna pairs
  • Doppler shifts indicating motion
  • Power spectral density (PSD) distribution

Human Presence Detection

The detect_human_presence() method aggregates these features into a confidence score. This score is smoothed over time and compared against human_detection_threshold from the configuration. Only frames exceeding this threshold proceed to pose estimation, reducing computational load and false positives.

Stage 4: Pose Estimation via Modality Translation

When human presence is confirmed, the pipeline enters the neural inference stage orchestrated by src/services/pose_service.py.

Phase Sanitization

Before neural processing, phase measurements undergo rigorous cleaning in src/core/phase_sanitizer.py. The PhaseSanitizer.sanitize_phase() method:

  • Unwraps phase discontinuities
  • Removes statistical outliers
  • Applies temporal smoothing
  • Implements low-pass filtering to remove high-frequency noise

CSI-to-Visual Feature Translation

The cleaned phase and normalized amplitude are concatenated and fed into ModalityTranslationNetwork (src/models/modality_translation.py). This encoder-decoder architecture bridges the gap between RF sensing and computer vision domains.

The network configuration uses:

  • input_channels=64
  • hidden_channels=[128, 256, 512]
  • output_channels=256

The output is a 256-channel visual-like feature map compatible with standard pose estimation architectures.

DensePose Head Inference

The feature map enters DensePoseHead (src/models/densepose_head.py), which produces:

  • Body-part segmentation logits (identifying anatomical regions)
  • UV-coordinate heatmaps (dense surface correspondence)

Post-processing converts these tensors into a list of pose dictionaries containing person IDs, confidence scores, and keypoint coordinates.

End-to-End Implementation Example

The process_csi_data() method in src/services/pose_service.py serves as the main entry point. Below is a complete example demonstrating the WiFi DensePose data flow with synthetic CSI data:

import numpy as np
from datetime import datetime
from src.config.settings import Settings
from src.config.domains import DomainConfig
from src.services.pose_service import PoseService

# ① Load configuration (replace with your own settings if needed)

settings = Settings()                # reads defaults from src/config/settings.py

domain_cfg = DomainConfig()          # domain‑specific limits (e.g., max persons)

# ② Initialise the service

pose_service = PoseService(settings, domain_cfg)
await pose_service.initialize()      # sets up CSIProcessor, PhaseSanitizer, models

# ③ Create a mock CSI frame (amplitude matrix)

csi_frame = np.random.rand(56, 3)   # 56 sub‑carriers × 3 antennas

metadata = {
    "timestamp": datetime.utcnow(),
    "frequency": 5.8e9,            # 5.8 GHz Wi‑Fi band

    "bandwidth": 20e6,
    "num_subcarriers": 56,
    "num_antennas": 3,
    "snr": 22.0,
}

# ④ Process the frame – the method returns a dict with pose data

result = await pose_service.process_csi_data(csi_frame, metadata)

print("Pose estimation result:")
print(f"  Time: {result['timestamp']}")
print(f"  Detected poses: {len(result['poses'])}")
for p in result['poses']:
    print(f"  • Person {p['person_id']} – confidence {p['confidence']:.2f}")

Key implementation details from the example:

  • The CSIProcessor runs automatically inside process_csi_data.
  • Phase sanitisation is applied via PhaseSanitizer before feeding the data to the translation network.
  • ModalityTranslationNetwork (config input_channels=64, hidden_channels=[128, 256, 512], output_channels=256) converts the CSI tensor into a visual-like representation that the DensePoseHead consumes.

Key Files in the WiFi DensePose Architecture

File Role in the data flow GitHub link
src/hardware/csi_extractor.py Parses raw bytes from ESP‑32 or router into CSIData. csi_extractor.py
src/core/csi_processor.py Noise removal, windowing, normalisation, feature extraction, human detection. csi_processor.py
src/core/phase_sanitizer.py Unwraps, outlier‑removes, smooths and low‑pass filters phase matrices. phase_sanitizer.py
src/services/pose_service.py Orchestrates CSI processing, phase sanitisation, modality translation, and DensePose inference. pose_service.py
src/models/modality_translation.py Neural encoder‑decoder that maps CSI tensors to visual‑domain feature maps (with optional attention). modality_translation.py
src/models/densepose_head.py DensePose head: body‑part segmentation + UV coordinate regression. densepose_head.py
src/api/websocket/pose_stream.py Exposes the end‑to‑end pipeline over a WebSocket endpoint for real‑time streaming. pose_stream.py
src/config/settings.py Central configuration (sampling rates, thresholds, model paths) that drives the whole flow. settings.py

These files together implement the end‑to‑end transformation (raw CSI) → (cleaned CSI) → (feature vector) → (visual feature map) → (pose estimation) that defines the WiFi DensePose architecture.

Summary

The WiFi DensePose data flow transforms invisible radio signals into structured human pose data through four deterministic stages:

  • Hardware abstraction via src/hardware/csi_extractor.py standardizes raw bytes from ESP-32 chips or routers into CSIData objects.
  • Signal conditioning in src/core/csi_processor.py applies noise masking, Hamming windowing, and unit-variance normalization to prepare clean amplitude and phase matrices.
  • Feature engineering and gating computes Doppler shifts, correlations, and PSD values to trigger human detection before expensive neural inference.
  • Modality translation uses ModalityTranslationNetwork and DensePoseHead to bridge RF sensing and computer vision domains, outputting body-part segmentation and UV coordinates.

All stages are orchestrated by src/services/pose_service.py, which exposes a unified process_csi_data() interface for real-time streaming applications.

Frequently Asked Questions

How does WiFi DensePose handle different hardware sources?

The architecture abstracts hardware differences through the CSIExtractor class in src/hardware/csi_extractor.py. The system supports both ESP-32-based sensors and standard routers by selecting the appropriate parser—ESP32CSIParser or RouterCSIParser—based on the configuration. Both parsers normalize raw bytes into the standard CSIData dataclass, ensuring downstream processing remains hardware-agnostic.

What preprocessing steps are applied to raw CSI signals?

Raw CSI undergoes three critical preprocessing steps in src/core/csi_processor.py. First, _remove_noise applies a dB threshold mask to eliminate low-power interference. Second, _apply_windowing uses a Hamming window to reduce spectral leakage in the frequency domain. Finally, _normalize_amplitude scales the amplitude matrix to unit variance, ensuring consistent signal magnitudes regardless of transmitter power or environmental attenuation.

How does the system determine if a human is present before running pose estimation?

The system uses a lightweight detection gate in src/core/csi_processor.py to avoid unnecessary neural inference. The detect_human_presence() method computes statistical features including amplitude mean/variance, phase-difference matrices, cross-correlation between antennas, Doppler shifts, and power spectral density. These features are aggregated into a confidence score that is temporally smoothed and compared against the human_detection_threshold from src/config/settings.py. Only frames exceeding this threshold proceed to the expensive modality translation and DensePose inference stages.

What is the role of the ModalityTranslationNetwork in the data flow?

The ModalityTranslationNetwork in src/models/modality_translation.py serves as the critical bridge between radio-frequency sensing and computer vision domains. This encoder-decoder architecture takes concatenated amplitude and sanitized phase tensors from the CSI domain and maps them into 256-channel visual-like feature maps that the DensePose head can process. The network configuration uses input_channels=64, hidden_channels=[128, 256, 512], and output_channels=256, effectively translating sparse RF signatures into dense spatial representations suitable for body-part segmentation and UV coordinate regression.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →