17-Keypoint Pose Estimation Architecture in RuView: Wi-Fi Dense Pose Implementation

RuView implements a three-stage neural network that converts raw Wi-Fi CSI tensors into 17 COCO-format human keypoints using a ResNet-18 backbone and a dedicated convolutional keypoint head.

The 17-keypoint pose estimation architecture in RuView enables Wi-Fi-based human pose detection without cameras. Built in Rust with the tch crate, the system processes channel state information (CSI) amplitude and phase data through a modality translator, extracts features via a lightweight ResNet-18-style backbone, and decodes 17 joint heatmaps through a specialized KeypointHead. This article examines the exact implementation found in the wifi-densepose-train crate, referencing specific file paths and line numbers from the RuView repository.

Overview of the 17-Keypoint Pose Estimation Pipeline

The architecture follows a modular encoder-decoder pattern optimized for Wi-Fi sensing. Raw CSI tensors undergo three transformations before producing normalized (x, y) coordinates for each joint.

Modality Translation Stage

The ModalityTranslator converts raw CSI amplitude and phase tensors into a 3-channel image-like representation with shape [B, 3, 48, 48]. This translation bridges the gap between RF signal processing and computer vision architectures, allowing standard convolutional networks to operate on Wi-Fi data.

ResNet-18 Backbone Feature Extractor

A lightweight ResNet-18-style network processes the translated modality. The backbone outputs a feature tensor with 256 channels, providing a rich spatial representation while maintaining computational efficiency suitable for real-time Wi-Fi sensing applications.

Keypoint Head Convolutional Decoder

The KeypointHead receives the 256-channel backbone features and decodes them into per-joint heatmaps. This dedicated head uses a three-layer convolutional architecture to progressively reduce channel dimensions while preserving spatial resolution, ultimately producing heatmaps for all 17 COCO keypoints.

KeypointHead Implementation Details

Located in rust-port/wifi-densepose-rs/crates/wifi-densepose-train/src/model.rs (lines 698-755), the KeypointHead struct implements the final decoding stage of the 17-keypoint pose estimation architecture.

Layer Architecture

The head consists of three convolution-batch normalization blocks defined at lines 698-704:

  • conv1 – 3×3 convolution with 256 output channels, padding 1, no bias
  • conv2 – 3×3 convolution with 128 output channels, padding 1, no bias
  • out_conv – 1×1 convolution mapping to num_keypoints channels (default 17)

During model construction (lines 96-100), the head is instantiated with the backbone channel count and configured keypoint count:

let kp_head = KeypointHead::new(
    &root / "kp_head",
    config.backbone_channels as i64,
    config.num_keypoints as i64,
);

Forward Pass and Heatmap Generation

The forward_t method (lines 744-755) processes backbone features through the three convolutional layers with batch normalization and ReLU activations:

let h = x
    .apply(&self.conv1)
    .apply_t(&self.bn1, train)
    .relu()
    .apply(&self.conv2)
    .apply_t(&self.bn2, train)
    .relu()
    .apply(&self.out_conv);

h.upsample_bilinear2d(&[heatmap_size, heatmap_size], false, None, None)

The output tensor has shape [B, 17, H, W], where H = W = heatmap_size (typically 8×8). These low-resolution heatmaps represent the probability distribution of each keypoint's location.

Converting Heatmaps to Normalized Coordinates

The final conversion from heatmaps to (x, y) coordinates occurs in trainer.rs via the heatmap_to_keypoints function (lines 608-627). This implementation uses argmax extraction followed by normalization:

let flat = heatmaps.reshape([batch, num_kp, h * w]);
let arg = flat.argmax(-1, false);
let row = (&arg / w).to_kind(Kind::Float);
let col = (&arg % w).to_kind(Kind::Float);
let x = col / (w - 1) as f64;
let y = row / (h - 1) as f64;
Tensor::stack(&[x, y], -1)

The resulting tensor has shape [B, 17, 2], where each keypoint is represented by normalized coordinates in the range [0, 1]. These coordinates correspond to the 17 COCO keypoints: nose, eyes, ears, shoulders, elbows, wrists, hips, knees, and ankles.

Configuration and Integration

The 17-keypoint default is configured in v1/src/config/domains.py, which sets num_keypoints = 17 for COCO-format pose estimation. The v1/src/services/pose_service.py exposes these keypoints via the HTTP API, allowing downstream applications to consume Wi-Fi-based pose data without processing raw heatmaps.

Practical Usage Example

The following Rust example demonstrates building the model and extracting 17-keypoint coordinates from dummy CSI data:

use wifi_densepose_train::{
    config::TrainingConfig,
    model::WiFiDensePoseModel,
};
use tch::{Device, Tensor};

fn main() -> anyhow::Result<()> {
    // 1️⃣ Load a training configuration (default includes 17 keypoints)
    let cfg = TrainingConfig::default();

    // 2️⃣ Build the model on the desired device (CPU or GPU)
    let model = WiFiDensePoseModel::new(&cfg, Device::Cpu);

    // 3️⃣ Dummy CSI tensors (replace with real amplitude/phase data)
    let amp = Tensor::rand(&[1, cfg.window_frames * cfg.num_antennas_tx * cfg.num_antennas_rx,
                            cfg.num_subcarriers as i64], (tch::Kind::Float, Device::Cpu));
    let phase = Tensor::rand_like(&amp);

    // 4️⃣ Forward pass (inference mode)
    let out = model.forward_inference(&amp, &phase);

    // 5️⃣ Extract the 17‑keypoint heatmaps
    let heatmaps = out.keypoints; // shape [1, 17, H, W]

    // 6️⃣ Convert to normalized coordinates
    let keypoints = wifi_densepose_train::trainer::heatmap_to_keypoints(&heatmaps);
    // keypoints shape [1, 17, 2] → (x, y) ∈ [0, 1]
    println!("Keypoints: {:?}", keypoints);
    Ok(())
}

This example shows the complete pipeline from raw Wi-Fi CSI tensors to normalized 17-keypoint coordinates, utilizing the KeypointHead architecture described above.

Summary

  • RuView's 17-keypoint pose estimation architecture processes Wi-Fi CSI data through a three-stage pipeline: modality translation, ResNet-18 backbone, and convolutional keypoint head.
  • The KeypointHead in model.rs (lines 698-755) uses three convolutional layers (256→128→17 channels) to generate low-resolution heatmaps.
  • Heatmap conversion occurs in trainer.rs (lines 608-627) via argmax extraction and normalization, producing [B, 17, 2] coordinate tensors.
  • Default configuration uses 17 COCO keypoints defined in domains.py, exposed via pose_service.py for API consumption.

Frequently Asked Questions

What are the 17 keypoints in RuView's pose estimation?

The 17 keypoints follow the COCO (Common Objects in Context) format: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. This standard format ensures compatibility with existing pose estimation datasets and visualization tools.

How does the KeypointHead generate heatmaps from Wi-Fi signals?

The KeypointHead receives 256-channel feature maps from the ResNet-18 backbone and processes them through three convolutional layers: a 3×3 conv with 256 channels, another 3×3 conv reducing to 128 channels, and a final 1×1 conv producing 17 output channels. Each layer includes batch normalization and ReLU activation. The output is bilinearly upsampled to the configured heatmap size (typically 8×8), resulting in per-joint probability distributions where each pixel represents the likelihood of a specific keypoint's location.

What is the output resolution of the pose estimation heatmaps?

By default, the KeypointHead produces heatmaps with spatial dimensions of 8×8 pixels (configurable via heatmap_size in the training configuration). Despite the low resolution, the model achieves sufficient accuracy because the subsequent heatmap_to_keypoints function in trainer.rs uses argmax extraction on the flattened heatmaps and normalizes the coordinates to the [0, 1] range, effectively sub-pixel precision through probabilistic localization.

How are the heatmaps converted to actual pixel coordinates?

The conversion happens in the heatmap_to_keypoints function within trainer.rs. The process reshapes the heatmap tensor from [B, 17, H, W] to [B, 17, H*W], then applies argmax along the last dimension to find the index of the maximum activation for each keypoint. These indices are converted to row and column coordinates via integer division and modulo operations, then normalized by dividing by (W - 1) and (H - 1) respectively. This yields coordinates in the range [0, 1], which can be scaled to any image resolution for visualization or downstream processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →