How the ModalityTranslator Converts CSI Features in WiFi DensePose

The ModalityTranslator converts CSI features by encoding raw Channel State Information through convolutional blocks, optionally applying multi-head attention for long-range dependencies, and decoding to visual-like tensors using transposed convolutions.

The ModalityTranslationNetwork in the ruvnet/wifi-densepose repository bridges the gap between wireless sensing and computer vision. It transforms raw Channel State Information (CSI) tensors into feature representations that downstream pose estimation modules can process as if they were visual inputs.

Architecture Overview of the ModalityTranslator

The translation pipeline implemented in v1/src/models/modality_translation.py follows an encoder-decoder pattern with an optional attention mechanism. The network processes input tensors through three distinct stages: spatial encoding, temporal/spatial attention, and visual decoding.

Stage 1: Encoding CSI into Abstract Feature Maps

The encoder, constructed by _build_encoder (lines 71-86), processes raw CSI tensors through a stack of convolutional blocks. Each block consists of:

  • A Conv2d layer for spatial feature extraction
  • Configurable normalization (BatchNorm by default)
  • ReLU activation
  • Dropout2d for regularization

The first encoder block maintains the original spatial resolution, while subsequent blocks use stride-2 convolutions to downsample the feature maps. This progressively increases the channel depth from input_channels to hidden_channels, creating abstract representations that capture high-level CSI patterns.

Stage 2: Capturing Long-Range Dependencies with Multi-Head Attention

When use_attention=True in the configuration, the network inserts a nn.MultiheadAttention module built by _build_attention (lines 47-53). This stage processes the encoded features as follows:

  1. The final encoder output is reshaped from (batch, channels, height, width) to a sequence format (batch, seq_len, embed_dim)
  2. The multi-head attention module processes this sequence to capture long-range spatial dependencies across the CSI tensor
  3. The output is reshaped back to 4-D feature maps before passing to the decoder

This attention mechanism allows the model to correlate distant subcarriers or antenna pairs that may jointly indicate specific body poses, addressing the limited receptive field of pure convolutional processing.

Stage 3: Decoding to Visual-Like Representations

The decoder, implemented in _build_decoder (lines 91-119), mirrors the encoder architecture but uses ConvTranspose2d layers for upsampling. The decoding process:

  • Progressively upsamples feature maps back to the target spatial dimensions
  • Maintains the same normalization, activation, and dropout pattern as the encoder
  • Applies a final Tanh activation (lines 14-19) to constrain output values to the range [-1, 1]

The resulting tensor has shape (batch, output_channels, H, W), compatible with standard computer vision backbones expecting visual feature maps.

Implementation Example: Converting CSI Tensors

The following example demonstrates how to instantiate and use the ModalityTranslationNetwork to convert raw CSI data into visual-like features:

import torch
from v1.src.models.modality_translation import ModalityTranslationNetwork

# Define configuration matching your CSI hardware setup

cfg = {
    "input_channels": 6,          # Number of CSI sub-carriers or antenna streams

    "hidden_channels": [32, 64],  # Two encoder stages: 32 then 64 channels

    "output_channels": 128,       # Dimensionality expected by pose estimation backbone

    "kernel_size": 3,
    "stride": 1,
    "padding": 1,
    "dropout_rate": 0.1,
    "activation": "relu",
    "normalization": "batch",
    "use_attention": True,        # Enable multi-head attention for long-range dependencies

    "attention_heads": 8
}

# Instantiate the translator

translator = ModalityTranslationNetwork(cfg)

# Create dummy CSI tensor: (batch, channels, height, width)

csi_tensor = torch.randn(4, cfg["input_channels"], 64, 64)

# Convert to visual-like features

visual_features = translator(csi_tensor)
print(visual_features.shape)  # Output: torch.Size([4, 128, 64, 64])

Inspecting Intermediate Representations

For debugging or visualization purposes, you can extract intermediate feature maps from the encoder, decoder, and attention modules:


# Get intermediate representations

intermediate = translator.get_intermediate_features(csi_tensor)

# Encoder outputs at each downsampling stage

encoder_features = intermediate["encoder_features"]  # List of tensors

# Decoder outputs at each upsampling stage  

decoder_features = intermediate["decoder_features"]  # List of tensors

# Attention weights (if use_attention=True)

attention_weights = intermediate.get("attention_weights")  # Shape: (batch, heads, seq, seq)

Summary

  • The ModalityTranslator (ModalityTranslationNetwork) converts raw CSI tensors into visual-like feature maps through an encoder-decoder architecture.
  • The encoder (_build_encoder) uses strided convolutions with ReLU activations and dropout to extract abstract features while reducing spatial dimensions.
  • Multi-head attention (_build_attention) optionally processes encoded features as sequences to capture long-range spatial dependencies across CSI subcarriers.
  • The decoder (_build_decoder) employs transposed convolutions to upsample features back to the target resolution, applying Tanh activation to normalize outputs to [-1, 1].
  • The resulting tensors match the expected input format of standard computer vision pose estimation backbones, enabling WiFi-based dense pose estimation.

Frequently Asked Questions

What input format does the ModalityTranslator expect?

The ModalityTranslator expects input tensors with shape (batch_size, input_channels, height, width), where input_channels corresponds to the number of CSI subcarriers or antenna streams from your WiFi hardware. The spatial dimensions (height and width) typically represent the subcarrier grouping or time-domain samples, depending on your preprocessing pipeline.

How does the attention mechanism improve CSI feature conversion?

When use_attention is enabled, the network reshapes the encoded feature maps into sequences and applies multi-head self-attention before decoding. This allows the model to explicitly model relationships between distant CSI measurements—such as correlations between subcarriers affected by different body parts—that pure convolutional layers might miss due to their limited receptive fields. The attention weights can be extracted for interpretability analysis.

Can I modify the encoder depth or channel dimensions?

Yes, the encoder and decoder architectures are fully configurable through the hidden_channels list parameter in the configuration dictionary. Each integer in the list adds an encoder stage with that many output channels, and the decoder automatically mirrors this structure in reverse. You can also adjust the kernel_size, dropout_rate, and normalization type (batch or instance) to suit your specific CSI dataset characteristics.

What is the output range of the converted features?

The final decoder layer applies a Tanh activation function, which constrains the output values to the range [-1, 1]. This normalization ensures compatibility with pre-trained computer vision backbones that expect visual feature maps within standard normalized ranges, allowing the WiFi-derived features to seamlessly replace or supplement RGB image features in pose estimation pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →