# How the ModalityTranslator Converts CSI Features in WiFi DensePose

> Learn how ModalityTranslator converts CSI features using convolutional blocks and attention mechanisms. Discover its process for WiFi DensePose.

- Repository: [rUv/wifi-densepose](https://github.com/ruvnet/wifi-densepose)
- Tags: tutorial
- Published: 2026-02-16

---

**The ModalityTranslator converts CSI features by encoding raw Channel State Information through convolutional blocks, optionally applying multi-head attention for long-range dependencies, and decoding to visual-like tensors using transposed convolutions.**

The `ModalityTranslationNetwork` in the [ruvnet/wifi-densepose](https://github.com/ruvnet/wifi-densepose) repository bridges the gap between wireless sensing and computer vision. It transforms raw Channel State Information (CSI) tensors into feature representations that downstream pose estimation modules can process as if they were visual inputs.

## Architecture Overview of the ModalityTranslator

The translation pipeline implemented in [`v1/src/models/modality_translation.py`](https://github.com/ruvnet/wifi-densepose/blob/main/v1/src/models/modality_translation.py) follows an encoder-decoder pattern with an optional attention mechanism. The network processes input tensors through three distinct stages: spatial encoding, temporal/spatial attention, and visual decoding.

## Stage 1: Encoding CSI into Abstract Feature Maps

The encoder, constructed by `_build_encoder` (lines 71-86), processes raw CSI tensors through a stack of convolutional blocks. Each block consists of:

- A `Conv2d` layer for spatial feature extraction
- Configurable normalization (BatchNorm by default)
- **ReLU** activation
- `Dropout2d` for regularization

The first encoder block maintains the original spatial resolution, while subsequent blocks use stride-2 convolutions to downsample the feature maps. This progressively increases the channel depth from `input_channels` to `hidden_channels`, creating abstract representations that capture high-level CSI patterns.

## Stage 2: Capturing Long-Range Dependencies with Multi-Head Attention

When `use_attention=True` in the configuration, the network inserts a `nn.MultiheadAttention` module built by `_build_attention` (lines 47-53). This stage processes the encoded features as follows:

1. The final encoder output is reshaped from `(batch, channels, height, width)` to a sequence format `(batch, seq_len, embed_dim)`
2. The multi-head attention module processes this sequence to capture long-range spatial dependencies across the CSI tensor
3. The output is reshaped back to 4-D feature maps before passing to the decoder

This attention mechanism allows the model to correlate distant subcarriers or antenna pairs that may jointly indicate specific body poses, addressing the limited receptive field of pure convolutional processing.

## Stage 3: Decoding to Visual-Like Representations

The decoder, implemented in `_build_decoder` (lines 91-119), mirrors the encoder architecture but uses `ConvTranspose2d` layers for upsampling. The decoding process:

- Progressively upsamples feature maps back to the target spatial dimensions
- Maintains the same normalization, activation, and dropout pattern as the encoder
- Applies a final **Tanh** activation (lines 14-19) to constrain output values to the range `[-1, 1]`

The resulting tensor has shape `(batch, output_channels, H, W)`, compatible with standard computer vision backbones expecting visual feature maps.

## Implementation Example: Converting CSI Tensors

The following example demonstrates how to instantiate and use the `ModalityTranslationNetwork` to convert raw CSI data into visual-like features:

```python
import torch
from v1.src.models.modality_translation import ModalityTranslationNetwork

# Define configuration matching your CSI hardware setup

cfg = {
    "input_channels": 6,          # Number of CSI sub-carriers or antenna streams

    "hidden_channels": [32, 64],  # Two encoder stages: 32 then 64 channels

    "output_channels": 128,       # Dimensionality expected by pose estimation backbone

    "kernel_size": 3,
    "stride": 1,
    "padding": 1,
    "dropout_rate": 0.1,
    "activation": "relu",
    "normalization": "batch",
    "use_attention": True,        # Enable multi-head attention for long-range dependencies

    "attention_heads": 8
}

# Instantiate the translator

translator = ModalityTranslationNetwork(cfg)

# Create dummy CSI tensor: (batch, channels, height, width)

csi_tensor = torch.randn(4, cfg["input_channels"], 64, 64)

# Convert to visual-like features

visual_features = translator(csi_tensor)
print(visual_features.shape)  # Output: torch.Size([4, 128, 64, 64])

```

## Inspecting Intermediate Representations

For debugging or visualization purposes, you can extract intermediate feature maps from the encoder, decoder, and attention modules:

```python

# Get intermediate representations

intermediate = translator.get_intermediate_features(csi_tensor)

# Encoder outputs at each downsampling stage

encoder_features = intermediate["encoder_features"]  # List of tensors

# Decoder outputs at each upsampling stage  

decoder_features = intermediate["decoder_features"]  # List of tensors

# Attention weights (if use_attention=True)

attention_weights = intermediate.get("attention_weights")  # Shape: (batch, heads, seq, seq)

```

## Summary

- The **ModalityTranslator** (`ModalityTranslationNetwork`) converts raw CSI tensors into visual-like feature maps through an encoder-decoder architecture.
- The **encoder** (`_build_encoder`) uses strided convolutions with ReLU activations and dropout to extract abstract features while reducing spatial dimensions.
- **Multi-head attention** (`_build_attention`) optionally processes encoded features as sequences to capture long-range spatial dependencies across CSI subcarriers.
- The **decoder** (`_build_decoder`) employs transposed convolutions to upsample features back to the target resolution, applying Tanh activation to normalize outputs to `[-1, 1]`.
- The resulting tensors match the expected input format of standard computer vision pose estimation backbones, enabling WiFi-based dense pose estimation.

## Frequently Asked Questions

### What input format does the ModalityTranslator expect?

The ModalityTranslator expects input tensors with shape `(batch_size, input_channels, height, width)`, where `input_channels` corresponds to the number of CSI subcarriers or antenna streams from your WiFi hardware. The spatial dimensions (height and width) typically represent the subcarrier grouping or time-domain samples, depending on your preprocessing pipeline.

### How does the attention mechanism improve CSI feature conversion?

When `use_attention` is enabled, the network reshapes the encoded feature maps into sequences and applies multi-head self-attention before decoding. This allows the model to explicitly model relationships between distant CSI measurements—such as correlations between subcarriers affected by different body parts—that pure convolutional layers might miss due to their limited receptive fields. The attention weights can be extracted for interpretability analysis.

### Can I modify the encoder depth or channel dimensions?

Yes, the encoder and decoder architectures are fully configurable through the `hidden_channels` list parameter in the configuration dictionary. Each integer in the list adds an encoder stage with that many output channels, and the decoder automatically mirrors this structure in reverse. You can also adjust the `kernel_size`, `dropout_rate`, and `normalization` type (batch or instance) to suit your specific CSI dataset characteristics.

### What is the output range of the converted features?

The final decoder layer applies a **Tanh** activation function, which constrains the output values to the range `[-1, 1]`. This normalization ensures compatibility with pre-trained computer vision backbones that expect visual feature maps within standard normalized ranges, allowing the WiFi-derived features to seamlessly replace or supplement RGB image features in pose estimation pipelines.