# SEANet Encoder/Decoder Configuration in NVIDIA PersonaPlex: Architecture and Implementation

> Explore the SEANet encoder/decoder configuration in NVIDIA PersonaPlex, featuring 320x compression, 128D latent representations, residual blocks, ELU activation, and streaming-compatible CNNs for real-time audio.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: architecture
- Published: 2026-04-07

---

**The SEANet encoder/decoder configuration in PersonaPlex defaults to a 320× compression ratio with 128-dimensional latent representations, utilizing configurable residual blocks, ELU activation, and streaming-compatible convolutional layers for real-time audio processing.**

The SEANet architecture serves as the backbone of NVIDIA's PersonaPlex audio codec, providing efficient neural audio coding through highly configurable encoder-decoder pairs. Implemented in [`moshi/moshi/modules/seanet.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/modules/seanet.py), both the `SEANetEncoder` and `SEANetDecoder` classes expose extensive hyperparameters controlling network depth, dilation strategies, and streaming behavior. This architecture enables high-quality audio compression with support for causal inference and block-wise processing without full-signal buffering.

## SEANet Architecture Overview

The SEANet implementation consists of three core components: the encoder, decoder, and residual blocks. The architecture leverages custom streaming primitives from [`moshi/moshi/modules/conv.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/modules/conv.py) to enable inference on arbitrary-length sequences.

**Key architectural elements include:**

- **Streaming-ready primitives**: Uses `StreamingConv1d`, `StreamingConvTranspose1d`, and `StreamingAdd` for block-wise processing
- **Residual blocks**: `SEANetResnetBlock` (lines 42-73 in [`seanet.py`](https://github.com/NVIDIA/personaplex/blob/main/seanet.py)) implements skip connections with configurable dilation and compression
- **Causal support**: All convolutions support strict causality via the `causal` parameter
- **Flexible normalization**: Optional normalization via the `norm` parameter (default `"none"`) with outer block disabling via `disable_norm_outer_blocks`

## SEANetEncoder Configuration Parameters

The `SEANetEncoder` compresses raw audio into a latent representation using downsampling stages interleaved with residual blocks.

**Default configuration values:**

- `channels`: `1` (mono input)
- `dimension`: `128` (latent vector size)
- `n_filters`: `32` (base filter count)
- `n_residual_layers`: `3` (residual blocks per stage)
- `ratios`: `[8, 5, 4, 2]` (downsampling factors)
- `activation`: `"ELU"` with `activation_params`: `{"alpha": 1.0}`
- `kernel_size` / `last_kernel_size`: `7`
- `residual_kernel_size`: `3`
- `dilation_base`: `2` (for exponential dilation growth)
- `causal`: `False`
- `pad_mode`: `"reflect"`
- `true_skip`: `True` (identity skip connections)
- `compress`: `2` (channel compression factor)
- `mask_fn` / `mask_position`: `None` (optional masking)

**Implementation specifics:**

The encoder reverses the ratios list during initialization: `self.ratios = list(reversed(ratios))` (lines 76-78). The total hop length calculates to 320 samples via `self.hop_length = int(np.prod(self.ratios))`, meaning a 16 kHz input produces latent frames at 50 Hz.

## SEANetDecoder Configuration Parameters

The `SEANetDecoder` reconstructs audio from latent representations, mirroring the encoder architecture using transposed convolutions for upsampling.

**Default configuration values:**

- `channels`: `1` (mono output)
- `dimension`: `128` (input latent size)
- `n_filters`: `32`
- `n_residual_layers`: `3`
- `ratios`: `[8, 5, 4, 2]` (upsampling factors, applied in forward order)
- `activation`: `"ELU"`
- `final_activation`: `None` (optional output non-linearity)
- `final_activation_params`: `None`
- `kernel_size` / `last_kernel_size`: `7`
- `residual_kernel_size`: `3`
- `dilation_base`: `2`
- `causal`: `False`
- `pad_mode`: `"reflect"`
- `true_skip`: `True`
- `compress`: `2`
- `disable_norm_outer_blocks`: `0`
- `trim_right_ratio`: `1.0` (controls right-side trimming for causal transposed convolutions)

The decoder begins with a projection from `dimension` to `2**len(ratios) * n_filters`, then iteratively upsamples through each ratio stage, halving the filter multiplier after each `StreamingConvTranspose1d` operation.

## Residual Block and Streaming Implementation

Each **SEANetResnetBlock** contains:

1. Two convolutional layers with kernel sizes `residual_kernel_size` (default 3) and `1`
2. Exponential dilation: `dilation_base ** layer_index`
3. Optional channel compression by the `compress` factor (default 2)
4. Identity skip connections when `true_skip=True`, otherwise 1×1 convolutions

**Streaming capabilities** derive from [`moshi/moshi/modules/streaming.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/modules/streaming.py), allowing models to process arbitrary-length sequences via `StreamingContainer` state management. When `causal=True`, all convolutions use appropriate padding to ensure no future dependencies, critical for real-time applications.

## Practical Configuration Examples

### Instantiating Default Encoder and Decoder

```python
import torch
from moshi.moshi.modules.seanet import SEANetEncoder, SEANetDecoder

# Default configuration: 320× compression, 128-dim latent

encoder = SEANetEncoder()
decoder = SEANetDecoder()

# Process 16 kHz mono audio (batch=1, channels=1, samples=16000)

waveform = torch.randn(1, 1, 16000)

# Encode to latent space

z = encoder(waveform)  # Shape: (1, 128, 50)

print(f"Latent shape: {z.shape}")  # 16000 / 320 = 50 frames

# Decode back to audio

recon = decoder(z)  # Shape: (1, 1, 16000)

print(f"Reconstructed shape: {recon.shape}")

```

### Custom Configuration for Streaming Applications

```python

# Causal configuration for real-time processing

encoder = SEANetEncoder(
    channels=2,              # Stereo input

    n_filters=64,            # Wider filters

    ratios=[4, 4, 4, 4],     # 256× total downsampling

    causal=True,             # Strict causality

    pad_mode="constant",     # Alternative padding

)

decoder = SEANetDecoder(
    channels=2,
    n_filters=64,
    ratios=[4, 4, 4, 4],
    causal=True,
    final_activation="Tanh", # Bound output to [-1, 1]

    trim_right_ratio=1.0,
)

```

### Accessing Hop Length for Alignment

```python

# Calculate temporal downsampling factor

hop_length = encoder.hop_length  # Returns 320 for default ratios

print(f"Hop length: {hop_length} samples")

# Useful for aligning latent features with spectrograms

```

### Adding Spectral Masking

```python
class EnergyBasedMask(torch.nn.Module):
    def forward(self, x):
        # Mask low-energy frames

        mask = (x.abs().mean(dim=1, keepdim=True) > 0.01).float()
        return x * mask

# Inject mask after specific layer

encoder = SEANetEncoder(
    mask_fn=EnergyBasedMask(),
    mask_position=2  # Apply after second downsampling block

)

```

## Summary

- **Location**: The SEANet encoder/decoder configuration resides in [`moshi/moshi/modules/seanet.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/modules/seanet.py) within the PersonaPlex repository
- **Compression**: Default configuration achieves 320× compression (ratios 8×5×4×2) with 128-dimensional latent vectors
- **Flexibility**: Both encoder and decoder support customizable channels, filter depths, dilation strategies, and normalization schemes
- **Streaming**: Native support for block-wise inference via `StreamingConv1d` and `StreamingConvTranspose1d` enables real-time processing
- **Causality**: Optional causal mode ensures no future dependencies for streaming applications

## Frequently Asked Questions

### What is the default compression ratio of the SEANet encoder in PersonaPlex?

The default SEANet encoder achieves a **320× compression ratio** calculated from the product of the default ratios `[8, 5, 4, 2]`. With a 16 kHz input sampling rate, this produces latent features at 50 Hz with a hop length of 320 samples, implemented via `self.hop_length = int(np.prod(self.ratios))` in the encoder initialization.

### How does the SEANet encoder handle the ratios parameter differently from the decoder?

The encoder **reverses the ratios list** during initialization (`self.ratios = list(reversed(ratios))` at lines 76-78) to ensure the downsampling order matches the decoder's upsampling order. While the decoder applies ratios `[8, 5, 4, 2]` sequentially for upsampling, the encoder applies them in reverse order `[2, 4, 5, 8]` for downsampling, maintaining architectural symmetry between compression and reconstruction paths.

### What streaming capabilities are built into the SEANet architecture?

SEANet utilizes **streaming-compatible convolution layers** including `StreamingConv1d`, `StreamingConvTranspose1d`, and `StreamingAdd` from [`moshi/moshi/modules/conv.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/modules/conv.py). These primitives, managed by `StreamingContainer` in [`moshi/moshi/modules/streaming.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/modules/streaming.py), enable the encoder and decoder to process audio in blocks without buffering entire sequences, supporting real-time inference with configurable causal constraints.

### Can normalization be disabled in specific layers of the SEANet encoder?

Yes, the `disable_norm_outer_blocks` parameter allows skipping normalization in the outermost blocks of both encoder and decoder. When set to a positive integer (e.g., `1` or `2`), normalization is disabled for that many blocks counting from the input/output edges toward the center, useful for stabilizing training or matching specific deployment constraints where normalization may cause artifacts.