SEANet Encoder/Decoder Configuration in NVIDIA PersonaPlex: Architecture and Implementation

The SEANet encoder/decoder configuration in PersonaPlex defaults to a 320× compression ratio with 128-dimensional latent representations, utilizing configurable residual blocks, ELU activation, and streaming-compatible convolutional layers for real-time audio processing.

The SEANet architecture serves as the backbone of NVIDIA's PersonaPlex audio codec, providing efficient neural audio coding through highly configurable encoder-decoder pairs. Implemented in moshi/moshi/modules/seanet.py, both the SEANetEncoder and SEANetDecoder classes expose extensive hyperparameters controlling network depth, dilation strategies, and streaming behavior. This architecture enables high-quality audio compression with support for causal inference and block-wise processing without full-signal buffering.

SEANet Architecture Overview

The SEANet implementation consists of three core components: the encoder, decoder, and residual blocks. The architecture leverages custom streaming primitives from moshi/moshi/modules/conv.py to enable inference on arbitrary-length sequences.

Key architectural elements include:

  • Streaming-ready primitives: Uses StreamingConv1d, StreamingConvTranspose1d, and StreamingAdd for block-wise processing
  • Residual blocks: SEANetResnetBlock (lines 42-73 in seanet.py) implements skip connections with configurable dilation and compression
  • Causal support: All convolutions support strict causality via the causal parameter
  • Flexible normalization: Optional normalization via the norm parameter (default "none") with outer block disabling via disable_norm_outer_blocks

SEANetEncoder Configuration Parameters

The SEANetEncoder compresses raw audio into a latent representation using downsampling stages interleaved with residual blocks.

Default configuration values:

  • channels: 1 (mono input)
  • dimension: 128 (latent vector size)
  • n_filters: 32 (base filter count)
  • n_residual_layers: 3 (residual blocks per stage)
  • ratios: [8, 5, 4, 2] (downsampling factors)
  • activation: "ELU" with activation_params: {"alpha": 1.0}
  • kernel_size / last_kernel_size: 7
  • residual_kernel_size: 3
  • dilation_base: 2 (for exponential dilation growth)
  • causal: False
  • pad_mode: "reflect"
  • true_skip: True (identity skip connections)
  • compress: 2 (channel compression factor)
  • mask_fn / mask_position: None (optional masking)

Implementation specifics:

The encoder reverses the ratios list during initialization: self.ratios = list(reversed(ratios)) (lines 76-78). The total hop length calculates to 320 samples via self.hop_length = int(np.prod(self.ratios)), meaning a 16 kHz input produces latent frames at 50 Hz.

SEANetDecoder Configuration Parameters

The SEANetDecoder reconstructs audio from latent representations, mirroring the encoder architecture using transposed convolutions for upsampling.

Default configuration values:

  • channels: 1 (mono output)
  • dimension: 128 (input latent size)
  • n_filters: 32
  • n_residual_layers: 3
  • ratios: [8, 5, 4, 2] (upsampling factors, applied in forward order)
  • activation: "ELU"
  • final_activation: None (optional output non-linearity)
  • final_activation_params: None
  • kernel_size / last_kernel_size: 7
  • residual_kernel_size: 3
  • dilation_base: 2
  • causal: False
  • pad_mode: "reflect"
  • true_skip: True
  • compress: 2
  • disable_norm_outer_blocks: 0
  • trim_right_ratio: 1.0 (controls right-side trimming for causal transposed convolutions)

The decoder begins with a projection from dimension to 2**len(ratios) * n_filters, then iteratively upsamples through each ratio stage, halving the filter multiplier after each StreamingConvTranspose1d operation.

Residual Block and Streaming Implementation

Each SEANetResnetBlock contains:

  1. Two convolutional layers with kernel sizes residual_kernel_size (default 3) and 1
  2. Exponential dilation: dilation_base ** layer_index
  3. Optional channel compression by the compress factor (default 2)
  4. Identity skip connections when true_skip=True, otherwise 1×1 convolutions

Streaming capabilities derive from moshi/moshi/modules/streaming.py, allowing models to process arbitrary-length sequences via StreamingContainer state management. When causal=True, all convolutions use appropriate padding to ensure no future dependencies, critical for real-time applications.

Practical Configuration Examples

Instantiating Default Encoder and Decoder

import torch
from moshi.moshi.modules.seanet import SEANetEncoder, SEANetDecoder

# Default configuration: 320× compression, 128-dim latent

encoder = SEANetEncoder()
decoder = SEANetDecoder()

# Process 16 kHz mono audio (batch=1, channels=1, samples=16000)

waveform = torch.randn(1, 1, 16000)

# Encode to latent space

z = encoder(waveform)  # Shape: (1, 128, 50)

print(f"Latent shape: {z.shape}")  # 16000 / 320 = 50 frames

# Decode back to audio

recon = decoder(z)  # Shape: (1, 1, 16000)

print(f"Reconstructed shape: {recon.shape}")

Custom Configuration for Streaming Applications


# Causal configuration for real-time processing

encoder = SEANetEncoder(
    channels=2,              # Stereo input

    n_filters=64,            # Wider filters

    ratios=[4, 4, 4, 4],     # 256× total downsampling

    causal=True,             # Strict causality

    pad_mode="constant",     # Alternative padding

)

decoder = SEANetDecoder(
    channels=2,
    n_filters=64,
    ratios=[4, 4, 4, 4],
    causal=True,
    final_activation="Tanh", # Bound output to [-1, 1]

    trim_right_ratio=1.0,
)

Accessing Hop Length for Alignment


# Calculate temporal downsampling factor

hop_length = encoder.hop_length  # Returns 320 for default ratios

print(f"Hop length: {hop_length} samples")

# Useful for aligning latent features with spectrograms

Adding Spectral Masking

class EnergyBasedMask(torch.nn.Module):
    def forward(self, x):
        # Mask low-energy frames

        mask = (x.abs().mean(dim=1, keepdim=True) > 0.01).float()
        return x * mask

# Inject mask after specific layer

encoder = SEANetEncoder(
    mask_fn=EnergyBasedMask(),
    mask_position=2  # Apply after second downsampling block

)

Summary

  • Location: The SEANet encoder/decoder configuration resides in moshi/moshi/modules/seanet.py within the PersonaPlex repository
  • Compression: Default configuration achieves 320× compression (ratios 8×5×4×2) with 128-dimensional latent vectors
  • Flexibility: Both encoder and decoder support customizable channels, filter depths, dilation strategies, and normalization schemes
  • Streaming: Native support for block-wise inference via StreamingConv1d and StreamingConvTranspose1d enables real-time processing
  • Causality: Optional causal mode ensures no future dependencies for streaming applications

Frequently Asked Questions

What is the default compression ratio of the SEANet encoder in PersonaPlex?

The default SEANet encoder achieves a 320× compression ratio calculated from the product of the default ratios [8, 5, 4, 2]. With a 16 kHz input sampling rate, this produces latent features at 50 Hz with a hop length of 320 samples, implemented via self.hop_length = int(np.prod(self.ratios)) in the encoder initialization.

How does the SEANet encoder handle the ratios parameter differently from the decoder?

The encoder reverses the ratios list during initialization (self.ratios = list(reversed(ratios)) at lines 76-78) to ensure the downsampling order matches the decoder's upsampling order. While the decoder applies ratios [8, 5, 4, 2] sequentially for upsampling, the encoder applies them in reverse order [2, 4, 5, 8] for downsampling, maintaining architectural symmetry between compression and reconstruction paths.

What streaming capabilities are built into the SEANet architecture?

SEANet utilizes streaming-compatible convolution layers including StreamingConv1d, StreamingConvTranspose1d, and StreamingAdd from moshi/moshi/modules/conv.py. These primitives, managed by StreamingContainer in moshi/moshi/modules/streaming.py, enable the encoder and decoder to process audio in blocks without buffering entire sequences, supporting real-time inference with configurable causal constraints.

Can normalization be disabled in specific layers of the SEANet encoder?

Yes, the disable_norm_outer_blocks parameter allows skipping normalization in the outermost blocks of both encoder and decoder. When set to a positive integer (e.g., 1 or 2), normalization is disabled for that many blocks counting from the input/output edges toward the center, useful for stabilizing training or matching specific deployment constraints where normalization may cause artifacts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →