SEANet Encoder/Decoder Configuration in NVIDIA PersonaPlex: Architecture and Implementation
The SEANet encoder/decoder configuration in PersonaPlex defaults to a 320× compression ratio with 128-dimensional latent representations, utilizing configurable residual blocks, ELU activation, and streaming-compatible convolutional layers for real-time audio processing.
The SEANet architecture serves as the backbone of NVIDIA's PersonaPlex audio codec, providing efficient neural audio coding through highly configurable encoder-decoder pairs. Implemented in moshi/moshi/modules/seanet.py, both the SEANetEncoder and SEANetDecoder classes expose extensive hyperparameters controlling network depth, dilation strategies, and streaming behavior. This architecture enables high-quality audio compression with support for causal inference and block-wise processing without full-signal buffering.
SEANet Architecture Overview
The SEANet implementation consists of three core components: the encoder, decoder, and residual blocks. The architecture leverages custom streaming primitives from moshi/moshi/modules/conv.py to enable inference on arbitrary-length sequences.
Key architectural elements include:
- Streaming-ready primitives: Uses
StreamingConv1d,StreamingConvTranspose1d, andStreamingAddfor block-wise processing - Residual blocks:
SEANetResnetBlock(lines 42-73 inseanet.py) implements skip connections with configurable dilation and compression - Causal support: All convolutions support strict causality via the
causalparameter - Flexible normalization: Optional normalization via the
normparameter (default"none") with outer block disabling viadisable_norm_outer_blocks
SEANetEncoder Configuration Parameters
The SEANetEncoder compresses raw audio into a latent representation using downsampling stages interleaved with residual blocks.
Default configuration values:
channels:1(mono input)dimension:128(latent vector size)n_filters:32(base filter count)n_residual_layers:3(residual blocks per stage)ratios:[8, 5, 4, 2](downsampling factors)activation:"ELU"withactivation_params:{"alpha": 1.0}kernel_size/last_kernel_size:7residual_kernel_size:3dilation_base:2(for exponential dilation growth)causal:Falsepad_mode:"reflect"true_skip:True(identity skip connections)compress:2(channel compression factor)mask_fn/mask_position:None(optional masking)
Implementation specifics:
The encoder reverses the ratios list during initialization: self.ratios = list(reversed(ratios)) (lines 76-78). The total hop length calculates to 320 samples via self.hop_length = int(np.prod(self.ratios)), meaning a 16 kHz input produces latent frames at 50 Hz.
SEANetDecoder Configuration Parameters
The SEANetDecoder reconstructs audio from latent representations, mirroring the encoder architecture using transposed convolutions for upsampling.
Default configuration values:
channels:1(mono output)dimension:128(input latent size)n_filters:32n_residual_layers:3ratios:[8, 5, 4, 2](upsampling factors, applied in forward order)activation:"ELU"final_activation:None(optional output non-linearity)final_activation_params:Nonekernel_size/last_kernel_size:7residual_kernel_size:3dilation_base:2causal:Falsepad_mode:"reflect"true_skip:Truecompress:2disable_norm_outer_blocks:0trim_right_ratio:1.0(controls right-side trimming for causal transposed convolutions)
The decoder begins with a projection from dimension to 2**len(ratios) * n_filters, then iteratively upsamples through each ratio stage, halving the filter multiplier after each StreamingConvTranspose1d operation.
Residual Block and Streaming Implementation
Each SEANetResnetBlock contains:
- Two convolutional layers with kernel sizes
residual_kernel_size(default 3) and1 - Exponential dilation:
dilation_base ** layer_index - Optional channel compression by the
compressfactor (default 2) - Identity skip connections when
true_skip=True, otherwise 1×1 convolutions
Streaming capabilities derive from moshi/moshi/modules/streaming.py, allowing models to process arbitrary-length sequences via StreamingContainer state management. When causal=True, all convolutions use appropriate padding to ensure no future dependencies, critical for real-time applications.
Practical Configuration Examples
Instantiating Default Encoder and Decoder
import torch
from moshi.moshi.modules.seanet import SEANetEncoder, SEANetDecoder
# Default configuration: 320× compression, 128-dim latent
encoder = SEANetEncoder()
decoder = SEANetDecoder()
# Process 16 kHz mono audio (batch=1, channels=1, samples=16000)
waveform = torch.randn(1, 1, 16000)
# Encode to latent space
z = encoder(waveform) # Shape: (1, 128, 50)
print(f"Latent shape: {z.shape}") # 16000 / 320 = 50 frames
# Decode back to audio
recon = decoder(z) # Shape: (1, 1, 16000)
print(f"Reconstructed shape: {recon.shape}")
Custom Configuration for Streaming Applications
# Causal configuration for real-time processing
encoder = SEANetEncoder(
channels=2, # Stereo input
n_filters=64, # Wider filters
ratios=[4, 4, 4, 4], # 256× total downsampling
causal=True, # Strict causality
pad_mode="constant", # Alternative padding
)
decoder = SEANetDecoder(
channels=2,
n_filters=64,
ratios=[4, 4, 4, 4],
causal=True,
final_activation="Tanh", # Bound output to [-1, 1]
trim_right_ratio=1.0,
)
Accessing Hop Length for Alignment
# Calculate temporal downsampling factor
hop_length = encoder.hop_length # Returns 320 for default ratios
print(f"Hop length: {hop_length} samples")
# Useful for aligning latent features with spectrograms
Adding Spectral Masking
class EnergyBasedMask(torch.nn.Module):
def forward(self, x):
# Mask low-energy frames
mask = (x.abs().mean(dim=1, keepdim=True) > 0.01).float()
return x * mask
# Inject mask after specific layer
encoder = SEANetEncoder(
mask_fn=EnergyBasedMask(),
mask_position=2 # Apply after second downsampling block
)
Summary
- Location: The SEANet encoder/decoder configuration resides in
moshi/moshi/modules/seanet.pywithin the PersonaPlex repository - Compression: Default configuration achieves 320× compression (ratios 8×5×4×2) with 128-dimensional latent vectors
- Flexibility: Both encoder and decoder support customizable channels, filter depths, dilation strategies, and normalization schemes
- Streaming: Native support for block-wise inference via
StreamingConv1dandStreamingConvTranspose1denables real-time processing - Causality: Optional causal mode ensures no future dependencies for streaming applications
Frequently Asked Questions
What is the default compression ratio of the SEANet encoder in PersonaPlex?
The default SEANet encoder achieves a 320× compression ratio calculated from the product of the default ratios [8, 5, 4, 2]. With a 16 kHz input sampling rate, this produces latent features at 50 Hz with a hop length of 320 samples, implemented via self.hop_length = int(np.prod(self.ratios)) in the encoder initialization.
How does the SEANet encoder handle the ratios parameter differently from the decoder?
The encoder reverses the ratios list during initialization (self.ratios = list(reversed(ratios)) at lines 76-78) to ensure the downsampling order matches the decoder's upsampling order. While the decoder applies ratios [8, 5, 4, 2] sequentially for upsampling, the encoder applies them in reverse order [2, 4, 5, 8] for downsampling, maintaining architectural symmetry between compression and reconstruction paths.
What streaming capabilities are built into the SEANet architecture?
SEANet utilizes streaming-compatible convolution layers including StreamingConv1d, StreamingConvTranspose1d, and StreamingAdd from moshi/moshi/modules/conv.py. These primitives, managed by StreamingContainer in moshi/moshi/modules/streaming.py, enable the encoder and decoder to process audio in blocks without buffering entire sequences, supporting real-time inference with configurable causal constraints.
Can normalization be disabled in specific layers of the SEANet encoder?
Yes, the disable_norm_outer_blocks parameter allows skipping normalization in the outermost blocks of both encoder and decoder. When set to a positive integer (e.g., 1 or 2), normalization is disabled for that many blocks counting from the input/output edges toward the center, useful for stabilizing training or matching specific deployment constraints where normalization may cause artifacts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →