StyleGAN 2 vs Original GAN: 8 Key Architectural Differences
StyleGAN 2 replaces the simple generator-discriminator pipeline of the original GAN with a mapping network, weight modulation and demodulation, residual skip connections, and path-length regularization to eliminate artifacts and enable high-fidelity image synthesis.
The evolution from the original Generative Adversarial Network to StyleGAN 2 represents a fundamental architectural overhaul. While the vanilla GAN implemented in labml_nn/gan/original/__init__.py uses a straightforward multilayer perceptron to map latent noise directly to image pixels, StyleGAN 2 introduces sophisticated mechanisms that decouple high-level attributes from stochastic variation. According to the labmlai/annotated_deep_learning_paper_implementations repository, these changes specifically target the "droplet artifacts" and training instabilities inherent in earlier designs.
Mapping Network and Disentangled Latent Space
The original GAN feeds the latent vector z directly into the generator stack. StyleGAN 2 introduces an 8-layer MLP MappingNetwork (defined at line 58 in labml_nn/gan/stylegan/__init__.py) that transforms z into an intermediate latent space w. This disentangles factors of variation and enables style mixing at different generator layers, allowing control over coarse, middle, and fine-grained image attributes independently.
Weight Modulation and Demodulation
Instead of Adaptive Instance Normalization (AdaIN) used in StyleGAN 1, StyleGAN 2 employs weight modulation and demodulation (implemented at line 173 in labml_nn/gan/stylegan/__init__.py). This mechanism directly scales convolution weights based on style vectors, then normalizes the output to maintain signal variance. By operating on weights rather than activations, this approach eliminates the droplet artifacts caused by AdaIN normalization statistics while preserving style control.
Skip Connections and Multi-Scale RGB Generation
The GeneratorBlock (line 497 in labml_nn/gan/stylegan/__init__.py) utilizes skip connections where each block contributes an RGB output that is summed with upsampled outputs from previous blocks. This contrasts with the vanilla GAN's single final output layer and improves gradient flow during training, allowing the model to learn multi-scale image generation effectively.
Residual Discriminator Architecture
Unlike the original GAN's simple feed-forward discriminator, StyleGAN 2 uses DiscriminatorBlock (line 563 in labml_nn/gan/stylegan/__init__.py) with residual connections and downsampling. This residual structure enables deeper architectures without training collapse, allowing the discriminator to better distinguish real from generated images at high resolutions.
Path-Length Regularization
StyleGAN 2 introduces PathLengthRegularizer (line 340 in labml_nn/gan/stylegan/__init__.py) that penalizes changes in image space relative to latent space steps. This regularization ensures stable, predictable latent space interpolation and prevents the generator from producing sudden, discontinuous changes when traversing the latent manifold.
Stochastic Noise Injection
Per-channel learned noise scaling (lines 99-102 in labml_nn/gan/stylegan/__init__.py) adds stochastic variation at each generator layer. This mechanism controls fine details like hair placement or skin pores separately from global structure, providing explicit control over image stochasticity that the original GAN lacks.
Removal of Progressive Growing
Unlike Progressive GAN, StyleGAN 2 trains at the target resolution from the start, eliminating the phased growing schedule that caused phase-specific artifacts. The architecture in labml_nn/gan/stylegan/experiment.py initializes the full-resolution network immediately, relying on skip connections and residual blocks rather than gradually growing layers.
Code Comparison: Vanilla GAN vs StyleGAN 2
The architectural differences become clear when comparing generator instantiation between the two implementations.
Original GAN Generator (simple MLP from labml_nn/gan/original/__init__.py):
from labml_nn.gan.original import Generator as VanillaGen
import torch
# Simple MLP-based generator
vanilla_G = VanillaGen(latent_dim=100, hidden_dim=256, image_dim=784)
z = torch.randn(1, 100)
img_flat = vanilla_G(z) # shape: (1, 784)
StyleGAN 2 Generator (mapping network + modulated convolutions from labml_nn/gan/stylegan/__init__.py):
from labml_nn.gan.stylegan import Generator
import torch
# 64×64 model with 8-layer mapping network
G = Generator(image_size=64, latent_size=512, mapping_layers=8)
# Sample latent vector z ~ N(0, 1)
z = torch.randn(1, 512)
# Forward pass through mapping network and synthesis network
img = G(z) # shape: (1, 3, 64, 64), RGB in [-1, 1]
Summary
- StyleGAN 2 introduces an 8-layer mapping network that transforms latent codes into a disentangled intermediate space w, enabling fine-grained style control.
- Weight modulation and demodulation replace AdaIN to eliminate droplet artifacts while maintaining style-based generation.
- Skip connections in the generator and residual blocks in the discriminator enable training at full resolution without progressive growing.
- Path-length regularization ensures smooth latent space interpolation by penalizing erratic image-space changes.
- Per-channel noise injection provides explicit control over stochastic variation for fine details like hair and textures.
Frequently Asked Questions
What is the main architectural difference between StyleGAN 2 and the original GAN?
The original GAN uses a simple feed-forward generator that maps latent noise directly to image pixels, while StyleGAN 2 employs a two-stage synthesis process: first, an 8-layer mapping network transforms the latent code z into an intermediate space w, then a synthesis network with weight modulation and skip connections generates the image. This decouples high-level attributes from stochastic variation and eliminates artifacts common in vanilla GANs.
Why did StyleGAN 2 remove AdaIN and replace it with weight modulation?
StyleGAN 2 removed Adaptive Instance Normalization (AdaIN) because it introduced droplet artifacts—blob-like patterns caused by the normalization statistics. The replacement, weight modulation and demodulation (implemented in labml_nn/gan/stylegan/__init__.py at line 173), directly scales convolution weights based on style vectors and normalizes the output feature maps. This preserves style control while maintaining signal variance and eliminating visual artifacts.
How does StyleGAN 2 handle training stability without progressive growing?
Instead of progressively growing network layers from low to high resolution, StyleGAN 2 trains the full architecture at the target resolution from the start, relying on residual connections in the discriminator (DiscriminatorBlock at line 563) and skip connections in the generator (GeneratorBlock at line 497) to maintain gradient flow. This approach, combined with path-length regularization (line 340), prevents the training instabilities and phase-specific artifacts that progressive growing often caused.
Can I use the StyleGAN 2 implementation in this repository for high-resolution image generation?
Yes, the implementation in labml_nn/gan/stylegan/ supports configurable image sizes via the image_size parameter, though the repository defaults to 64×64 for demonstration purposes. To train at higher resolutions (e.g., 1024×1024), you would adjust image_size, increase the channel capacities in Generator and Discriminator, and ensure adequate computational resources, as the architecture in experiment.py already handles the full-resolution training pipeline without progressive growing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →