Autoencoder/VAE Architecture for Image Generation in Microsoft’s AI-For-Beginners
The microsoft/AI-For-Beginners repository implements a fully-connected Variational Autoencoder featuring a dual-output encoder that predicts Gaussian parameters μ and log-variance, a reparameterization trick for latent sampling, and a symmetric decoder with sigmoid activation trained with binary cross-entropy and KL-divergence losses to generate 28×28 MNIST digits.
Lesson 09 of the computer vision curriculum (lessons/4-ComputerVision/09-Autoencoders/) provides complete implementations in both TensorFlow/Keras and PyTorch. These notebooks demonstrate how a dense-layer autoencoder/VAE architecture learns compressed latent representations and synthesizes new handwritten digit images through probabilistic decoding.
Encoder Architecture
The encoder compresses the 784-dimensional flattened MNIST input (28×28 pixels) through a series of fully-connected layers. According to the source code in AutoencodersTF.ipynb and AutoEncodersPyTorch.ipynb, the architecture follows this specification:
- Input layer: Accepts 784 grayscale pixel values (flattened from 28×28)
- Hidden layer 1: 512 units with ReLU activation
- Hidden layer 2: 512 units with ReLU activation
- Output split: Two parallel dense layers producing
mu(mean μ) andlog_var(log-variance) vectors
The final layers do not output a single vector but rather parameterize the latent Gaussian distribution. The latent dimension is typically configured as 2 for visualization purposes or 10 for more expressive generation capacity.
# PyTorch encoder implementation (from AutoEncodersPyTorch.ipynb)
def encode(self, x):
x = torch.relu(self.fc1(x)) # fc1: Linear(784, 512)
x = torch.relu(self.fc2(x)) # fc2: Linear(512, 512)
return self.fc_mu(x), self.fc_logvar(x) # Linear(512, latent_dim)
Latent Space and the Reparameterization Trick
To enable backpropagation through the stochastic sampling operation, the implementation uses the reparameterization trick. Instead of sampling directly from the distribution defined by μ and σ, the model samples epsilon (ε) from a standard normal distribution and transforms it deterministically.
The operation follows: z = μ + σ·ε, where σ = exp(0.5 * log_var) and ε ~ N(0, I).
In the PyTorch implementation (AutoEncodersPyTorch.ipynb), this logic resides in a dedicated method:
def reparameterize(self, mu, logvar):
std = torch.exp(0.5 * logvar)
eps = torch.randn_like(std)
return mu + eps * std
This separates the stochastic component (epsilon) from the learnable parameters (μ and log-variance), allowing gradient flow during training.
Decoder Architecture
The decoder mirrors the encoder structure to reconstruct the original image dimensions from the sampled latent vector z. The architecture specifications found in both framework implementations include:
- Input: Sampled latent vector
z(dimension 2 or 10) - Hidden layer 1: 512 units with ReLU activation
- Hidden layer 2: 512 units with ReLU activation
- Output layer: 784 units with sigmoid activation, reshaped to 28×28
The sigmoid activation ensures pixel values fall within the [0, 1] range, matching the normalized MNIST input.
# TensorFlow/Keras decoder (from AutoencodersTF.ipynb)
latent_inputs = keras.Input(shape=(latent_dim,))
x = keras.layers.Dense(512, activation='relu')(latent_inputs)
x = keras.layers.Dense(512, activation='relu')(x)
outputs = keras.layers.Dense(784, activation='sigmoid')(x)
decoder = keras.Model(latent_inputs, outputs, name="decoder")
Loss Function and Training Configuration
The training objective combines reconstruction accuracy with distribution regularization. As implemented in the lesson notebooks, the total loss is the sum of two components:
- Reconstruction loss: Binary cross-entropy between the input image and the decoder output (measuring pixel-wise reconstruction error)
- KL Divergence:
0.5 * sum(exp(log_var) + mu^2 - 1 - log_var), which forces the learned latent distribution toward a standard normal distribution
Training hyperparameters used in the source code:
- Optimizer: Adam with learning rate 1e-3
- Batch size: 128
- Epochs: 30–50 (sufficient for MNIST convergence)
- Dataset: MNIST handwritten digits normalized to [0, 1]
Framework-Specific Implementation Differences
While the architecture remains identical, the organizational structure differs between frameworks.
TensorFlow/Keras (AutoencodersTF.ipynb): Uses the Functional API to define clear encoder and decoder sub-models, connecting them via the reparameterization layer in the training step.
PyTorch (AutoEncodersPyTorch.ipynb): Encapsulates the logic in a VAE class inheriting from nn.Module with explicit methods:
encode(x)→ returns mu, logvarreparameterize(mu, logvar)→ returns zdecode(z)→ returns reconstructed imageforward(x)→ orchestrates the full inference pass
Visual diagrams of this architecture are provided in the repository at lessons/4-ComputerVision/09-Autoencoders/images/autoencoder_schema.jpg and lessons/4-ComputerVision/09-Autoencoders/images/vae.png.
Summary
- The VAE uses a symmetric dense architecture (784→512→512→latent→512→512→784) with ReLU activations in hidden layers and sigmoid on the output.
- The encoder outputs μ and log-variance parameters rather than a single deterministic vector.
- The reparameterization trick enables gradient descent through the stochastic latent sampling operation.
- Training optimizes a composite binary cross-entropy + KL-divergence loss using the Adam optimizer.
- Complete implementations are available in
AutoencodersTF.ipynb(TensorFlow) andAutoEncodersPyTorch.ipynb(PyTorch) under the09-Autoencoderslesson directory.
Frequently Asked Questions
What dataset does the autoencoder/VAE implementation use?
The implementation uses the MNIST dataset of handwritten digits. Each image is 28×28 pixels, grayscale, and flattened to a 784-dimensional vector before being fed into the encoder.
Why is the latent dimension typically set to 2 in the examples?
A latent dimension of 2 allows for easy visualization of the learned distribution in a 2D plane, making it straightforward to plot the μ vectors and sample points for educational purposes. The architecture supports larger dimensions (e.g., 10) for higher-quality generation.
What is the purpose of the reparameterization trick in the VAE?
The reparameterization trick moves the random sampling operation outside of the main computational graph. By sampling epsilon (ε) from a standard normal distribution and transforming it via μ + σ·ε, the model enables backpropagation to flow through the mean and variance parameters while maintaining the stochastic nature of the latent space.
Where can I find the complete model definitions and training loops?
The complete TensorFlow implementation is located in lessons/4-ComputerVision/09-Autoencoders/AutoencodersTF.ipynb, while the PyTorch version resides in lessons/4-ComputerVision/09-Autoencoders/AutoEncodersPyTorch.ipynb. Both files contain the full VAE class definitions, loss computations, and training epochs necessary to reproduce the image generation results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →