# How Encoder-Decoder Architecture Differs in U-Net, FCN, and DeepLab

> Understand the encoder-decoder architecture differences in U-Net, FCN, and DeepLab. Explore skip connections, transposed convolutions, and atrous convolutions to improve image segmentation models.

- Repository: [scutan90/DeepLearning-500-questions](https://github.com/scutan90/DeepLearning-500-questions)
- Tags: deep-dive
- Published: 2026-03-06

---

**U-Net uses symmetric skip connections to concatenate encoder and decoder features, FCN relies on a single transposed convolution for upsampling without skip connections, and DeepLab employs atrous convolutions with ASPP and selective low-level feature fusion.**

The encoder-decoder architecture serves as the foundational pattern for semantic segmentation, yet these three landmark implementations diverge significantly in how they preserve spatial resolution and combine hierarchical features. According to the scutan90/DeepLearning-500-questions repository—specifically documented in `ch09_图像分割/第九章_图像分割.md`—each architecture represents a distinct trade-off between computational efficiency, implementation complexity, and boundary localization accuracy.

## Encoder Design: Backbone Strategies

### FCN: Classification Network Conversion

FCN repurposes standard classification CNNs (e.g., VGG, ResNet) by replacing fully-connected layers with 1×1 convolutions to enable dense prediction. The encoder maintains traditional max-pooling operations, progressively reducing spatial resolution to 1/32 of the input size while extracting deep semantic features.

### U-Net: Contracting Path with Repeated Convolutions

U-Net implements a dedicated contracting path consisting of repeated Conv-ReLU-Conv blocks followed by 2×2 max-pooling operations. As detailed in the repository, this symmetric encoder "mirrors" the decoder structure, with each downsampling step doubling the feature channel count to preserve information capacity.

### DeepLab: Atrous Convolutions for Resolution Preservation

DeepLab modifies standard backbones by removing the final pooling layers and implementing **atrous (dilated) convolutions** to enlarge the receptive field without downsampling. According to `第九章_图像分割.md` lines 496-513, this approach keeps a finer feature grid throughout the encoder, avoiding the aggressive resolution loss inherent in traditional pooling-based architectures.

## Decoder Architectures: Upsampling and Feature Fusion

### FCN: Single-Stage Transposed Convolution

FCN employs the simplest decoder design, utilizing a single transposed convolution (deconvolution) layer to upsample the final feature map directly to the original image resolution. As documented in lines 45-48 of `第九章_图像分割.md`, "FCN对图像进行像素级的分类 … 采用反卷积层对最后一个卷积层的feature map进行上采样". This design lacks skip connections, forcing the network to rely entirely on deep, low-resolution features for pixel-level predictions.

### U-Net: Symmetric Expansion with Skip Concatenation

U-Net implements a symmetric decoder that reverses the encoder's contractions through upsampling blocks paired with **skip connections**. The repository's architecture description (lines 52-55) specifies that "左边的网络是收缩路径 … 右边的网络是扩张路径 … 与左侧收缩路径对应层产生的特征图进行concatenate操作". Each decoder level concatenates high-resolution encoder features with upsampled decoder outputs, followed by two 1×1 convolutions to fuse multi-scale information before the next upsampling step.

### DeepLab: ASPP and Low-Level Feature Fusion

DeepLab v3+ introduces a sophisticated two-stage decoder. First, **Atrous Spatial Pyramid Pooling (ASPP)** processes encoder outputs at multiple dilation rates (6, 12, 18) to capture multi-scale context. The decoder then concatenates low-level features from early encoder layers (specifically the `conv2_block3_out` layer at 1/4 resolution) with the ASPP output. As noted in lines 511-515 of the source file, this v3+ modification "将修改之前提出的带孔空间金字塔池化模块" and refines boundaries through 1×1 convolutions and bilinear upsampling rather than learned transposed convolutions.

## Code Implementation Examples

### U-Net Implementation with Skip Connections

The following implementation from `第九章_图像分割.md` (lines 60-100) demonstrates the symmetric skip-connection pattern:

```python
def get_unet():
    inputs = Input((img_rows, img_cols, 1))

    # Encoder

    conv1 = Conv2D(32, (3, 3), activation='relu', padding='same')(inputs)
    conv1 = Conv2D(32, (3, 3), activation='relu', padding='same')(conv1)
    pool1 = MaxPooling2D(pool_size=(2, 2))(conv1)

    conv2 = Conv2D(64, (3, 3), activation='relu', padding='same')(pool1)
    conv2 = Conv2D(64, (3, 3), activation='relu', padding='same')(conv2)
    pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)

    # … (continue down to conv5)

    # Decoder (upsampling + concat)

    up6 = concatenate([
        Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same')(conv5),
        conv4], axis=3)
    conv6 = Conv2D(256, (3, 3), activation='relu', padding='same')(up6)
    conv6 = Conv2D(256, (3, 3), activation='relu', padding='same')(conv6)

    # … (continue up‑sampling to final 1×1 conv)

    output = Conv2D(num_classes, (1, 1), activation='softmax')(conv_final)
    return Model(inputs=inputs, outputs=output)

```

### FCN Minimal Decoder Pattern

This TensorFlow implementation reflects FCN's single-stage upsampling approach without skip connections:

```python
def fcn_8s(num_classes, input_shape=(None, None, 3)):
    base = tf.keras.applications.VGG16(weights='imagenet',
                                       include_top=False,
                                       input_shape=input_shape)

    x = base.output                               # shape: 1/32 of input

    x = Conv2D(num_classes, (1, 1), activation=None)(x)

    # Upsample 8× with transposed convolution

    x = Conv2DTranspose(num_classes, kernel_size=16,
                        strides=8, padding='same')(x)
    x = Activation('softmax')(x)
    return Model(inputs=base.input, outputs=x)

```

### DeepLab v3+ with ASPP and Feature Fusion

This implementation pattern matches the DeepLab v3+ architecture described in the repository:

```python
import tensorflow as tf
from tensorflow.keras.applications import ResNet50
from tensorflow.keras.layers import Conv2D, UpSampling2D, Concatenate

def deeplabv3_plus(num_classes, input_shape=(None, None, 3)):
    # Encoder – ResNet‑50 with atrous rates 6,12,18

    backbone = ResNet50(weights='imagenet', include_top=False,
                        input_shape=input_shape)
    low_level_feat = backbone.get_layer('conv2_block3_out').output   # 1/4 resolution

    # Atrous Spatial Pyramid Pooling (ASPP)

    aspp = ASPP(backbone.output)   # custom layer that builds 1x1, 3x3 (rate 6/12/18), image‑pool

    # Decoder

    x = UpSampling2D(size=(4, 4), interpolation='bilinear')(aspp)
    low = Conv2D(48, (1, 1), padding='same')(low_level_feat)
    x = Concatenate()([x, low])
    x = Conv2D(256, (3, 3), padding='same', activation='relu')(x)
    x = Conv2D(256, (3, 3), padding='same', activation='relu')(x)

    # Final up‑sample to original size

    x = UpSampling2D(size=(4, 4), interpolation='bilinear')(x)
    output = Conv2D(num_classes, (1, 1), activation='softmax')(x)
    return tf.keras.Model(inputs=backbone.input, outputs=output)

```

## Summary

- **U-Net** achieves precise localization through **symmetric skip connections** that concatenate encoder and decoder features at multiple scales, using `Conv2DTranspose` followed by `concatenate` operations.
- **FCN** provides a **minimal decoder** using only a single transposed convolution layer, trading spatial precision for implementation simplicity and direct upsampling from deep features.
- **DeepLab** preserves resolution via **atrous convolutions** in the encoder and aggregates multi-scale context through **ASPP**, using selective low-level feature fusion only in the final decoder stage rather than full symmetric skip connections.
- All three architectures derive their encoders from classification networks, but differ fundamentally in how they recover spatial resolution: FCN through aggressive learned upsampling, U-Net through hierarchical feature concatenation, and DeepLab through dilated encoding plus shallow refinement.

## Frequently Asked Questions

### What is the main disadvantage of FCN's decoder compared to U-Net?

FCN's reliance on a single transposed convolution without skip connections causes significant loss of fine-grained spatial details during encoding. This produces blurrier boundary predictions compared to U-Net's feature concatenation approach, which directly injects high-resolution encoder information into the decoder at multiple levels.

### Why does DeepLab use atrous convolutions instead of standard pooling in the encoder?

Atrous convolutions maintain higher spatial resolution throughout the encoder by removing final downsampling operations while still expanding the receptive field through dilation. This eliminates the need for aggressive upsampling in the decoder and preserves fine details for better segmentation accuracy, particularly at object boundaries.

### Can U-Net's skip connections be integrated into FCN or DeepLab architectures?

Yes, skip connections can be added to FCN-style architectures (creating variants like FCN-8s with skip pathways), and DeepLab v3+ already incorporates low-level skip connections from early encoder layers such as `conv2_block3_out`. However, U-Net requires precise dimensional matching between symmetric encoder and decoder stages, which constrains architectural flexibility compared to DeepLab's asymmetric design.

### Which architecture is best for medical image segmentation?

U-Net is generally preferred for medical imaging due to its symmetric skip connections that preserve anatomical boundaries and precise localization. The architecture was specifically designed for biomedical applications where pixel-accurate delineation of small structures is critical, making it superior to FCN's blunt upsampling or DeepLab's focus on natural scene multi-scale context.