# How to Use Transfer Learning and Fine-Tuning with Limited Datasets: A Practical Guide

> Master transfer learning and fine-tuning with limited datasets. Freeze early layers and retrain later ones for efficient model adaptation and improved performance.

- Repository: [scutan90/DeepLearning-500-questions](https://github.com/scutan90/DeepLearning-500-questions)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Transfer learning with limited datasets works by freezing the early convolutional layers that capture general features (edges, textures) while fine-tuning the later layers and classification head on your specific data, effectively leveraging pre-trained knowledge without overfitting.**

Transfer learning and fine-tuning with limited datasets is a critical technique in modern deep learning, allowing practitioners to build accurate models even when labeled data is scarce. This approach, extensively documented in the `scutan90/DeepLearning-500-questions` repository, leverages pre-trained networks to overcome the data hunger typical of deep neural networks. By strategically freezing and adapting specific layers, you can repurpose models trained on millions of images for highly specialized tasks with only hundreds of samples.

## Why Transfer Learning Works for Small Datasets

### General vs. Task-Specific Features

Early convolutional layers learn **domain-independent** representations such as edges, colors, and textures that generalize across almost any visual task. Later layers become increasingly **task-specific**, capturing complex patterns like object parts or semantic concepts. According to 【§11.3.8】 in the repository (lines 57‑59), freezing the early layers preserves this general knowledge while allowing the unfrozen later layers to adapt to your small target dataset.

### Avoiding Negative Transfer

When the source domain (e.g., ImageNet) and target domain are too dissimilar, transferring all layers can actually hurt performance—a phenomenon called **negative transfer**. Selectively freezing layers mitigates this risk by preventing overfitting to irrelevant source features. The repository discusses this extensively in 【§11.1.8】 (lines 85‑90), recommending careful layer selection based on domain similarity.

## The Standard Fine-Tuning Workflow

Follow these five steps to implement transfer learning effectively with limited data:

1. **Select a pre-trained backbone** such as ResNet‑50, VGG‑16, or MobileNetV2 trained on a large source dataset.
2. **Freeze the bottom *k* layers** (typically 70–80 % of the network) to preserve general feature extractors.
3. **Replace the classification head** with a new fully-connected layer matching your number of target classes.
4. **Fine-tune the unfrozen layers** using differential learning rates—lower for the backbone, higher for the new head.
5. **(Optional) Add a domain-adaptation block** such as an MMD loss layer or gradient-reversal layer to align source and target distributions.

## Practical Strategies for Limited Data

### Aggressive Data Augmentation

Expand your effective dataset size without collecting new samples by applying **random crops, horizontal flips, color jitter, and MixUp**. These transformations prevent overfitting by exposing the model to diverse variations of the same limited data.

### Learning Rate Discrimination

Use a **smaller learning rate for the frozen backbone** (e.g., `1e‑4`) and a **larger rate for the new classification head** (e.g., `1e‑3`). This preserves the pre-trained knowledge while allowing rapid adaptation of the task-specific parameters.

### Regularization Techniques

Apply **weight decay** (L2 regularization) and **early stopping** based on validation loss to prevent overfitting on small datasets. Monitor validation metrics closely and stop training when performance plateaus for more than five epochs.

## Implementation: PyTorch and TensorFlow Examples

### PyTorch: Freezing ResNet-50 Layers

This example demonstrates freezing early ResNet blocks while training a new binary classifier head:

```python
import torch, torchvision
from torchvision import transforms, datasets, models
from torch import nn, optim

# 1️⃣ Data – heavy augmentation for the small set

train_tf = transforms.Compose([
    transforms.RandomResizedCrop(224),
    transforms.RandomHorizontalFlip(),
    transforms.ColorJitter(brightness=0.2, contrast=0.2),
    transforms.ToTensor(),
])
train_ds = datasets.ImageFolder('data/train', transform=train_tf)
train_loader = torch.utils.data.DataLoader(train_ds, batch_size=16, shuffle=True)

# 2️⃣ Model – load pretrained, freeze early blocks

model = models.resnet50(pretrained=True)
for name, param in model.named_parameters():
    # Freeze everything up to layer3 (≈70% of layers)

    if "layer4" not in name:
        param.requires_grad = False

# 3️⃣ Replace classifier head

num_ftrs = model.fc.in_features
model.fc = nn.Linear(num_ftrs, 2)       # 2 target classes

# 4️⃣ Optimizer – separate LR for backbone vs. head

optimizer = optim.Adam([
    {"params": [p for n, p in model.named_parameters() if p.requires_grad and "fc" not in n],
     "lr": 1e-4},
    {"params": model.fc.parameters(), "lr": 1e-3}
])
criterion = nn.CrossEntropyLoss()

# 5️⃣ Fine‑tuning loop

model.train()
for epoch in range(30):
    for imgs, lbls in train_loader:
        optimizer.zero_grad()
        out = model(imgs)
        loss = criterion(out, lbls)
        loss.backward()
        optimizer.step()

```

### TensorFlow/Keras: Progressive Unfreezing

This Keras example uses MobileNetV2 with a two-stage training approach:

```python
import tensorflow as tf
from tensorflow.keras import layers, models, applications, optimizers

# 1️⃣ Data pipeline with augmentation

train_ds = tf.keras.preprocessing.image_dataset_from_directory(
    "data/train",
    image_size=(224, 224),
    batch_size=32,
    label_mode="categorical"
)
aug = tf.keras.Sequential([
    layers.RandomFlip("horizontal"),
    layers.RandomRotation(0.2),
    layers.RandomZoom(0.1),
])
train_ds = train_ds.map(lambda x, y: (aug(x), y))

# 2️⃣ Load pretrained backbone, freeze early layers

base = applications.MobileNetV2(input_shape=(224, 224, 3),
                               include_top=False,
                               weights="imagenet")
base.trainable = False                     # freeze whole backbone

# Optionally unfreeze last block for deeper adaptation

base.get_layer("block_13_expand").trainable = True

# 3️⃣ Build classifier head

inputs = layers.Input(shape=(224, 224, 3))
x = base(inputs, training=False)
x = layers.GlobalAveragePooling2D()(x)
x = layers.Dropout(0.2)(x)
outputs = layers.Dense(3, activation="softmax")(x)   # 3 target classes

model = models.Model(inputs, outputs)

# 4️⃣ Compile with two‑stage LR schedule

model.compile(optimizer=optimizers.Adam(learning_rate=1e-3),
              loss="categorical_crossentropy",
              metrics=["accuracy"])

# 5️⃣ Fine‑tune (first train head only, then unfreeze some backbone layers)

model.fit(train_ds, epochs=5)  # head only

base.trainable = True           # unfreeze selected layers

model.compile(optimizer=optimizers.Adam(learning_rate=1e-4),
              loss="categorical_crossentropy",
              metrics=["accuracy"])
model.fit(train_ds, epochs=20)

```

## Key Source Files in DeepLearning-500-questions

The techniques described above are grounded in the following documentation from the `scutan90/DeepLearning-500-questions` repository:

- **`ch11_迁移学习/第十一章_迁移学习.md`** — Contains the full theoretical overview of transfer learning, fine-tuning, domain adaptation, and negative transfer (sections 11.1‑11.3), providing the conceptual foundation for freezing strategies.

- **`ch05_卷积神经网络(CNN)/第五章_卷积神经网络(CNN).md`** — Explains CNN layer hierarchies and feature extraction principles, helping identify which early layers capture generalizable low-level features safe to freeze.

- **`ch03_深度学习基础/第三章_深度学习基础.md`** — Covers gradient descent, overfitting, and regularization fundamentals that underpin early stopping and weight decay strategies for small datasets.

## Summary

- **Freeze early layers** to preserve general feature extractors (edges, textures) while adapting later layers to your specific domain.
- **Use differential learning rates**—lower for the backbone (1e‑4) and higher for the new head (1e‑3)—to prevent catastrophic forgetting.
- **Apply aggressive data augmentation** (random crops, color jitter, MixUp) to artificially expand limited datasets and prevent overfitting.
- **Monitor for negative transfer** when source and target domains differ significantly; selectively unfreeze layers or add domain adaptation modules if performance degrades.
- **Implement early stopping and weight decay** to regularize small-dataset training and avoid overfitting to the limited samples.

## Frequently Asked Questions

### How many layers should I freeze when fine-tuning with limited data?

Typically freeze 70–80 % of the network, keeping early convolutional blocks frozen while unfreezing the final one or two blocks plus the classification head. According to 【§11.3.8】 in the DeepLearning-500-questions repository, this preserves domain-independent features (edges, colors) while allowing task-specific adaptation. If your target domain closely matches the source (e.g., natural images), freeze more layers; for dissimilar domains, freeze fewer.

### What is negative transfer and how do I avoid it?

Negative transfer occurs when knowledge from the source domain degrades performance on the target domain because the domains are too dissimilar. As noted in 【§11.1.8】 (lines 85‑90), you can mitigate this by selectively freezing layers rather than transferring the entire network, or by adding domain adaptation modules (e.g., gradient-reversal layers) to align feature distributions before fine-tuning.

### Should I use different learning rates for different layers?

Yes. Use a **discriminative learning rate** strategy: apply a smaller learning rate (e.g., `1e‑4`) to the pre-trained backbone to preserve general features, and a larger rate (e.g., `1e‑3`) to the randomly initialized classification head so it learns quickly. This prevents catastrophic forgetting while accelerating convergence on the new task, as demonstrated in both the PyTorch and TensorFlow code examples above.

### When should I consider domain adaptation techniques?

Consider domain adaptation when you observe poor validation performance despite proper fine-tuning, indicating a significant distribution shift between source and target data. According to 【§11.3.10】 (lines 41‑45), techniques like adding an MMD loss layer or adversarial gradient-reversal modules can align the feature distributions before or during fine-tuning, effectively bridging the gap between dissimilar domains when simple freezing strategies prove insufficient.