# Differences Between PyTorch and TensorFlow Versions of Dual-Framework Notebooks

> Explore PyTorch vs TensorFlow notebooks. Learn how PyTorch needs manual loops and device management, while TensorFlow uses Keras for automated training and optimization.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-27

---

**The PyTorch notebooks require explicit manual training loops and device management, while the TensorFlow notebooks leverage high-level Keras APIs that automate gradient computation and model optimization.**

The microsoft/AI-For-Beginners repository provides dual-framework notebooks that teach identical machine learning concepts using both PyTorch and TensorFlow 2.x implementations. While both versions achieve the same learning objectives—loading datasets, defining convolutional architectures, and evaluating model performance—the code structure differs significantly due to each framework's underlying design philosophy. Understanding these differences helps learners choose the approach that best fits their workflow, whether they prefer TensorFlow's abstraction or PyTorch's explicit control.

## API Architecture and Model Definition

The primary divergence between the PyTorch and TensorFlow versions begins with how models are constructed and imported.

### Import Conventions and Module Structure

In the TensorFlow notebooks—such as `lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb`—the code relies on Keras layers bundled within TensorFlow:

```python
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Dense

```

Conversely, the PyTorch notebooks found in `lessons/4-ComputerVision/07-ConvNets/ConvNetsPyTorch.ipynb` require separate imports for network layers, optimization utilities, and data handling:

```python
import torch
import torch.nn as nn
import torch.optim as optim
import torchvision

```

### Model Definition: Subclassing vs Sequential API

**TensorFlow** utilizes the **Keras Sequential API** for rapid prototyping. Layers stack declaratively, and the framework handles forward-pass logic internally:

```python
model = tf.keras.Sequential([
    tf.keras.layers.Conv2D(32, 3, activation='relu', input_shape=(28,28,1)),
    tf.keras.layers.MaxPooling2D(),
    tf.keras.layers.Flatten(),
    tf.keras.layers.Dense(128, activation='relu'),
    tf.keras.layers.Dense(10)
])

```

**PyTorch** requires explicit subclassing of `nn.Module`. You must define a `forward(self, x)` method that specifies exactly how data flows through the network:

```python
class ConvNet(nn.Module):
    def __init__(self):
        super(ConvNet, self).__init__()
        self.conv1 = nn.Conv2d(1, 32, 3)
        self.pool = nn.MaxPool2d(2, 2)
        self.fc1 = nn.Linear(64*5*5, 128)
        self.fc2 = nn.Linear(128, 10)

    def forward(self, x):
        x = self.pool(torch.relu(self.conv1(x)))
        x = x.view(-1, 64*5*5)
        x = torch.relu(self.fc1(x))
        return self.fc2(x)

```

## Data Loading and Pipeline Differences

The dual-framework notebooks demonstrate distinct approaches to data ingestion that reflect each library's ecosystem.

**TensorFlow** versions typically use `tf.keras.datasets` helpers that return NumPy arrays immediately ready for training:

```python
(train_x, train_y), (test_x, test_y) = keras.datasets.mnist.load_data()
train_x = train_x[..., np.newaxis] / 255.0

```

**PyTorch** implementations rely on **torchvision.datasets** combined with `torch.utils.data.DataLoader` for batched iteration. This requires defining transformation pipelines and explicitly shuffling data:

```python
transform = transforms.Compose([transforms.ToTensor()])
train_set = torchvision.datasets.MNIST(root='./data', train=True, download=True, transform=transform)
train_loader = torch.utils.data.DataLoader(train_set, batch_size=64, shuffle=True)

```

The PyTorch approach provides fine-grained control over batching and augmentation, while the TensorFlow method prioritizes immediate usability.

## Training Loop Implementation

The most significant practical difference lies in how training executes.

### TensorFlow High-Level Abstraction

TensorFlow notebooks compile the model with loss functions and optimizers, then invoke `model.fit()` to handle epochs, batching, and gradient updates automatically:

```python
model.compile(
    loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
    optimizer=tf.keras.optimizers.Adam(),
    metrics=['accuracy']
)

model.fit(train_x, train_y, epochs=5, batch_size=64, validation_split=0.1)
test_loss, test_acc = model.evaluate(test_x, test_y)

```

### PyTorch Explicit Training Loop

PyTorch notebooks require manual iteration over the `DataLoader`, explicit loss computation, and manual backpropagation steps. The `ConvNetsPyTorch.ipynb` file demonstrates this low-level pattern:

```python
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters())

for epoch in range(5):
    running_loss = 0.0
    for images, labels in train_loader:
        images, labels = images.to(device), labels.to(device)
        
        optimizer.zero_grad()
        outputs = model(images)
        loss = criterion(outputs, labels)
        loss.backward()
        optimizer.step()
        running_loss += loss.item()

```

Evaluation also requires manual computation:

```python
correct = 0
total = 0
with torch.no_grad():
    for images, labels in test_loader:
        outputs = model(images.to(device))
        _, predicted = torch.max(outputs, 1)
        total += labels.size(0)
        correct += (predicted == labels.cpu()).sum().item()

```

## Device Management and Execution Context

Both frameworks support GPU acceleration, but handle device placement differently.

**TensorFlow** abstracts device management unless explicitly scoped. Tensors typically reside on the default device, and you can force GPU usage via `tf.device('/GPU:0')` context managers.

**PyTorch** requires explicit device handling. The notebooks include standard boilerplate to detect CUDA availability and move both models and tensors manually:

```python
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = ConvNet().to(device)

```

Every batch must explicitly transfer to the target device: `images, labels = images.to(device), labels.to(device)`.

## Model Persistence and Serialization

Saving trained models differs significantly between the two frameworks.

**TensorFlow** uses the Keras H5 format or SavedModel format:

```python
model.save('my_model.h5')
loaded_model = tf.keras.models.load_model('my_model.h5')

```

**PyTorch** saves only the model's `state_dict` (learned parameters), requiring you to reinstantiate the architecture class before loading weights:

```python
torch.save(model.state_dict(), 'model.pt')
model = ConvNet()
model.load_state_dict(torch.load('model.pt'))

```

## Summary

- **TensorFlow notebooks** utilize high-level Keras APIs with `model.fit()` and `model.compile()`, minimizing boilerplate code for standard training loops.
- **PyTorch notebooks** require explicit manual training loops with `loss.backward()` and `optimizer.step()`, offering granular control over gradient computation.
- **Data loading** in TensorFlow relies on `tf.keras.datasets` returning NumPy arrays, while PyTorch uses `torchvision.datasets` with `DataLoader` for batch management.
- **Device management** is automatic in TensorFlow but requires explicit `.to(device)` calls for models and tensors in PyTorch.
- **Model saving** in TensorFlow preserves the complete architecture and weights in H5 format, whereas PyTorch saves only parameters via `state_dict` dictionaries.

## Frequently Asked Questions

### Which framework version should beginners start with?

Beginners often find the TensorFlow/Keras versions more approachable because the `model.fit()` API abstracts away the training loop complexity. However, the PyTorch versions provide clearer insight into how gradient descent actually works, which proves valuable when debugging or customizing architectures later.

### Do the PyTorch and TensorFlow notebooks produce identical results?

Yes, both implementations train the same neural network architectures on identical datasets and should converge to similar accuracy levels. The microsoft/AI-For-Beginners curriculum specifically designs both versions to achieve equivalent learning outcomes despite differing syntax.

### Can I mix code from both frameworks in a single notebook?

While technically possible using interoperability libraries like ONNX, it is not recommended within the AI-For-Beginners curriculum. Each notebook is self-contained and assumes a single framework environment to avoid conflicts between TensorFlow's graph management and PyTorch's dynamic autograd system.

### Why does the PyTorch version require more code for the same task?

PyTorch follows an **explicit programming** philosophy where developers manually control the forward pass, backward pass, and parameter updates. TensorFlow's Keras API embraces **implicit automation**, handling gradient computation and batch iteration internally. This trade-off gives PyTorch greater flexibility for research and custom loss functions, while TensorFlow prioritizes rapid prototyping and standard use cases.