# PyTorch vs TensorFlow Notebook Implementations: A Side-by-Side Comparison

> Compare PyTorch and TensorFlow notebook implementations. See how TensorFlow's model.fit() differs from PyTorch's manual training loops and tensor operations for researcher control.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-25

---

**TensorFlow uses a high-level declarative API with `model.fit()`, while PyTorch requires manual training loops and explicit tensor operations, giving researchers finer control over the training process.**

Both frameworks teach identical deep learning concepts—convolutional neural networks for MNIST digit recognition—but their notebook implementations in Microsoft's AI-For-Beginners repository reveal fundamentally different philosophies. This guide examines the actual code in `ConvNetsTF.ipynb` and `ConvNetsPyTorch.ipynb` to show exactly how these PyTorch vs TensorFlow notebook implementations diverge in practice.

## How Model Definitions Differ

The most immediate distinction appears in how you architect neural networks.

### TensorFlow/Keras: Sequential API

TensorFlow leverages Keras's **Sequential API** for compact, readable model definitions. From `lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb`:

```python
import tensorflow as tf
from tensorflow import keras

model = keras.models.Sequential([
    keras.layers.Conv2D(filters=9, kernel_size=(5,5), 
                        input_shape=(28,28,1), activation='relu'),
    keras.layers.Flatten(),
    keras.layers.Dense(10)
])
model.compile(loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),
              metrics=['acc'])

```

**Key characteristics:**
- Layer declaration and activation binding happen in one line
- `input_shape` defined once at the first layer
- Compilation separates architecture from optimizer/loss configuration

### PyTorch: Module Subclassing

PyTorch demands explicit class inheritance from `nn.Module`. From the PyTorch notebook implementation:

```python
import torch
import torch.nn as nn
from torchinfo import summary

class OneConv(nn.Module):
    def __init__(self):
        super(OneConv, self).__init__()
        self.conv = nn.Conv2d(in_channels=1, out_channels=9, kernel_size=(5,5))
        self.flatten = nn.Flatten()
        self.fc = nn.Linear(5184, 10)

    def forward(self, x):
        x = nn.functional.relu(self.conv(x))
        x = self.flatten(x)
        x = nn.functional.log_softmax(self.fc(x), dim=1)
        return x

net = OneConv()

```

**Key characteristics:**
- `forward()` method manually wires data flow
- Activations applied explicitly (not layer-bound)
- `in_channels/out_channels` naming versus TensorFlow's `filters`

## Training Loop Implementation

This is where PyTorch vs TensorFlow notebook implementations show their deepest architectural split.

### TensorFlow: Declarative `model.fit()`

One line handles the entire epoch loop, batching, gradient computation, and metric tracking:

```python
history = model.fit(x_train_c, y_train,
                    validation_data=(x_test_c, y_test),
                    epochs=5)

```

`model.fit()` returns a **History object** containing loss and accuracy arrays for plotting.

### PyTorch: Explicit Training Function

The PyTorch notebook delegates to a custom `train()` helper from the `pytorchcv` library:

```python
from pytorchcv import train, plot_results

hist = train(net, train_loader, test_loader, epochs=5)
plot_results(hist)

```

Behind this abstraction, the actual training loop (implemented in [`pytorchcv.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/pytorchcv.py)) manually:
- Iterates batches from `DataLoader`
- Computes `nn.CrossEntropyLoss()`
- Calls `optimizer.zero_grad()`, `.backward()`, and `.step()`

This imperative style matches how research papers describe algorithms—step by step, fully visible.

## Data Loading Architecture

| Aspect | TensorFlow Notebook | PyTorch Notebook |
|--------|---------------------|------------------|
| **Import** | `keras.datasets.mnist.load_data()` | `from pytorchcv import load_mnist` |
| **Returns** | NumPy arrays `(60000, 28, 28)` | `DataLoader` objects with batching |
| **Preprocessing** | Manual NumPy reshape/normalize | Encapsulated in helper |
| **GPU handling** | Automatic via `tf.config` | Explicit `torch.device()` assignment |

The TensorFlow approach gives raw data you manipulate directly. PyTorch's `DataLoader` provides **lazy batching with shuffling**—critical for large datasets that don't fit in memory.

## GPU and Device Management

TensorFlow automatically detects and uses GPUs through:

```python
tf.config.list_physical_devices('GPU')

```

PyTorch requires explicit device management throughout [`pytorchcv.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/pytorchcv.py) helper functions:

```python
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')

# Then: tensor.to(device), model.to(device)

```

This explicitness prevents silent CPU fallback bugs but adds boilerplate.

## Model Inspection and Visualization

Both notebooks generate layer summaries, but with different tools:

**TensorFlow** (built-in):

```python
model.summary()

```

**PyTorch** (requires external package):

```python
from torchinfo import summary
summary(net, input_size=(1,1,28,28))

```

Visualization helpers like `plot_convolution` exist in both [`tfcv.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/tfcv.py) and [`pytorchcv.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/pytorchcv.py), but PyTorch needs explicit tensor conversion before plotting.

## When to Choose Each Implementation Style

- **TensorFlow notebooks** suit production engineering, rapid prototyping, and teams prioritizing standardized pipelines
- **PyTorch notebooks** fit research experimentation, custom loss functions, and educational contexts where understanding internals matters

The Microsoft AI-For-Beginners repository deliberately maintains both to demonstrate that framework choice involves trade-offs between **abstraction and control**.

## Summary

- **Model definition**: TensorFlow uses `keras.Sequential`; PyTorch subclasses `nn.Module` with explicit `forward()`
- **Training**: TensorFlow's `model.fit()` vs PyTorch's manual batch iteration and `optimizer.step()`
- **Data flow**: NumPy arrays versus `DataLoader` with automatic batching
- **GPU handling**: Automatic configuration versus explicit `torch.device()` management
- **File locations**: `lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb` and `ConvNetsPyTorch.ipynb` in the microsoft/AI-For-Beginners repository

Both implementations achieve comparable MNIST accuracy, validating that framework choice affects *how* you code more than *what* you can build.

## Frequently Asked Questions

### Which notebook is better for beginners?

The **TensorFlow notebook** has gentler initial complexity due to Keras's high-level API. However, the **PyTorch notebook** builds deeper understanding of how training actually works—valuable when you need to customize beyond standard patterns.

### Can I mix TensorFlow and PyTorch code in the same project?

Technically possible through ONNX conversion or frameworks like Hugging Face's `accelerate`, but Microsoft's notebooks keep them separate to demonstrate pure-idiom implementations. Each notebook in AI-For-Beginners is self-contained.

### Why does the PyTorch notebook need a helper library (`pytorchcv`) while TensorFlow doesn't?

PyTorch's lower-level design requires more boilerplate for common tasks. The [`pytorchcv.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/pytorchcv.py) helper provides `load_mnist`, `train`, and `plot_results` to match Keras's convenience—without hiding the framework's imperative nature. TensorFlow's built-in `keras.datasets` and `model.fit()` already provide this abstraction.