PyTorch vs TensorFlow Notebook Implementations: A Side-by-Side Comparison
TensorFlow uses a high-level declarative API with model.fit(), while PyTorch requires manual training loops and explicit tensor operations, giving researchers finer control over the training process.
Both frameworks teach identical deep learning concepts—convolutional neural networks for MNIST digit recognition—but their notebook implementations in Microsoft's AI-For-Beginners repository reveal fundamentally different philosophies. This guide examines the actual code in ConvNetsTF.ipynb and ConvNetsPyTorch.ipynb to show exactly how these PyTorch vs TensorFlow notebook implementations diverge in practice.
How Model Definitions Differ
The most immediate distinction appears in how you architect neural networks.
TensorFlow/Keras: Sequential API
TensorFlow leverages Keras's Sequential API for compact, readable model definitions. From lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb:
import tensorflow as tf
from tensorflow import keras
model = keras.models.Sequential([
keras.layers.Conv2D(filters=9, kernel_size=(5,5),
input_shape=(28,28,1), activation='relu'),
keras.layers.Flatten(),
keras.layers.Dense(10)
])
model.compile(loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),
metrics=['acc'])
Key characteristics:
- Layer declaration and activation binding happen in one line
input_shapedefined once at the first layer- Compilation separates architecture from optimizer/loss configuration
PyTorch: Module Subclassing
PyTorch demands explicit class inheritance from nn.Module. From the PyTorch notebook implementation:
import torch
import torch.nn as nn
from torchinfo import summary
class OneConv(nn.Module):
def __init__(self):
super(OneConv, self).__init__()
self.conv = nn.Conv2d(in_channels=1, out_channels=9, kernel_size=(5,5))
self.flatten = nn.Flatten()
self.fc = nn.Linear(5184, 10)
def forward(self, x):
x = nn.functional.relu(self.conv(x))
x = self.flatten(x)
x = nn.functional.log_softmax(self.fc(x), dim=1)
return x
net = OneConv()
Key characteristics:
forward()method manually wires data flow- Activations applied explicitly (not layer-bound)
in_channels/out_channelsnaming versus TensorFlow'sfilters
Training Loop Implementation
This is where PyTorch vs TensorFlow notebook implementations show their deepest architectural split.
TensorFlow: Declarative model.fit()
One line handles the entire epoch loop, batching, gradient computation, and metric tracking:
history = model.fit(x_train_c, y_train,
validation_data=(x_test_c, y_test),
epochs=5)
model.fit() returns a History object containing loss and accuracy arrays for plotting.
PyTorch: Explicit Training Function
The PyTorch notebook delegates to a custom train() helper from the pytorchcv library:
from pytorchcv import train, plot_results
hist = train(net, train_loader, test_loader, epochs=5)
plot_results(hist)
Behind this abstraction, the actual training loop (implemented in pytorchcv.py) manually:
- Iterates batches from
DataLoader - Computes
nn.CrossEntropyLoss() - Calls
optimizer.zero_grad(),.backward(), and.step()
This imperative style matches how research papers describe algorithms—step by step, fully visible.
Data Loading Architecture
| Aspect | TensorFlow Notebook | PyTorch Notebook |
|---|---|---|
| Import | keras.datasets.mnist.load_data() |
from pytorchcv import load_mnist |
| Returns | NumPy arrays (60000, 28, 28) |
DataLoader objects with batching |
| Preprocessing | Manual NumPy reshape/normalize | Encapsulated in helper |
| GPU handling | Automatic via tf.config |
Explicit torch.device() assignment |
The TensorFlow approach gives raw data you manipulate directly. PyTorch's DataLoader provides lazy batching with shuffling—critical for large datasets that don't fit in memory.
GPU and Device Management
TensorFlow automatically detects and uses GPUs through:
tf.config.list_physical_devices('GPU')
PyTorch requires explicit device management throughout pytorchcv.py helper functions:
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
# Then: tensor.to(device), model.to(device)
This explicitness prevents silent CPU fallback bugs but adds boilerplate.
Model Inspection and Visualization
Both notebooks generate layer summaries, but with different tools:
TensorFlow (built-in):
model.summary()
PyTorch (requires external package):
from torchinfo import summary
summary(net, input_size=(1,1,28,28))
Visualization helpers like plot_convolution exist in both tfcv.py and pytorchcv.py, but PyTorch needs explicit tensor conversion before plotting.
When to Choose Each Implementation Style
- TensorFlow notebooks suit production engineering, rapid prototyping, and teams prioritizing standardized pipelines
- PyTorch notebooks fit research experimentation, custom loss functions, and educational contexts where understanding internals matters
The Microsoft AI-For-Beginners repository deliberately maintains both to demonstrate that framework choice involves trade-offs between abstraction and control.
Summary
- Model definition: TensorFlow uses
keras.Sequential; PyTorch subclassesnn.Modulewith explicitforward() - Training: TensorFlow's
model.fit()vs PyTorch's manual batch iteration andoptimizer.step() - Data flow: NumPy arrays versus
DataLoaderwith automatic batching - GPU handling: Automatic configuration versus explicit
torch.device()management - File locations:
lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynbandConvNetsPyTorch.ipynbin the microsoft/AI-For-Beginners repository
Both implementations achieve comparable MNIST accuracy, validating that framework choice affects how you code more than what you can build.
Frequently Asked Questions
Which notebook is better for beginners?
The TensorFlow notebook has gentler initial complexity due to Keras's high-level API. However, the PyTorch notebook builds deeper understanding of how training actually works—valuable when you need to customize beyond standard patterns.
Can I mix TensorFlow and PyTorch code in the same project?
Technically possible through ONNX conversion or frameworks like Hugging Face's accelerate, but Microsoft's notebooks keep them separate to demonstrate pure-idiom implementations. Each notebook in AI-For-Beginners is self-contained.
Why does the PyTorch notebook need a helper library (pytorchcv) while TensorFlow doesn't?
PyTorch's lower-level design requires more boilerplate for common tasks. The pytorchcv.py helper provides load_mnist, train, and plot_results to match Keras's convenience—without hiding the framework's imperative nature. TensorFlow's built-in keras.datasets and model.fit() already provide this abstraction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →