Differences Between PyTorch and TensorFlow Versions of Dual-Framework Notebooks

The PyTorch notebooks require explicit manual training loops and device management, while the TensorFlow notebooks leverage high-level Keras APIs that automate gradient computation and model optimization.

The microsoft/AI-For-Beginners repository provides dual-framework notebooks that teach identical machine learning concepts using both PyTorch and TensorFlow 2.x implementations. While both versions achieve the same learning objectives—loading datasets, defining convolutional architectures, and evaluating model performance—the code structure differs significantly due to each framework's underlying design philosophy. Understanding these differences helps learners choose the approach that best fits their workflow, whether they prefer TensorFlow's abstraction or PyTorch's explicit control.

API Architecture and Model Definition

The primary divergence between the PyTorch and TensorFlow versions begins with how models are constructed and imported.

Import Conventions and Module Structure

In the TensorFlow notebooks—such as lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb—the code relies on Keras layers bundled within TensorFlow:

import tensorflow as tf
from tensorflow import keras
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Dense

Conversely, the PyTorch notebooks found in lessons/4-ComputerVision/07-ConvNets/ConvNetsPyTorch.ipynb require separate imports for network layers, optimization utilities, and data handling:

import torch
import torch.nn as nn
import torch.optim as optim
import torchvision

Model Definition: Subclassing vs Sequential API

TensorFlow utilizes the Keras Sequential API for rapid prototyping. Layers stack declaratively, and the framework handles forward-pass logic internally:

model = tf.keras.Sequential([
    tf.keras.layers.Conv2D(32, 3, activation='relu', input_shape=(28,28,1)),
    tf.keras.layers.MaxPooling2D(),
    tf.keras.layers.Flatten(),
    tf.keras.layers.Dense(128, activation='relu'),
    tf.keras.layers.Dense(10)
])

PyTorch requires explicit subclassing of nn.Module. You must define a forward(self, x) method that specifies exactly how data flows through the network:

class ConvNet(nn.Module):
    def __init__(self):
        super(ConvNet, self).__init__()
        self.conv1 = nn.Conv2d(1, 32, 3)
        self.pool = nn.MaxPool2d(2, 2)
        self.fc1 = nn.Linear(64*5*5, 128)
        self.fc2 = nn.Linear(128, 10)

    def forward(self, x):
        x = self.pool(torch.relu(self.conv1(x)))
        x = x.view(-1, 64*5*5)
        x = torch.relu(self.fc1(x))
        return self.fc2(x)

Data Loading and Pipeline Differences

The dual-framework notebooks demonstrate distinct approaches to data ingestion that reflect each library's ecosystem.

TensorFlow versions typically use tf.keras.datasets helpers that return NumPy arrays immediately ready for training:

(train_x, train_y), (test_x, test_y) = keras.datasets.mnist.load_data()
train_x = train_x[..., np.newaxis] / 255.0

PyTorch implementations rely on torchvision.datasets combined with torch.utils.data.DataLoader for batched iteration. This requires defining transformation pipelines and explicitly shuffling data:

transform = transforms.Compose([transforms.ToTensor()])
train_set = torchvision.datasets.MNIST(root='./data', train=True, download=True, transform=transform)
train_loader = torch.utils.data.DataLoader(train_set, batch_size=64, shuffle=True)

The PyTorch approach provides fine-grained control over batching and augmentation, while the TensorFlow method prioritizes immediate usability.

Training Loop Implementation

The most significant practical difference lies in how training executes.

TensorFlow High-Level Abstraction

TensorFlow notebooks compile the model with loss functions and optimizers, then invoke model.fit() to handle epochs, batching, and gradient updates automatically:

model.compile(
    loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
    optimizer=tf.keras.optimizers.Adam(),
    metrics=['accuracy']
)

model.fit(train_x, train_y, epochs=5, batch_size=64, validation_split=0.1)
test_loss, test_acc = model.evaluate(test_x, test_y)

PyTorch Explicit Training Loop

PyTorch notebooks require manual iteration over the DataLoader, explicit loss computation, and manual backpropagation steps. The ConvNetsPyTorch.ipynb file demonstrates this low-level pattern:

criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters())

for epoch in range(5):
    running_loss = 0.0
    for images, labels in train_loader:
        images, labels = images.to(device), labels.to(device)
        
        optimizer.zero_grad()
        outputs = model(images)
        loss = criterion(outputs, labels)
        loss.backward()
        optimizer.step()
        running_loss += loss.item()

Evaluation also requires manual computation:

correct = 0
total = 0
with torch.no_grad():
    for images, labels in test_loader:
        outputs = model(images.to(device))
        _, predicted = torch.max(outputs, 1)
        total += labels.size(0)
        correct += (predicted == labels.cpu()).sum().item()

Device Management and Execution Context

Both frameworks support GPU acceleration, but handle device placement differently.

TensorFlow abstracts device management unless explicitly scoped. Tensors typically reside on the default device, and you can force GPU usage via tf.device('/GPU:0') context managers.

PyTorch requires explicit device handling. The notebooks include standard boilerplate to detect CUDA availability and move both models and tensors manually:

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = ConvNet().to(device)

Every batch must explicitly transfer to the target device: images, labels = images.to(device), labels.to(device).

Model Persistence and Serialization

Saving trained models differs significantly between the two frameworks.

TensorFlow uses the Keras H5 format or SavedModel format:

model.save('my_model.h5')
loaded_model = tf.keras.models.load_model('my_model.h5')

PyTorch saves only the model's state_dict (learned parameters), requiring you to reinstantiate the architecture class before loading weights:

torch.save(model.state_dict(), 'model.pt')
model = ConvNet()
model.load_state_dict(torch.load('model.pt'))

Summary

  • TensorFlow notebooks utilize high-level Keras APIs with model.fit() and model.compile(), minimizing boilerplate code for standard training loops.
  • PyTorch notebooks require explicit manual training loops with loss.backward() and optimizer.step(), offering granular control over gradient computation.
  • Data loading in TensorFlow relies on tf.keras.datasets returning NumPy arrays, while PyTorch uses torchvision.datasets with DataLoader for batch management.
  • Device management is automatic in TensorFlow but requires explicit .to(device) calls for models and tensors in PyTorch.
  • Model saving in TensorFlow preserves the complete architecture and weights in H5 format, whereas PyTorch saves only parameters via state_dict dictionaries.

Frequently Asked Questions

Which framework version should beginners start with?

Beginners often find the TensorFlow/Keras versions more approachable because the model.fit() API abstracts away the training loop complexity. However, the PyTorch versions provide clearer insight into how gradient descent actually works, which proves valuable when debugging or customizing architectures later.

Do the PyTorch and TensorFlow notebooks produce identical results?

Yes, both implementations train the same neural network architectures on identical datasets and should converge to similar accuracy levels. The microsoft/AI-For-Beginners curriculum specifically designs both versions to achieve equivalent learning outcomes despite differing syntax.

Can I mix code from both frameworks in a single notebook?

While technically possible using interoperability libraries like ONNX, it is not recommended within the AI-For-Beginners curriculum. Each notebook is self-contained and assumes a single framework environment to avoid conflicts between TensorFlow's graph management and PyTorch's dynamic autograd system.

Why does the PyTorch version require more code for the same task?

PyTorch follows an explicit programming philosophy where developers manually control the forward pass, backward pass, and parameter updates. TensorFlow's Keras API embraces implicit automation, handling gradient computation and batch iteration internally. This trade-off gives PyTorch greater flexibility for research and custom loss functions, while TensorFlow prioritizes rapid prototyping and standard use cases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →