Setting up GPU Acceleration for PyTorch vs TensorFlow in AI-For-Beginners

Both PyTorch and TensorFlow require explicit device configuration to leverage GPU acceleration, but PyTorch uses manual tensor placement while TensorFlow needs memory growth configuration to prevent out-of-memory errors.

The Microsoft AI-For-Beginners curriculum provides production-ready implementations demonstrating how to enable GPU support for deep learning workflows. Setting up GPU acceleration for PyTorch vs TensorFlow involves distinct patterns for device detection and memory management, which this guide extracts directly from the repository's source code.

PyTorch GPU Setup: Explicit Device Management

PyTorch follows an imperative, dynamic graph approach where you manually specify the compute device for every tensor and model.

Device Detection and Default Configuration

In lessons/4-ComputerVision/08-TransferLearning/pytorchcv.py, the curriculum defines a reusable default_device variable at line 17 that automatically detects CUDA availability:


# https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/08-TransferLearning/pytorchcv.py#L17-L34

import torch

default_device = 'cuda' if torch.cuda.is_available() else 'cpu'

The torch.cuda.is_available() function returns a Boolean indicating whether CUDA-capable hardware and drivers are present. This single variable centralizes device selection, allowing the same code to run on CPU-only laptops or GPU-powered cloud instances without modification.

Tensor-wise GPU Placement

PyTorch requires explicit movement of data to the GPU using the .to(device) method. The repository demonstrates this pattern in training loops:


# Move data and model to the selected device

lbls = labels.to(default_device)
out = net(features.to(default_device))

Once tensors reside on the GPU, all subsequent operations—including torch.max and torch.nn.functional calls—execute on the accelerator automatically. This explicit approach yields typical speedups of 5–10× for MNIST-style computer vision workloads compared to CPU execution.

TensorFlow GPU Setup: Device Enumeration and Memory Control

TensorFlow 2.x uses eager execution by default but requires specific configuration to manage GPU memory allocation behavior.

Detecting GPU Devices

The TensorFlow lessons in lessons/4-ComputerVision/07-ConvNets/tfcv.py begin with device enumeration at lines 3-7:


# https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/07-ConvNets/tfcv.py#L3-L7

import tensorflow as tf

physical_devices = tf.config.list_physical_devices('GPU')

Unlike PyTorch's simple Boolean check, tf.config.list_physical_devices('GPU') returns a list of available GPU hardware, providing granular control over multi-GPU systems.

Configuring Memory Growth

By default, TensorFlow allocates all available GPU memory upfront, which causes out-of-memory errors when loading multiple models or running several notebooks simultaneously. The curriculum implements the recommended pattern at lines 8-11:


# https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/07-ConvNets/tfcv.py#L8-L11

if physical_devices:
    tf.config.experimental.set_memory_growth(physical_devices[0], True)

Enabling memory growth allows TensorFlow to allocate GPU memory incrementally as needed rather than reserving the entire device memory at startup.

Implicit Device Placement

After configuration, TensorFlow automatically places Keras layers on the first visible GPU. You can also use explicit device scopes when needed:

with tf.device('/GPU:0'):
    model = tf.keras.applications.ResNet50(weights='imagenet')

Comparing GPU Setup Patterns

PyTorch requires manual tensor placement via .to(device) but handles memory dynamically without configuration flags. TensorFlow automatically places tensors after device enumeration but requires set_memory_growth to prevent memory allocation conflicts.

Aspect PyTorch Implementation TensorFlow Implementation
Device Detection torch.cuda.is_available() returns Boolean tf.config.list_physical_devices('GPU') returns device list
Memory Management Dynamic allocation as needed Must explicitly enable set_memory_growth to avoid OOM errors
Code Footprint 5–7 lines for detection and tensor movement 8–10 lines for enumeration and memory configuration
Placement Style Explicit .to(device) calls on every tensor Implicit once GPU is configured; optional tf.device contexts

Both frameworks automatically fall back to CPU execution when torch.cuda.is_available() returns False or when list_physical_devices('GPU') returns an empty list, ensuring the notebooks run seamlessly across local laptops, Azure Data Science VMs, and Google Colab environments.

Key Implementation Files

The following source files in the AI-For-Beginners repository demonstrate these GPU acceleration patterns:

Summary

  • PyTorch GPU setup relies on torch.cuda.is_available() to select between 'cuda' and 'cpu', then requires explicit .to(device) calls to move tensors and models to the accelerator
  • TensorFlow GPU setup uses tf.config.list_physical_devices('GPU') to enumerate hardware and must configure set_memory_growth to prevent aggressive memory pre-allocation
  • Both approaches achieve 5–10× speedups on computer vision workloads but differ in memory management philosophy: PyTorch allocates on demand while TensorFlow defaults to reserving all GPU memory
  • The AI-For-Beginners curriculum provides working implementations in pytorchcv.py and tfcv.py that handle CPU fallback automatically

Frequently Asked Questions

How do I check if my GPU is available in PyTorch according to the AI-For-Beginners code?

Use torch.cuda.is_available() as implemented in lessons/4-ComputerVision/08-TransferLearning/pytorchcv.py. This Boolean function returns True when CUDA drivers and compatible hardware are detected, allowing the script to set default_device = 'cuda'; otherwise it falls back to 'cpu'.

Why does the TensorFlow code in AI-For-Beginners enable memory growth?

The tf.config.experimental.set_memory_growth(physical_devices[0], True) configuration prevents TensorFlow from allocating all GPU memory at startup. Without this setting, loading multiple models or running several notebooks simultaneously triggers out-of-memory errors, as shown in lessons/4-ComputerVision/07-ConvNets/tfcv.py.

Can I run the AI-For-Beginners notebooks on a CPU-only machine?

Yes. Both frameworks automatically handle CPU fallback. PyTorch checks torch.cuda.is_available() and defaults to 'cpu', while TensorFlow's list_physical_devices('GPU') returns an empty list when no GPU is present, causing the code to execute on the CPU without modification.

What is the performance difference between CPU and GPU in these examples?

According to the repository's implementations, GPU acceleration provides typical speedups of 5–10× for MNIST-style computer vision workloads. The actual improvement depends on batch size and model complexity, but both PyTorch and TensorFlow versions in the curriculum demonstrate this performance gain once tensors are properly placed on the GPU.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →