Best Practices for Managing Large Datasets in Jupyter Notebooks: 6 Proven Techniques
Use high-level dataset libraries, chunked reading, caching, and memory-mapping to keep notebooks responsive when working with multi-gigabyte datasets.
Managing large datasets in Jupyter notebooks presents a common challenge for data scientists and machine learning practitioners. Without proper techniques, notebooks can crash from memory exhaustion or grind to a halt during iterative experimentation. The microsoft/AI-For-Beginners curriculum demonstrates production-ready patterns for handling substantial data efficiently while maintaining an interactive workflow.
Leverage High-Level Dataset Libraries
Rather than manually downloading and parsing raw files, use libraries that provide streaming, sharding, and caching out of the box.
TensorFlow Datasets (tfds.load) automatically downloads, caches, and splits datasets into train/test partitions. In lessons/5-NLP/13-TextRep/TextRepresentationTF.ipynb, the tfds.load function returns a tf.data.Dataset object that streams records on demand:
import tensorflow_datasets as tfds
# Loads the 'ag_news_subset' dataset, caches it locally, and returns a tf.data.Dataset
train_ds, test_ds = tfds.load('ag_news_subset', split=['train', 'test'], as_supervised=True)
# Iterate in batches without ever loading the whole dataset into memory
train_batch = train_ds.shuffle(1024).batch(32)
for text, label in train_batch.take(1):
print(text.shape, label.shape)
PyTorch DataLoader couples a dataset object with lazy batching, optional multi-process loading, and automatic shuffling. The CBoW-PyTorch.ipynb notebook in lessons/5-NLP/15-LanguageModeling/ demonstrates this pattern:
import torch
from torch.utils.data import DataLoader
import torchtext.datasets as datasets
# torchtext returns an iterator; we wrap it in a DataLoader for batching
train_iter, test_iter = datasets.AG_NEWS(root='./data')
train_loader = DataLoader(list(train_iter), batch_size=64, shuffle=True, num_workers=4)
for batch in train_loader:
texts, labels = batch
# `texts` is a list of raw strings; you can tokenize on-the-fly here
break
These abstractions eliminate boilerplate while ensuring memory remains bounded regardless of dataset size.
Read Data in Chunks or Batches
For tabular formats like CSV, avoid loading entire files into memory. Pandas supports chunked iteration via the chunksize parameter, returning an iterator of DataFrames processed sequentially.
The NER-TF.ipynb notebook in lessons/5-NLP/19-NER/ applies this technique to handle large NER datasets:
import pandas as pd
chunks = pd.read_csv('large_dataset.csv', chunksize=10_000) # 10k rows per chunk
for i, chunk in enumerate(chunks):
# Perform preprocessing on each chunk
processed = chunk.assign(new_col=chunk['old_col'] * 2)
# Optionally write to an HDF5 store for later fast access
processed.to_hdf('processed.h5', key=f'chunk_{i}', mode='a')
This pattern keeps peak memory usage constant relative to chunk size rather than file size.
Cache Intermediate Results
Expensive preprocessing steps—tokenization, image augmentation, feature extraction—should persist results to disk. Subsequent notebook runs load cached files instead of recomputing, dramatically reducing I/O time and ensuring reproducibility.
Common cache formats include:
- TFRecord for TensorFlow pipelines
- NumPy
.npyor.npzfor dense numeric arrays - HDF5 for hierarchical datasets with metadata
- Parquet for columnar tabular data
Store processed data alongside raw sources in a dedicated directory structure.
Use Memory-Mapped Arrays for Large Numeric Data
NumPy's memmap treats a file on disk as an array without loading the entire file into memory. This technique excels for high-resolution image collections or large feature matrices that exceed available RAM.
import numpy as np
# Creates a memmap that reads data on demand
mm = np.memmap('features.npy', dtype='float32', mode='r', shape=(100_000, 512))
# Access a slice without loading everything
sample = mm[0:10] # loads only the first 10 rows
Access patterns remain identical to standard NumPy arrays, with paging handled transparently by the operating system.
Persist Datasets in a Dedicated Folder
Centralize all downloaded datasets in a single, clearly named directory (conventionally ./data). This prevents repeated downloads, simplifies cleanup, and enables sharing across multiple notebooks.
The BodySegmentation.ipynb lab in lessons/4-ComputerVision/12-Segmentation/lab/ follows this convention:
dataset_path = 'segmentation_full_body_mads_dataset_1192_img'
Absolute or relative paths are abstracted into variables, making notebooks portable and data dependencies explicit.
Lazy-Load Large Image Datasets with Generators
For computer vision workflows, wrap image loading logic in a generator that yields batches each iteration. Both Keras and PyTorch support this pattern natively.
- Keras: Legacy
fit_generator(superseded bytf.datapipelines) or moderntf.data.Datasetfrom generator - PyTorch:
DataLoaderwith customDatasetsubclass implementing__getitem__and__len__
The StyleTransfer_Keras.ipynb notebook in lessons/4-ComputerVision/10-GANs/ demonstrates loading MNIST via keras.datasets with batch training, while the segmentation lab shows generator-based image loading for custom datasets.
Summary
- Adopt high-level APIs like
tfds.loadandDataLoaderto offload streaming, caching, and batching to battle-tested libraries - Process data in chunks using Pandas
chunksizeor equivalent iterators to bound memory consumption - Cache expensive preprocessing to disk in formats like TFRecord, HDF5, or NumPy binary
- Employ memory-mapping via
numpy.memmapfor array data larger than physical RAM - Centralize data storage in dedicated folders with consistent path conventions across notebooks
- Use generators or
tf.datapipelines for lazy-loading image datasets during model training
Frequently Asked Questions
How do I prevent Jupyter from crashing when loading a multi-gigabyte CSV?
Use Pandas read_csv with the chunksize parameter, as demonstrated in lessons/5-NLP/19-NER/NER-TF.ipynb. This returns an iterator yielding DataFrames of specified row counts, keeping memory usage constant. Process each chunk independently or aggregate results incrementally.
What is the difference between TensorFlow Datasets and PyTorch DataLoader for large data?
Both provide lazy loading, but TensorFlow Datasets (tfds.load) additionally handles download, extraction, versioning, and canonical train/test splits automatically. PyTorch DataLoader requires manual dataset implementation but offers more flexible custom sampling and multi-process loading via num_workers.
When should I use NumPy memory-mapping versus standard array loading?
Use numpy.memmap when your array exceeds 50-70% of available RAM or when you need random access to arbitrary slices without loading the entire file. Memory-mapped arrays trade some performance for dramatically reduced memory footprint and faster initialization times.
How does caching improve notebook reproducibility?
Caching stores deterministic preprocessing outputs, ensuring identical inputs across notebook restarts. This eliminates variability from random augmentation, timestamp-based splits, or external API calls. Store caches alongside version-controlled code or document their generation steps for full reproducibility.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →