GPU Memory Requirements for Batch Sizes and Context Sizes in RPD‑DNN
The RPD‑DNN model requires approximately 300 MB–4.8 GB of GPU memory for activation tensors (depending on batch sizes from 32–256 and context sizes from 200–400), plus a static 450 MB for the ELMo encoder and LSTM parameters.
The RPD‑DNN (Rumor Detection Deep Neural Network) repository at jerrygaolondon/rpdnn implements a hybrid architecture that combines a large‑scale ELMo encoder with bidirectional LSTM social‑context encoders. Understanding the GPU memory requirements for different batch sizes and context sizes is critical when configuring training jobs for hardware ranging from 12 GB Tesla K40M GPUs to 24 GB K80 accelerators.
What Drives GPU Memory Usage in RPD‑DNN
GPU memory consumption in this codebase is dominated by four distinct components. In src/allennlp_rumor_classifier.py, the model allocates tensors for ELMo embeddings, context features, and LSTM hidden states that scale directly with your chosen batch and sequence dimensions.
| Component | Bytes per Element (float32) | Scaling Factor |
|---|---|---|
| ELMo embeddings (per token) | 1 024 × 4 B = 4 KB | batch_size × max_sentence_len |
| Context embeddings (content + metadata) | ~4 KB (content) + ~0.1 KB (metadata) | batch_size × max_cxt_size |
| LSTM hidden states (bidirectional, 2×dim) | 2 048 × 4 B ≈ 8 KB | batch_size × max_cxt_size |
| Model parameters (ELMo + LSTMs + FF) | ~450 MB total | Fixed (independent of batch) |
The constant MAXIMUM_CONTEXT_SEQ_SIZE = 200 defined at line 90 of src/allennlp_rumor_classifier.py establishes the default context window, while the comment at lines 1020‑1023 warns that very long context sequences can exhaust GPU memory. Additionally, src/rumour_dnn_trainer.py (lines 148‑152) explicitly recommends setting the batch size as large as the GPU memory allows to maximize throughput.
Memory Estimates by Batch and Context Configuration
The following estimates cover only activation tensors (embeddings, LSTM outputs, and masks) and exclude the permanent parameter storage (~450 MB). Actual usage may be higher due to PyTorch’s caching allocator and temporary gradient buffers.
| Batch Size | Max Context Size | Approximate Activation Memory |
|---|---|---|
| 32 | 200 (default) | ~300 MB |
| 64 | 200 | ~600 MB |
| 128 | 200 | ~1.2 GB |
| 256 | 200 | ~2.4 GB |
| 128 | 400 (doubled) | ~2.4 GB |
| 256 | 400 | ~4.8 GB |
Doubling the context size from 200 to 400 effectively doubles the memory required for context‑related tensors and LSTM states, making it equivalent to doubling the batch size in terms of GPU pressure.
Hardware Compatibility and Limits
According to README.md (lines 49‑50), the repository was developed and tested on two GPU configurations:
- NVIDIA Tesla K40M – 12 GB per GPU
- NVIDIA Tesla K80 – 24 GB per GPU
Practical limits:
- On a K40M (12 GB), you can safely train with
batch_size = 128and the defaultmax_cxt_size = 200(total footprint ~1.7 GB including parameters, leaving ample headroom for gradients and optimizer states). - On a K80 (24 GB), you may increase either dimension: use
batch_size = 256withmax_cxt_size = 200, or maintainbatch_size = 128and raisemax_cxt_sizeto 400 without exceeding memory capacity.
Adjusting Batch Size and Context Length in Code
You can override the default context length using the --max_cxt_size command‑line argument defined in src/rumour_dnn_trainer.py (lines 60‑62). Batch size is hard‑coded in the training script but can be modified directly in the source.
Example: Default Training on 12 GB GPU
python src/rumour_dnn_trainer.py \
-t data/train/bostonbombings/aug_rnr_train_set_combined.csv \
--heldout data/train/bostonbombings/aug_rnr_heldout_set_combined.csv \
-e data/test/bostonbombings.csv \
-p bostonbombings \
-g 0 \
-f -1 \
--max_cxt_size 200
The default train_batch_size = 128 at line 152 of src/rumour_dnn_trainer.py is optimized for 12 GB devices.
Example: Large Context Training on 24 GB GPU
To process longer rumor threads, increase the context size and reduce the batch size to stay within memory limits:
# Edit src/rumour_dnn_trainer.py line 152 to set train_batch_size = 64
python src/rumour_dnn_trainer.py \
-t data/train/bostonbombings/aug_rnr_train_set_combined.csv \
--heldout data/train/bostonbombings/aug_rnr_heldout_set_combined.csv \
-e data/test/bostonbombings.csv \
-p bostonbombings \
-g 0 \
--max_cxt_size 400
Example: Programmatic Inference with Custom Sizes
from src.allennlp_rumor_classifier import instantiate_rumour_model, config_gpu_use
import torch
config_gpu_use(0) # Enable CUDA device 0
model = instantiate_rumour_model(
n_gpu=0,
vocab=my_vocabulary, # Load your Vocabulary instance
feature_setting=1,
max_cxt_size=300 # Override default 200
)
model.eval()
# Prepare batch dicts and run model(**batch)
Summary
- GPU memory in RPD‑DNN is consumed primarily by ELMo embeddings (~4 KB/token), bidirectional LSTM hidden states (~8 KB/token), and context feature tensors, while model parameters occupy a fixed ~450 MB.
- Default settings (
batch_size = 128,max_cxt_size = 200) require roughly 1.2 GB for activations, fitting comfortably within a 12 GB K40M GPU. - Scaling to
batch_size = 256ormax_cxt_size = 400can push activation memory to 4.8 GB, necessitating a 24 GB K80 or careful memory management. - Adjust
MAXIMUM_CONTEXT_SEQ_SIZEinsrc/allennlp_rumor_classifier.pyor use the--max_cxt_sizeCLI flag to control sequence length; modify the hard‑codedtrain_batch_sizeinsrc/rumour_dnn_trainer.pyto control batch throughput.
Frequently Asked Questions
What is the default maximum context size in RPD‑DNN?
The default value is 200 tokens, defined as the constant MAXIMUM_CONTEXT_SEQ_SIZE = 200 at line 90 of src/allennlp_rumor_classifier.py. You can override this at runtime using the --max_cxt_size argument in src/rumour_dnn_trainer.py (lines 60‑62).
How can I train RPD‑DNN if my GPU has less than 12 GB of memory?
Reduce the batch size below 128 in src/rumour_dnn_trainer.py (line 152) and decrease max_cxt_size below 200. Alternatively, implement gradient accumulation by running multiple forward passes with smaller batches before calling optimizer.step(), though this requires minor scripting modifications not provided in the original codebase.
Why does increasing context size consume more memory than increasing batch size?
Both dimensions increase memory linearly, but the LSTM hidden states (which scale with batch_size × max_cxt_size) consume 8 KB per element in float32 due to the bidirectional 2 048‑dimensional hidden layer. Doubling the context length doubles the sequence dimension for all context‑related tensors simultaneously, creating the same multiplicative pressure as doubling the batch size.
What happens if I set a batch size that exceeds available GPU memory?
As noted in src/allennlp_rumor_classifier.py (lines 1020‑1023), very long context sequences or excessively large batches will cause CUDA out‑of‑memory errors during the forward pass through the social‑context encoders. The PyTorch caching allocator may also fragment memory, causing sporadic failures even when theoretical usage appears below the hardware limit.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →