Common Failure Modes and Debugging Strategies for RP-DNN Training
The most common RP-DNN training failures stem from missing dataset files, invalid feature settings, ELMo weight loading errors, GPU out-of-memory issues, and numerical instabilities like NaN gradients, all of which can be diagnosed by verifying file paths in src/rumour_dnn_trainer.py, checking enum values in src/allennlp_rumor_classifier.py, and enabling gradient clipping in src/training_util.py.
RP-DNN (Rumour-Propagation Deep Neural Network) is an AllenNLP-based deep learning framework for rumor detection that combines ELMo embeddings, LSTM or Transformer encoders, and handcrafted social-context features. Because the training pipeline integrates external resources like ELMo weights, the PHEME social-context corpus, and CSV-based tweet datasets, RP-DNN training failures often occur at the intersection of data validation, configuration parsing, and GPU memory management.
Common RP-DNN Training Failure Modes
Missing or Malformed Dataset Files
Before the model is instantiated, src/rumour_dnn_trainer.py (lines 109‑116) attempts to load CSV files specified by the command-line arguments -t/--trainset, --heldout, and -e/--evaluationset. If these paths are invalid, src/data_loader.py (lines 65‑77) raises a FileNotFoundError.
Symptom: Immediate crash with FileNotFoundError before any GPU allocation occurs.
Debugging strategy: Verify paths programmatically using os.path.isfile or the helper load_abs_path to resolve symlinks in the data directory. Ensure the CSV files exist before invoking model_training().
Invalid Feature Setting or Attention Options
The CLI arguments --feature_setting and --attention_option must match enums defined at the top of src/allennlp_rumor_classifier.py. Validation logic in rumour_dnn_trainer.py (lines 118‑124) raises ValueError with messages like “Supported training (feature_setting option) …” when unsupported values are provided.
Symptom: ValueError during argument parsing before training begins.
Debugging strategy: Restrict values to documented constants such as FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY = 1 or ATTENTION_OPTION_HIERARCHICAL = 1. Print the constant list before calling model_training() to verify compatibility.
ELMo Model Loading Failures
RP-DNN relies on ELMo embeddings loaded via ElmoTokenEmbedder. The weight path is configured in rumour_dnn_trainer.py (lines 136‑141), pointing to resource/embedding/elmo_model/elmo_credbank_2x4096_512_2048cnn_2xhighway_weights_10052019.hdf5.
Symptom: Runtime error when the embedder attempts to load the weight file, often referencing missing HDF5 files.
Debugging strategy: Confirm the HDF5 file exists and that symlinks in resource/embedding/ resolve correctly. The load_abs_path helper handles Windows shortcuts and relative paths.
GPU Device Mismatch and Out-of-Memory Errors
GPU configuration occurs via config_gpu_use (called at line 42 of rumour_dnn_trainer.py), and the model is moved to the device in instantiate_rumour_model (lines 63‑66). Out-of-memory errors occur when train_batch_size (default 128) exceeds available VRAM.
Symptom: “CUDA out of memory” or silent CPU fallback causing extremely slow training throughput.
Debugging strategy: Run with -g 0 for GPU 0 or -g -1 for CPU-only mode. Monitor nvidia-smi output (the script prints this at lines 99‑104). If OOM persists, lower train_batch_size or reduce max_cxt_size_option to limit context window memory usage.
NaN Gradients and Loss Explosion
Numerical instability manifests as nan loss or RuntimeError about inconsistent loss production. Gradient clipping utilities reside in src/training_util.py (sparse_clip_norm), with implementation hints in rumour_dnn_trainer.py (lines 92‑96).
Symptom: Training halts with nan loss or gradient explosion errors after several batches.
Debugging strategy: Enable gradient clipping by passing grad_norm=5.0 to the Trainer constructor, or set grad_clipping=1.0 for per-parameter clipping. Uncomment the relevant block in rumour_dnn_trainer.py and experiment with thresholds between 1.0 and 5.0.
Layer-Norm Numerical Instability
A custom LayerNorm implementation in src/my_layer_norm.py (see comment around line 71) can produce NaNs after a few epochs, particularly with stacked LSTM encoders.
Symptom: NaNs appear mid-training despite stable initial epochs, often affecting deep encoder stacks.
Debugging strategy: Replace the custom MyLayerNorm with PyTorch’s built-in torch.nn.LayerNorm. Update instantiate_rumour_model in src/allennlp_rumor_classifier.py to use the standard implementation, which handles edge cases in variance calculation more robustly.
Inconsistent Social-Context Directory Structure
The PHEME corpus loader in src/data_loader.py (load_tweets_context_dataset_dir) expects a strict layout where each rumor thread is a numeric folder containing source-tweets and reactions subdirectories. The global social_context_data_dir variable (line 43) must point to the unpacked corpus.
Symptom: Exception: source tweet (id :…) is not found in current social context dataset!
Debugging strategy: Verify that data/social_context/aug-rnr-annotated-threads-retweets symlinks to the unpacked PHEME corpus. Ensure only numeric tweet ID folders exist at the root—any stray files break the mapping logic in load_tweets_context_dataset_dir.
Serialization Directory Conflicts on Resume
When resuming training, training_util.create_serialization_dir (lines 165‑227) checks for existing directories and config mismatches.
Symptom: ConfigurationError stating “Serialization directory already exists”.
Debugging strategy: Either delete the old serialization_dir before restarting, or invoke the script with --recover alongside identical configuration files. The helper validates key consistency and aborts if the archived config differs from the current run parameters.
Debugging Strategies and Code Solutions
Verify All Input Paths Before Launching Training
Prevent early crashes by validating CSV and model weight paths programmatically.
from pathlib import Path
import sys
def assert_file(p: Path, name: str) -> None:
if not p.is_file():
sys.exit(f"[ERROR] {name} not found: {p}")
train = Path("/data/train/bostonbombings/aug_rnr_train_set_combined.csv")
heldout = Path("/data/train/bostonbombings/aug_rnr_heldout_set_combined.csv")
test = Path("/data/test/bostonbombings.csv")
for p, n in [(train, "train set"), (heldout, "held‑out set"), (test, "test set")]:
assert_file(p, n)
print("✅ All CSVs exist – you can safely call model_training()")
Reference: rumour_dnn_trainer.py (lines 109‑116)
Enable Gradient Clipping to Avoid NaNs
Stabilize training by clipping gradients before the optimizer step.
# Inside rumour_dnn_trainer.py, before Trainer creation:
optimizer = optim.Adam(model.parameters(), lr=1e-4, weight_decay=1e-5)
# Add clipping via the grad_norm argument
trainer = Trainer(
model=model,
optimizer=optimizer,
iterator=iterator,
train_dataset=dev_set,
validation_dataset=heldout_set,
patience=10,
num_epochs=num_epochs,
cuda_device=n_gpu,
grad_norm=5.0, # <<< NEW: clip global gradient norm
# grad_clipping=1.0 # optional per‑parameter clipping
)
Reference: training_util.py (sparse_clip_norm) and comments in rumour_dnn_trainer.py (lines 92‑96)
Switch to Built‑in LayerNorm to Eliminate NaNs
Replace the custom implementation with PyTorch’s stable version.
# Replace custom MyLayerNorm usage (see src/my_layer_norm.py) with:
from torch.nn import LayerNorm
# Example inside instantiate_rumour_model:
context_metadata_encoder = torch.nn.LSTM(CM_LSTM_INPUT_DIM,
CM_LSTM_HIDDEN_DIM,
num_layers=2,
batch_first=True)
# Apply LayerNorm after LSTM output:
ln = LayerNorm(CM_LSTM_HIDDEN_DIM)
def forward(...):
lstm_out, _ = context_metadata_encoder(...)
normed = ln(lstm_out)
# continue with normed tensor
Reference: my_layer_norm.py (line 71 comment) and the PyTorch forum discussion linked therein
Monitor GPU Usage and Memory Before Training
Prevent OOM errors by checking device availability.
import torch
print("🔧 CUDA available :", torch.cuda.is_available())
if torch.cuda.is_available():
print("🔧 GPU name :", torch.cuda.get_device_name(0))
print("🔧 Total memory (MiB):", torch.cuda.get_device_properties(0).total_memory // 2**20)
Reference: rumour_dnn_trainer.py already prints nvidia‑smi output (lines 99‑104)
Catch Mismatched Serialization Directories
Handle resume conflicts gracefully.
from src.training_util import create_serialization_dir
try:
create_serialization_dir(params, "my_experiment_dir", recover=False, force=False)
except Exception as e:
print("❌ Serialization error:", e)
# Inspect the offending key:
# e.args[0] contains the mismatched config name.
Reference: training_util.py (lines 165‑227)
Key Files in the RP-DNN Repository
Understanding the codebase layout accelerates debugging by mapping errors to their source locations:
src/rumour_dnn_trainer.py– CLI entry point that parses arguments, validates files, and launchesmodel_training. Lines 109‑116 handle CSV validation, while lines 136‑141 configure ELMo paths.src/allennlp_rumor_classifier.py– Model definition, feature-setting logic, and themodel_trainingroutine. Contains enum definitions forFEATURE_SETTING_OPTION_*andATTENTION_OPTION_*at the top of the file, with validation at line 376.src/data_loader.py– Loads the PHEME social-context corpus, resolves symlinks viaload_abs_path, and yields JSON tweets. Theload_tweets_context_dataset_dirfunction expects numeric folder names and raises exceptions at line 188 if source tweets are missing.src/training_util.py– Helper utilities includingsparse_clip_normfor gradient clipping andcreate_serialization_dir(lines 165‑227) for handling experiment resumes.src/context_features_extractor.py– Extracts numeric and textual context features based on the chosenfeature_setting, with validation logic at line 92.src/my_layer_norm.py– Custom LayerNorm implementation (source of NaN bugs) with a documented workaround comment at line 71.src/preprocessing/CredbankProcessor.py– Pre-processes CredBank text files and contains guards for'nan'string literals.requirements.txt– Pins external libraries (AllenNLP, PyTorch, pandas) whose version changes often cause hidden crashes.
Summary
- Validate inputs early by checking CSV paths and ELMo weight files in
src/rumour_dnn_trainer.pybefore the training loop starts. - Use enum values for
feature_settingandattention_optionas defined insrc/allennlp_rumor_classifier.pyto avoid configuration errors. - Monitor GPU resources via
nvidia-smiand reducetrain_batch_sizeormax_cxt_size_optionif you encounter CUDA OOM errors. - Stabilize training by enabling gradient clipping (
grad_norm=5.0) in theTrainerconstructor and replacing the customMyLayerNormwithtorch.nn.LayerNormto prevent NaN gradients. - Handle resumes carefully by either deleting old serialization directories or using the
--recoverflag with identical config files to avoidConfigurationError.
Frequently Asked Questions
What causes "source tweet not found" errors during RP-DNN training?
This error originates in src/data_loader.py (line 188) when the load_tweets_context_dataset_dir function cannot locate a tweet ID in the PHEME social-context directory. It typically occurs when the symlink data/social_context/aug-rnr-annotated-threads-retweets points to the wrong location or when stray non-numeric files exist in the corpus root. Verify that only numeric tweet ID folders are present and that the path resolves correctly via load_abs_path.
How do I fix NaN loss during RP-DNN training?
NaN loss usually results from gradient explosion or the custom layer normalization implementation in src/my_layer_norm.py (line 71). First, enable gradient clipping by passing grad_norm=5.0 to the AllenNLP Trainer constructor in src/rumour_dnn_trainer.py. If NaNs persist, replace the custom MyLayerNorm with torch.nn.LayerNorm in src/allennlp_rumor_classifier.py, as the built-in implementation handles variance calculation edge cases more robustly.
Why does RP-DNN crash with a serialization directory error?
The ConfigurationError stating “Serialization directory already exists” is raised by create_serialization_dir in src/training_util.py (lines 165‑227) when a previous run’s output folder conflicts with a new training invocation. This protects against accidental overwrites. To resolve, either delete the existing serialization_dir manually, or resume training using the --recover flag with configuration files identical to the original run.
How do I resolve ELMo weight loading failures in RP-DNN?
ELMo loading fails when rumour_dnn_trainer.py (lines 136‑141) cannot locate the HDF5 weight file at resource/embedding/elmo_model/elmo_credbank_2x4096_512_2048cnn_2xhighway_weights_10052019.hdf5. This typically occurs when symlinks are broken or the file was not downloaded. Confirm the file exists and that the symlink in resource/embedding/ resolves correctly using the load_abs_path helper, which handles both Unix symlinks and Windows shortcuts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →