# Common Failure Modes and Debugging Strategies for RP-DNN Training

> Debug RP-DNN training failures like missing files, OOM errors, and NaN gradients. Learn common failure modes and effective debugging strategies for your RP-DNN models with this guide.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: debugging-strategies
- Published: 2026-03-04

---

**The most common RP-DNN training failures stem from missing dataset files, invalid feature settings, ELMo weight loading errors, GPU out-of-memory issues, and numerical instabilities like NaN gradients, all of which can be diagnosed by verifying file paths in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py), checking enum values in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py), and enabling gradient clipping in [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py).**

RP-DNN (Rumour-Propagation Deep Neural Network) is an AllenNLP-based deep learning framework for rumor detection that combines ELMo embeddings, LSTM or Transformer encoders, and handcrafted social-context features. Because the training pipeline integrates external resources like ELMo weights, the PHEME social-context corpus, and CSV-based tweet datasets, RP-DNN training failures often occur at the intersection of data validation, configuration parsing, and GPU memory management.

## Common RP-DNN Training Failure Modes

### Missing or Malformed Dataset Files

Before the model is instantiated, [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py) (lines 109‑116) attempts to load CSV files specified by the command-line arguments `-t/--trainset`, `--heldout`, and `-e/--evaluationset`. If these paths are invalid, [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) (lines 65‑77) raises a `FileNotFoundError`.

**Symptom:** Immediate crash with `FileNotFoundError` before any GPU allocation occurs.

**Debugging strategy:** Verify paths programmatically using `os.path.isfile` or the helper `load_abs_path` to resolve symlinks in the data directory. Ensure the CSV files exist before invoking `model_training()`.

### Invalid Feature Setting or Attention Options

The CLI arguments `--feature_setting` and `--attention_option` must match enums defined at the top of [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py). Validation logic in [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) (lines 118‑124) raises `ValueError` with messages like “Supported training (feature_setting option) …” when unsupported values are provided.

**Symptom:** `ValueError` during argument parsing before training begins.

**Debugging strategy:** Restrict values to documented constants such as `FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY = 1` or `ATTENTION_OPTION_HIERARCHICAL = 1`. Print the constant list before calling `model_training()` to verify compatibility.

### ELMo Model Loading Failures

RP-DNN relies on ELMo embeddings loaded via `ElmoTokenEmbedder`. The weight path is configured in [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) (lines 136‑141), pointing to `resource/embedding/elmo_model/elmo_credbank_2x4096_512_2048cnn_2xhighway_weights_10052019.hdf5`.

**Symptom:** Runtime error when the embedder attempts to load the weight file, often referencing missing HDF5 files.

**Debugging strategy:** Confirm the HDF5 file exists and that symlinks in `resource/embedding/` resolve correctly. The `load_abs_path` helper handles Windows shortcuts and relative paths.

### GPU Device Mismatch and Out-of-Memory Errors

GPU configuration occurs via `config_gpu_use` (called at line 42 of [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py)), and the model is moved to the device in `instantiate_rumour_model` (lines 63‑66). Out-of-memory errors occur when `train_batch_size` (default 128) exceeds available VRAM.

**Symptom:** “CUDA out of memory” or silent CPU fallback causing extremely slow training throughput.

**Debugging strategy:** Run with `-g 0` for GPU 0 or `-g -1` for CPU-only mode. Monitor `nvidia-smi` output (the script prints this at lines 99‑104). If OOM persists, lower `train_batch_size` or reduce `max_cxt_size_option` to limit context window memory usage.

### NaN Gradients and Loss Explosion

Numerical instability manifests as `nan` loss or `RuntimeError` about inconsistent loss production. Gradient clipping utilities reside in [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py) (`sparse_clip_norm`), with implementation hints in [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) (lines 92‑96).

**Symptom:** Training halts with `nan` loss or gradient explosion errors after several batches.

**Debugging strategy:** Enable gradient clipping by passing `grad_norm=5.0` to the `Trainer` constructor, or set `grad_clipping=1.0` for per-parameter clipping. Uncomment the relevant block in [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) and experiment with thresholds between 1.0 and 5.0.

### Layer-Norm Numerical Instability

A custom LayerNorm implementation in [`src/my_layer_norm.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/my_layer_norm.py) (see comment around line 71) can produce NaNs after a few epochs, particularly with stacked LSTM encoders.

**Symptom:** NaNs appear mid-training despite stable initial epochs, often affecting deep encoder stacks.

**Debugging strategy:** Replace the custom `MyLayerNorm` with PyTorch’s built-in `torch.nn.LayerNorm`. Update `instantiate_rumour_model` in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) to use the standard implementation, which handles edge cases in variance calculation more robustly.

### Inconsistent Social-Context Directory Structure

The PHEME corpus loader in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) (`load_tweets_context_dataset_dir`) expects a strict layout where each rumor thread is a numeric folder containing `source-tweets` and `reactions` subdirectories. The global `social_context_data_dir` variable (line 43) must point to the unpacked corpus.

**Symptom:** `Exception: source tweet (id :…) is not found in current social context dataset!`

**Debugging strategy:** Verify that `data/social_context/aug-rnr-annotated-threads-retweets` symlinks to the unpacked PHEME corpus. Ensure only numeric tweet ID folders exist at the root—any stray files break the mapping logic in `load_tweets_context_dataset_dir`.

### Serialization Directory Conflicts on Resume

When resuming training, `training_util.create_serialization_dir` (lines 165‑227) checks for existing directories and config mismatches.

**Symptom:** `ConfigurationError` stating “Serialization directory already exists”.

**Debugging strategy:** Either delete the old `serialization_dir` before restarting, or invoke the script with `--recover` alongside identical configuration files. The helper validates key consistency and aborts if the archived config differs from the current run parameters.

## Debugging Strategies and Code Solutions

### Verify All Input Paths Before Launching Training

Prevent early crashes by validating CSV and model weight paths programmatically.

```python
from pathlib import Path
import sys

def assert_file(p: Path, name: str) -> None:
    if not p.is_file():
        sys.exit(f"[ERROR] {name} not found: {p}")

train = Path("/data/train/bostonbombings/aug_rnr_train_set_combined.csv")
heldout = Path("/data/train/bostonbombings/aug_rnr_heldout_set_combined.csv")
test = Path("/data/test/bostonbombings.csv")

for p, n in [(train, "train set"), (heldout, "held‑out set"), (test, "test set")]:
    assert_file(p, n)

print("✅ All CSVs exist – you can safely call model_training()")

```

*Reference:* [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) (lines 109‑116)

### Enable Gradient Clipping to Avoid NaNs

Stabilize training by clipping gradients before the optimizer step.

```python

# Inside rumour_dnn_trainer.py, before Trainer creation:

optimizer = optim.Adam(model.parameters(), lr=1e-4, weight_decay=1e-5)

# Add clipping via the grad_norm argument

trainer = Trainer(
    model=model,
    optimizer=optimizer,
    iterator=iterator,
    train_dataset=dev_set,
    validation_dataset=heldout_set,
    patience=10,
    num_epochs=num_epochs,
    cuda_device=n_gpu,
    grad_norm=5.0,          # <<< NEW: clip global gradient norm

    # grad_clipping=1.0    # optional per‑parameter clipping

)

```

*Reference:* [`training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/training_util.py) (`sparse_clip_norm`) and comments in [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) (lines 92‑96)

### Switch to Built‑in LayerNorm to Eliminate NaNs

Replace the custom implementation with PyTorch’s stable version.

```python

# Replace custom MyLayerNorm usage (see src/my_layer_norm.py) with:

from torch.nn import LayerNorm

# Example inside instantiate_rumour_model:

context_metadata_encoder = torch.nn.LSTM(CM_LSTM_INPUT_DIM,
                                        CM_LSTM_HIDDEN_DIM,
                                        num_layers=2,
                                        batch_first=True)

# Apply LayerNorm after LSTM output:

ln = LayerNorm(CM_LSTM_HIDDEN_DIM)

def forward(...):
    lstm_out, _ = context_metadata_encoder(...)
    normed = ln(lstm_out)
    # continue with normed tensor

```

*Reference:* [`my_layer_norm.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/my_layer_norm.py) (line 71 comment) and the PyTorch forum discussion linked therein

### Monitor GPU Usage and Memory Before Training

Prevent OOM errors by checking device availability.

```python
import torch

print("🔧 CUDA available :", torch.cuda.is_available())
if torch.cuda.is_available():
    print("🔧 GPU name :", torch.cuda.get_device_name(0))
    print("🔧 Total memory (MiB):", torch.cuda.get_device_properties(0).total_memory // 2**20)

```

*Reference:* [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) already prints `nvidia‑smi` output (lines 99‑104)

### Catch Mismatched Serialization Directories

Handle resume conflicts gracefully.

```python
from src.training_util import create_serialization_dir

try:
    create_serialization_dir(params, "my_experiment_dir", recover=False, force=False)
except Exception as e:
    print("❌ Serialization error:", e)
    # Inspect the offending key:

    #   e.args[0] contains the mismatched config name.

```

*Reference:* [`training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/training_util.py) (lines 165‑227)

## Key Files in the RP-DNN Repository

Understanding the codebase layout accelerates debugging by mapping errors to their source locations:

- **[`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py)** – CLI entry point that parses arguments, validates files, and launches `model_training`. Lines 109‑116 handle CSV validation, while lines 136‑141 configure ELMo paths.
- **[`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py)** – Model definition, feature-setting logic, and the `model_training` routine. Contains enum definitions for `FEATURE_SETTING_OPTION_*` and `ATTENTION_OPTION_*` at the top of the file, with validation at line 376.
- **[`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py)** – Loads the PHEME social-context corpus, resolves symlinks via `load_abs_path`, and yields JSON tweets. The `load_tweets_context_dataset_dir` function expects numeric folder names and raises exceptions at line 188 if source tweets are missing.
- **[`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py)** – Helper utilities including `sparse_clip_norm` for gradient clipping and `create_serialization_dir` (lines 165‑227) for handling experiment resumes.
- **[`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py)** – Extracts numeric and textual context features based on the chosen `feature_setting`, with validation logic at line 92.
- **[`src/my_layer_norm.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/my_layer_norm.py)** – Custom LayerNorm implementation (source of NaN bugs) with a documented workaround comment at line 71.
- **[`src/preprocessing/CredbankProcessor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/CredbankProcessor.py)** – Pre-processes CredBank text files and contains guards for `'nan'` string literals.
- **[`requirements.txt`](https://github.com/jerrygaolondon/rpdnn/blob/main/requirements.txt)** – Pins external libraries (AllenNLP, PyTorch, pandas) whose version changes often cause hidden crashes.

## Summary

- **Validate inputs early** by checking CSV paths and ELMo weight files in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py) before the training loop starts.
- **Use enum values** for `feature_setting` and `attention_option` as defined in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) to avoid configuration errors.
- **Monitor GPU resources** via `nvidia-smi` and reduce `train_batch_size` or `max_cxt_size_option` if you encounter CUDA OOM errors.
- **Stabilize training** by enabling gradient clipping (`grad_norm=5.0`) in the `Trainer` constructor and replacing the custom `MyLayerNorm` with `torch.nn.LayerNorm` to prevent NaN gradients.
- **Handle resumes carefully** by either deleting old serialization directories or using the `--recover` flag with identical config files to avoid `ConfigurationError`.

## Frequently Asked Questions

### What causes "source tweet not found" errors during RP-DNN training?

This error originates in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) (line 188) when the `load_tweets_context_dataset_dir` function cannot locate a tweet ID in the PHEME social-context directory. It typically occurs when the symlink `data/social_context/aug-rnr-annotated-threads-retweets` points to the wrong location or when stray non-numeric files exist in the corpus root. Verify that only numeric tweet ID folders are present and that the path resolves correctly via `load_abs_path`.

### How do I fix NaN loss during RP-DNN training?

NaN loss usually results from gradient explosion or the custom layer normalization implementation in [`src/my_layer_norm.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/my_layer_norm.py) (line 71). First, enable gradient clipping by passing `grad_norm=5.0` to the AllenNLP `Trainer` constructor in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py). If NaNs persist, replace the custom `MyLayerNorm` with `torch.nn.LayerNorm` in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py), as the built-in implementation handles variance calculation edge cases more robustly.

### Why does RP-DNN crash with a serialization directory error?

The `ConfigurationError` stating “Serialization directory already exists” is raised by `create_serialization_dir` in [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py) (lines 165‑227) when a previous run’s output folder conflicts with a new training invocation. This protects against accidental overwrites. To resolve, either delete the existing `serialization_dir` manually, or resume training using the `--recover` flag with configuration files identical to the original run.

### How do I resolve ELMo weight loading failures in RP-DNN?

ELMo loading fails when [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) (lines 136‑141) cannot locate the HDF5 weight file at `resource/embedding/elmo_model/elmo_credbank_2x4096_512_2048cnn_2xhighway_weights_10052019.hdf5`. This typically occurs when symlinks are broken or the file was not downloaded. Confirm the file exists and that the symlink in `resource/embedding/` resolves correctly using the `load_abs_path` helper, which handles both Unix symlinks and Windows shortcuts.