How RPDNN Handles Class Imbalance in Rumor Detection Datasets
RPDNN mitigates class imbalance by preprocessing the PHEME dataset through a hybrid of optional oversampling for the minority class and mandatory undersampling for the majority class, ensuring balanced training data before the neural network ever sees a single batch.
The jerrygaolondon/rpdnn repository implements a deep neural network for rumor detection that addresses the typical 4:1 imbalance between rumor and non-rumor tweets through data-level resampling rather than algorithm-level adjustments. By handling class imbalance in rumor detection datasets during the preprocessing phase, the model trains on an even distribution without requiring weighted loss functions or specialized sampling layers.
Hybrid Resampling Strategy in PHEME Data Processor
The core balancing logic resides in src/preprocessing/pheme_data_processor.py, which implements a four-stage pipeline to equalize class distributions before training.
Separating Positive and Negative Instances
The helper function separate_pos_neg splits the raw dataset into two distinct NumPy arrays: pos_samples containing rumor tweets and neg_samples containing non-rumor tweets. This separation occurs at line 63 and enables independent manipulation of each class population.
Optional Oversampling of Rumor Tweets
When the positive class shortage exceeds acceptable thresholds, oversampling_pos (implemented at line 98) randomly repeats positive examples until the count matches the calculated difference (train_diffs). This step executes only before the train/validation split, preserving the integrity of the validation set by excluding synthetic duplicates.
Mandatory Undersampling of Non-Rumor Tweets
The undersampling_neg function at line 30 randomly discards a computed number of negative examples to bring the majority class into parity with the minority class. Unlike oversampling, this undersampling applies consistently across training, validation, and test splits to maintain balanced evaluation metrics throughout the entire pipeline.
Verifying Dataset Balance
After resampling, check_dataset_balance (line 43) prints the final counts of positive and negative instances, providing developers with explicit confirmation that the class distributions have been successfully equalized.
Preprocessing Pipeline Implementation
The following code demonstrates how to apply these resampling utilities to a raw PHEME CSV file before feeding the data into the trainer.
from src.preprocessing.pheme_data_processor import (
load_matrix_from_csv,
undersampling_neg,
oversampling_pos,
separate_pos_neg,
check_dataset_balance,
)
# Load raw source-tweet matrix (columns: …, label)
raw_path = "data/train/pheme_6392078_train_set_combined.csv"
raw_X = load_matrix_from_csv(raw_path, header=0, start_col_index=0, end_col_index=4)
# 1) split into positive / negative
pos, neg = separate_pos_neg(raw_X)
# 2) compute how many positives we need to add / negatives to drop
train_diffs = len(neg) - len(pos) # positive shortage
# 3) optionally oversample positives (uncomment if you want)
# raw_X = oversampling_pos(raw_X, shuffling_options, raw_X, train_diffs)
# 4) undersample negatives to obtain balance
balanced_X = undersampling_neg(raw_X)
# 5) verify balance
check_dataset_balance(balanced_X)
Integration with the Rumor Detection Trainer
Once balanced, the matrices move directly into src/rumour_dnn_trainer.py without additional class-weight configurations. The trainer consumes the balanced datasets through standard AllenNLP data loaders, as shown in the following implementation pattern.
from src.rumour_dnn_trainer import RumourDNNTrainer
from src.training_util import create_serialization_dir
# Assume `balanced_X` is a NumPy array (features + label column)
trainer = RumourDNNTrainer(
train_data_path="balanced_train.npy",
validation_data_path="balanced_val.npy",
test_data_path="balanced_test.npy",
# other hyper-parameters …
)
serialization_dir = "output/rumor_exp"
create_serialization_dir(params=trainer.params,
serialization_dir=serialization_dir,
recover=False,
force=True)
trainer.train()
Why Pre-Balancing Eliminates Loss-Function Workarounds
Because pheme_data_processor.py enforces balance before training begins, the downstream DNN defined in rumour_dnn_trainer.py requires no special loss weighting or class-weight arguments. While src/training_util.py contains weighted loss handling logic (lines 295-370) for alternative workflows, the RPDNN approach renders these utilities unnecessary by ensuring the model learns from an even distribution of rumors and non-rumors from the first epoch.
Summary
- Data-level resourcing: RPDNN handles class imbalance through preprocessing in
src/preprocessing/pheme_data_processor.pyrather than algorithmic compensation during training. - Hybrid approach: The pipeline combines optional oversampling (
oversampling_pos, line 98) with mandatory undersampling (undersampling_neg, line 30) to achieve parity. - Validation integrity: Oversampling applies only to training data, while undersampling affects all splits (train, validation, test) uniformly.
- Zero configuration: Pre-balanced inputs eliminate the need for class weights or specialized loss functions in
src/rumour_dnn_trainer.py.
Frequently Asked Questions
Does RPDNN use class weights in the loss function to handle imbalance?
No. According to the source code in jerrygaolondon/rpdnn, the model relies entirely on preprocessing-based balancing. Because pheme_data_processor.py delivers already-equalized datasets to the trainer, the AllenNLP training loop in rumour_dnn_trainer.py uses standard unweighted loss calculations.
Where does the optional oversampling occur in the codebase?
The optional oversampling logic resides in src/preprocessing/pheme_data_processor.py at line 98 within the oversampling_pos function. This utility randomly duplicates positive examples until the minority class count matches the specified difference, but developers must explicitly invoke it before the train/validation split occurs.
Why undersample the majority class instead of only oversampling the minority?
The undersampling_neg function (line 30) provides a deterministic path to balance by reducing the majority class size to match the minority. This approach prevents the data bloat and potential overfitting risks associated with excessive duplication of rumor tweets, while still offering oversampling as an optional supplementary step when dataset size constraints permit.
How can I verify that my resampling succeeded?
Call check_dataset_balance from src/preprocessing/pheme_data_processor.py (line 43) after running the undersampling or oversampling functions. This utility prints the final positive and negative instance counts, allowing immediate visual confirmation that the 4:1 imbalance has been neutralized before model training commences.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →