How RPDNN Handles Class Imbalance in Rumor Detection Datasets

RPDNN mitigates class imbalance by preprocessing the PHEME dataset through a hybrid of optional oversampling for the minority class and mandatory undersampling for the majority class, ensuring balanced training data before the neural network ever sees a single batch.

The jerrygaolondon/rpdnn repository implements a deep neural network for rumor detection that addresses the typical 4:1 imbalance between rumor and non-rumor tweets through data-level resampling rather than algorithm-level adjustments. By handling class imbalance in rumor detection datasets during the preprocessing phase, the model trains on an even distribution without requiring weighted loss functions or specialized sampling layers.

Hybrid Resampling Strategy in PHEME Data Processor

The core balancing logic resides in src/preprocessing/pheme_data_processor.py, which implements a four-stage pipeline to equalize class distributions before training.

Separating Positive and Negative Instances

The helper function separate_pos_neg splits the raw dataset into two distinct NumPy arrays: pos_samples containing rumor tweets and neg_samples containing non-rumor tweets. This separation occurs at line 63 and enables independent manipulation of each class population.

Optional Oversampling of Rumor Tweets

When the positive class shortage exceeds acceptable thresholds, oversampling_pos (implemented at line 98) randomly repeats positive examples until the count matches the calculated difference (train_diffs). This step executes only before the train/validation split, preserving the integrity of the validation set by excluding synthetic duplicates.

Mandatory Undersampling of Non-Rumor Tweets

The undersampling_neg function at line 30 randomly discards a computed number of negative examples to bring the majority class into parity with the minority class. Unlike oversampling, this undersampling applies consistently across training, validation, and test splits to maintain balanced evaluation metrics throughout the entire pipeline.

Verifying Dataset Balance

After resampling, check_dataset_balance (line 43) prints the final counts of positive and negative instances, providing developers with explicit confirmation that the class distributions have been successfully equalized.

Preprocessing Pipeline Implementation

The following code demonstrates how to apply these resampling utilities to a raw PHEME CSV file before feeding the data into the trainer.

from src.preprocessing.pheme_data_processor import (
    load_matrix_from_csv,
    undersampling_neg,
    oversampling_pos,
    separate_pos_neg,
    check_dataset_balance,
)

# Load raw source-tweet matrix (columns: …, label)

raw_path = "data/train/pheme_6392078_train_set_combined.csv"
raw_X = load_matrix_from_csv(raw_path, header=0, start_col_index=0, end_col_index=4)

# 1) split into positive / negative

pos, neg = separate_pos_neg(raw_X)

# 2) compute how many positives we need to add / negatives to drop

train_diffs = len(neg) - len(pos)      # positive shortage

# 3) optionally oversample positives (uncomment if you want)

# raw_X = oversampling_pos(raw_X, shuffling_options, raw_X, train_diffs)

# 4) undersample negatives to obtain balance

balanced_X = undersampling_neg(raw_X)

# 5) verify balance

check_dataset_balance(balanced_X)

Integration with the Rumor Detection Trainer

Once balanced, the matrices move directly into src/rumour_dnn_trainer.py without additional class-weight configurations. The trainer consumes the balanced datasets through standard AllenNLP data loaders, as shown in the following implementation pattern.

from src.rumour_dnn_trainer import RumourDNNTrainer
from src.training_util import create_serialization_dir

# Assume `balanced_X` is a NumPy array (features + label column)

trainer = RumourDNNTrainer(
    train_data_path="balanced_train.npy",
    validation_data_path="balanced_val.npy",
    test_data_path="balanced_test.npy",
    # other hyper-parameters …

)

serialization_dir = "output/rumor_exp"
create_serialization_dir(params=trainer.params,
                         serialization_dir=serialization_dir,
                         recover=False,
                         force=True)

trainer.train()

Why Pre-Balancing Eliminates Loss-Function Workarounds

Because pheme_data_processor.py enforces balance before training begins, the downstream DNN defined in rumour_dnn_trainer.py requires no special loss weighting or class-weight arguments. While src/training_util.py contains weighted loss handling logic (lines 295-370) for alternative workflows, the RPDNN approach renders these utilities unnecessary by ensuring the model learns from an even distribution of rumors and non-rumors from the first epoch.

Summary

  • Data-level resourcing: RPDNN handles class imbalance through preprocessing in src/preprocessing/pheme_data_processor.py rather than algorithmic compensation during training.
  • Hybrid approach: The pipeline combines optional oversampling (oversampling_pos, line 98) with mandatory undersampling (undersampling_neg, line 30) to achieve parity.
  • Validation integrity: Oversampling applies only to training data, while undersampling affects all splits (train, validation, test) uniformly.
  • Zero configuration: Pre-balanced inputs eliminate the need for class weights or specialized loss functions in src/rumour_dnn_trainer.py.

Frequently Asked Questions

Does RPDNN use class weights in the loss function to handle imbalance?

No. According to the source code in jerrygaolondon/rpdnn, the model relies entirely on preprocessing-based balancing. Because pheme_data_processor.py delivers already-equalized datasets to the trainer, the AllenNLP training loop in rumour_dnn_trainer.py uses standard unweighted loss calculations.

Where does the optional oversampling occur in the codebase?

The optional oversampling logic resides in src/preprocessing/pheme_data_processor.py at line 98 within the oversampling_pos function. This utility randomly duplicates positive examples until the minority class count matches the specified difference, but developers must explicitly invoke it before the train/validation split occurs.

Why undersample the majority class instead of only oversampling the minority?

The undersampling_neg function (line 30) provides a deterministic path to balance by reducing the majority class size to match the minority. This approach prevents the data bloat and potential overfitting risks associated with excessive duplication of rumor tweets, while still offering oversampling as an optional supplementary step when dataset size constraints permit.

How can I verify that my resampling succeeded?

Call check_dataset_balance from src/preprocessing/pheme_data_processor.py (line 43) after running the undersampling or oversampling functions. This utility prints the final positive and negative instance counts, allowing immediate visual confirmation that the 4:1 imbalance has been neutralized before model training commences.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →