# How RPDNN Handles Class Imbalance in Rumor Detection Datasets

> Discover how RPDNN effectively handles class imbalance in rumor detection datasets using hybrid oversampling and undersampling techniques for balanced training data.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: deep-dive
- Published: 2026-03-04

---

**RPDNN mitigates class imbalance by preprocessing the PHEME dataset through a hybrid of optional oversampling for the minority class and mandatory undersampling for the majority class, ensuring balanced training data before the neural network ever sees a single batch.**

The `jerrygaolondon/rpdnn` repository implements a deep neural network for rumor detection that addresses the typical 4:1 imbalance between rumor and non-rumor tweets through data-level resampling rather than algorithm-level adjustments. By handling class imbalance in rumor detection datasets during the preprocessing phase, the model trains on an even distribution without requiring weighted loss functions or specialized sampling layers.

## Hybrid Resampling Strategy in PHEME Data Processor

The core balancing logic resides in [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py), which implements a four-stage pipeline to equalize class distributions before training.

### Separating Positive and Negative Instances

The helper function `separate_pos_neg` splits the raw dataset into two distinct NumPy arrays: `pos_samples` containing rumor tweets and `neg_samples` containing non-rumor tweets. This separation occurs at line 63 and enables independent manipulation of each class population.

### Optional Oversampling of Rumor Tweets

When the positive class shortage exceeds acceptable thresholds, `oversampling_pos` (implemented at line 98) randomly repeats positive examples until the count matches the calculated difference (`train_diffs`). This step executes only before the train/validation split, preserving the integrity of the validation set by excluding synthetic duplicates.

### Mandatory Undersampling of Non-Rumor Tweets

The `undersampling_neg` function at line 30 randomly discards a computed number of negative examples to bring the majority class into parity with the minority class. Unlike oversampling, this undersampling applies consistently across training, validation, and test splits to maintain balanced evaluation metrics throughout the entire pipeline.

### Verifying Dataset Balance

After resampling, `check_dataset_balance` (line 43) prints the final counts of positive and negative instances, providing developers with explicit confirmation that the class distributions have been successfully equalized.

## Preprocessing Pipeline Implementation

The following code demonstrates how to apply these resampling utilities to a raw PHEME CSV file before feeding the data into the trainer.

```python
from src.preprocessing.pheme_data_processor import (
    load_matrix_from_csv,
    undersampling_neg,
    oversampling_pos,
    separate_pos_neg,
    check_dataset_balance,
)

# Load raw source-tweet matrix (columns: …, label)

raw_path = "data/train/pheme_6392078_train_set_combined.csv"
raw_X = load_matrix_from_csv(raw_path, header=0, start_col_index=0, end_col_index=4)

# 1) split into positive / negative

pos, neg = separate_pos_neg(raw_X)

# 2) compute how many positives we need to add / negatives to drop

train_diffs = len(neg) - len(pos)      # positive shortage

# 3) optionally oversample positives (uncomment if you want)

# raw_X = oversampling_pos(raw_X, shuffling_options, raw_X, train_diffs)

# 4) undersample negatives to obtain balance

balanced_X = undersampling_neg(raw_X)

# 5) verify balance

check_dataset_balance(balanced_X)

```

## Integration with the Rumor Detection Trainer

Once balanced, the matrices move directly into [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py) without additional class-weight configurations. The trainer consumes the balanced datasets through standard AllenNLP data loaders, as shown in the following implementation pattern.

```python
from src.rumour_dnn_trainer import RumourDNNTrainer
from src.training_util import create_serialization_dir

# Assume `balanced_X` is a NumPy array (features + label column)

trainer = RumourDNNTrainer(
    train_data_path="balanced_train.npy",
    validation_data_path="balanced_val.npy",
    test_data_path="balanced_test.npy",
    # other hyper-parameters …

)

serialization_dir = "output/rumor_exp"
create_serialization_dir(params=trainer.params,
                         serialization_dir=serialization_dir,
                         recover=False,
                         force=True)

trainer.train()

```

## Why Pre-Balancing Eliminates Loss-Function Workarounds

Because [`pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/pheme_data_processor.py) enforces balance before training begins, the downstream DNN defined in [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) requires no special loss weighting or class-weight arguments. While [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py) contains weighted loss handling logic (lines 295-370) for alternative workflows, the RPDNN approach renders these utilities unnecessary by ensuring the model learns from an even distribution of rumors and non-rumors from the first epoch.

## Summary

- **Data-level resourcing**: RPDNN handles class imbalance through preprocessing in [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py) rather than algorithmic compensation during training.
- **Hybrid approach**: The pipeline combines optional oversampling (`oversampling_pos`, line 98) with mandatory undersampling (`undersampling_neg`, line 30) to achieve parity.
- **Validation integrity**: Oversampling applies only to training data, while undersampling affects all splits (train, validation, test) uniformly.
- **Zero configuration**: Pre-balanced inputs eliminate the need for class weights or specialized loss functions in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py).

## Frequently Asked Questions

### Does RPDNN use class weights in the loss function to handle imbalance?

No. According to the source code in `jerrygaolondon/rpdnn`, the model relies entirely on preprocessing-based balancing. Because [`pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/pheme_data_processor.py) delivers already-equalized datasets to the trainer, the AllenNLP training loop in [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) uses standard unweighted loss calculations.

### Where does the optional oversampling occur in the codebase?

The optional oversampling logic resides in [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py) at line 98 within the `oversampling_pos` function. This utility randomly duplicates positive examples until the minority class count matches the specified difference, but developers must explicitly invoke it before the train/validation split occurs.

### Why undersample the majority class instead of only oversampling the minority?

The `undersampling_neg` function (line 30) provides a deterministic path to balance by reducing the majority class size to match the minority. This approach prevents the data bloat and potential overfitting risks associated with excessive duplication of rumor tweets, while still offering oversampling as an optional supplementary step when dataset size constraints permit.

### How can I verify that my resampling succeeded?

Call `check_dataset_balance` from [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py) (line 43) after running the undersampling or oversampling functions. This utility prints the final positive and negative instance counts, allowing immediate visual confirmation that the 4:1 imbalance has been neutralized before model training commences.