# Required Preprocessing Steps for Credbank and PHEME Datasets in RPDNN

> Learn the essential preprocessing steps for Credbank and PHEME datasets in RPDNN. Discover normalization, deduplication, and balancing techniques for cleaner data.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: how-to-guide
- Published: 2026-03-04

---

**The RPDNN repository normalizes both Credbank and PHEME datasets through lower-casing, URL/mention removal, de-accenting, and tokenization using the shared `preprocessing_tweet_text` function, with Credbank processed from CSV batches into deduplicated text corpora and PHEME processed from JSON threads into balanced, labeled CSV files.**

The `jerrygaolondon/rpdnn` repository implements a dual-track preprocessing pipeline to prepare raw Twitter archives for rumor detection. Understanding the required preprocessing steps for Credbank and PHEME datasets is essential before feeding data into the ELMo fine-tuning stage or the RPDNN classification model.

## Common Text Normalization Routine

Both datasets rely on a single text-cleaning implementation to ensure consistent embedding inputs.

### The `preprocessing_tweet_text` Function

Implemented in [`src/preprocessing/CredbankProcessor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/CredbankProcessor.py) (lines 70-84), this function performs an eight-step normalization sequence on every tweet string:

- **Type validation** – Raises `ValueError` if the input is not a `str` instance.
- **Lower-casing** – Converts the entire string via `norm_tweet.lower()`.
- **URL removal** – Strips standard links using `re.sub(r'http\S+', '', norm_tweet)`.
- **Picture URL removal** – Eliminates `pic.twitter.com\S+` patterns.
- **Mention stripping** – Removes `@username` references via `re.sub(r"(?:\@|https?\://)\S+", "", norm_tweet)`.
- **De-accenting** – Normalizes accented characters using `gensim.utils.deaccent`.
- **Tokenization** – Splits text into tokens with NLTK's `TweetTokenizer()`.
- **Length filtering** – Discards tweets containing fewer than 4 tokens or fewer than 2 unique tokens.

The function returns a `List[str]` of clean tokens ready for vectorization.

## CredBank Preprocessing Workflow

The `CredbankProcessor` class in [`src/preprocessing/CredbankProcessor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/CredbankProcessor.py) handles large-scale CSV archives through a sequential pipeline designed for language model corpora generation.

### 1. File Discovery and Loading

The `load_all_files_path(dataset_dir)` utility scans the target directory for files containing the substring "csv" (lines 89-95). Each discovered file is processed by `load_tweets_from_credbank_csv(file)`, which reads tab-separated lines and extracts column 8 (the tweet text field).

### 2. Text Cleaning and Deduplication

Every extracted tweet passes through `preprocessing_tweet_text`. To prevent data leakage, the processor accumulates tweets into a Python `set` named `dedup_tweet_corpus_batch_i` (lines 48-52), removing duplicate sentences before persistence.

### 3. Export and Splitting

Cleaned, deduplicated tweets are written line-by-line to [`credbank_dataset_corpus_v2.txt`](https://github.com/jerrygaolondon/rpdnn/blob/main/credbank_dataset_corpus_v2.txt) (lines 40-68). For perplexity evaluation, the optional `generate_train_held_out_set` function employs `sklearn.model_selection.ShuffleSplit` to reserve a held-out subset (lines 85-108).

## PHEME Dataset Preprocessing Pipeline

The [`pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/pheme_data_processor.py) script converts threaded JSON conversations into structured, balanced classification datasets.

### 1. JSON Thread Traversal

The processor targets the `Data/all-rnr-annotated-threads` directory structure, walking through event folders and their `rumours` or `non-rumours` subdirectories (lines 24-27). For each thread, it loads only the source tweet from `source_tweet[0]` (lines 60-66), ignoring reply chains.

### 2. Metadata Extraction and Labeling

The script parses `created_at` timestamps using `pd.DatetimeIndex`, assigns binary labels (`1` for rumours, `0` for non-rumours), and extracts `user_id` and `user_name` metadata. Raw text is retrieved from either `full_text` or `text` JSON fields.

### 3. Normalization and CSV Generation

Extracted text passes through the identical `preprocessing_tweet_text` routine imported from `CredbankProcessor` (lines 10-12). Results are stored in event-specific CSV files with columns: `tweet_id, created_at, text, label, user_id, user_name` (lines 78-82).

### 4. Aggregation and Class Balancing

The `generate_combined_dev_set` function (lines 29-33) concatenates per-event CSVs and applies optional undersampling via `undersampling_neg` to balance negative examples. `ShuffleSplit` generates train/validation folds with 10% held-out, exporting final balanced datasets to `data/test-balance/<event>/` (lines 84-92).

## Implementation Examples

### Processing CredBank for ELMo Fine-tuning

```python
from src.preprocessing.CredbankProcessor import export_credbank_trainset

# Directory containing raw CredBank CSV batches

dataset_dir = "/path/to/credbank/tweets_corpus"
export_credbank_trainset(dataset_dir)

# Generates: credbank_dataset_corpus_v2.txt in the same folder

```

### Generating PHEME Development Sets

```python
from src.preprocessing.pheme_data_processor import generate_development_set

# Root of downloaded PHEME "all-rnr-annotated-threads" folder

pheme_root = "/path/to/Data/all-rnr-annotated-threads"
generate_development_set(pheme_root, output_dir="aug_rnr_training")

# Produces: <event>.csv files under aug_rnr_training/

```

### Creating Balanced Train/Validation Splits

```python
from src.preprocessing.pheme_data_processor import generate_combined_dev_set

train_dir = "/path/to/pheme_training"    # per-event CSVs from previous step

test_dir = "/path/to/pheme_test"         # optional separate test data

generate_combined_dev_set(train_dir, test_dir)

# Writes balanced CSVs to data/test-balance/<event>/

```

## Summary

- Both datasets utilize the **shared `preprocessing_tweet_text` function** in [`src/preprocessing/CredbankProcessor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/CredbankProcessor.py) for identical text cleaning, including lower-casing, URL removal, de-accenting, and token filtering.
- **CredBank** processing emphasizes CSV batch loading, column 8 extraction, set-based deduplication, and plain-text corpus export for language model pre-training.
- **PHEME** processing focuses on JSON thread traversal, source tweet extraction, metadata preservation, binary label assignment, and class balancing through optional undersampling before CSV export.
- The pipelines diverge in data handling—CredBank prioritizes deduplication for monolithic corpus generation, while PHEME implements event-based aggregation with train/validation splitting for supervised classification.

## Frequently Asked Questions

### What is the minimum token requirement for tweets in the RPDNN preprocessing pipeline?

The `preprocessing_tweet_text` function enforces two strict filtering thresholds: tweets must contain at least **4 tokens** and at least **2 unique tokens**. Content failing either criterion is discarded to eliminate extremely short or repetitive noise from the training data.

### How does the CredbankProcessor handle duplicate tweets across batches?

Before writing the final output file, the processor accumulates all tweets from a batch into a Python `set` named `dedup_tweet_corpus_batch_i` (lines 48-52 in [`src/preprocessing/CredbankProcessor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/CredbankProcessor.py)). This set operation automatically removes identical sentences, ensuring the exported [`credbank_dataset_corpus_v2.txt`](https://github.com/jerrygaolondon/rpdnn/blob/main/credbank_dataset_corpus_v2.txt) contains only unique instances.

### Why does the PHEME pipeline process only source tweets instead of full conversation threads?

The `generate_development_set` function specifically loads `source_tweet[0]` from each thread directory (lines 60-66 in [`pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/pheme_data_processor.py)), extracting only the initial claim that started the rumor. This design focuses the RPDNN model on detecting the root veracity of claims rather than classifying the sentiment or stance of subsequent replies in the conversation tree.

### Can the preprocessing scripts handle class imbalance in the PHEME dataset?

Yes. The `generate_combined_dev_set` function includes an `undersampling_neg` parameter that reduces the majority class (non-rumors) to match the minority class count. This optional balancing step occurs during CSV aggregation (lines 29-33) before the ShuffleSplit operation creates the final training and validation folds.