Required Preprocessing Steps for Credbank and PHEME Datasets in RPDNN
The RPDNN repository normalizes both Credbank and PHEME datasets through lower-casing, URL/mention removal, de-accenting, and tokenization using the shared preprocessing_tweet_text function, with Credbank processed from CSV batches into deduplicated text corpora and PHEME processed from JSON threads into balanced, labeled CSV files.
The jerrygaolondon/rpdnn repository implements a dual-track preprocessing pipeline to prepare raw Twitter archives for rumor detection. Understanding the required preprocessing steps for Credbank and PHEME datasets is essential before feeding data into the ELMo fine-tuning stage or the RPDNN classification model.
Common Text Normalization Routine
Both datasets rely on a single text-cleaning implementation to ensure consistent embedding inputs.
The preprocessing_tweet_text Function
Implemented in src/preprocessing/CredbankProcessor.py (lines 70-84), this function performs an eight-step normalization sequence on every tweet string:
- Type validation – Raises
ValueErrorif the input is not astrinstance. - Lower-casing – Converts the entire string via
norm_tweet.lower(). - URL removal – Strips standard links using
re.sub(r'http\S+', '', norm_tweet). - Picture URL removal – Eliminates
pic.twitter.com\S+patterns. - Mention stripping – Removes
@usernamereferences viare.sub(r"(?:\@|https?\://)\S+", "", norm_tweet). - De-accenting – Normalizes accented characters using
gensim.utils.deaccent. - Tokenization – Splits text into tokens with NLTK's
TweetTokenizer(). - Length filtering – Discards tweets containing fewer than 4 tokens or fewer than 2 unique tokens.
The function returns a List[str] of clean tokens ready for vectorization.
CredBank Preprocessing Workflow
The CredbankProcessor class in src/preprocessing/CredbankProcessor.py handles large-scale CSV archives through a sequential pipeline designed for language model corpora generation.
1. File Discovery and Loading
The load_all_files_path(dataset_dir) utility scans the target directory for files containing the substring "csv" (lines 89-95). Each discovered file is processed by load_tweets_from_credbank_csv(file), which reads tab-separated lines and extracts column 8 (the tweet text field).
2. Text Cleaning and Deduplication
Every extracted tweet passes through preprocessing_tweet_text. To prevent data leakage, the processor accumulates tweets into a Python set named dedup_tweet_corpus_batch_i (lines 48-52), removing duplicate sentences before persistence.
3. Export and Splitting
Cleaned, deduplicated tweets are written line-by-line to credbank_dataset_corpus_v2.txt (lines 40-68). For perplexity evaluation, the optional generate_train_held_out_set function employs sklearn.model_selection.ShuffleSplit to reserve a held-out subset (lines 85-108).
PHEME Dataset Preprocessing Pipeline
The pheme_data_processor.py script converts threaded JSON conversations into structured, balanced classification datasets.
1. JSON Thread Traversal
The processor targets the Data/all-rnr-annotated-threads directory structure, walking through event folders and their rumours or non-rumours subdirectories (lines 24-27). For each thread, it loads only the source tweet from source_tweet[0] (lines 60-66), ignoring reply chains.
2. Metadata Extraction and Labeling
The script parses created_at timestamps using pd.DatetimeIndex, assigns binary labels (1 for rumours, 0 for non-rumours), and extracts user_id and user_name metadata. Raw text is retrieved from either full_text or text JSON fields.
3. Normalization and CSV Generation
Extracted text passes through the identical preprocessing_tweet_text routine imported from CredbankProcessor (lines 10-12). Results are stored in event-specific CSV files with columns: tweet_id, created_at, text, label, user_id, user_name (lines 78-82).
4. Aggregation and Class Balancing
The generate_combined_dev_set function (lines 29-33) concatenates per-event CSVs and applies optional undersampling via undersampling_neg to balance negative examples. ShuffleSplit generates train/validation folds with 10% held-out, exporting final balanced datasets to data/test-balance/<event>/ (lines 84-92).
Implementation Examples
Processing CredBank for ELMo Fine-tuning
from src.preprocessing.CredbankProcessor import export_credbank_trainset
# Directory containing raw CredBank CSV batches
dataset_dir = "/path/to/credbank/tweets_corpus"
export_credbank_trainset(dataset_dir)
# Generates: credbank_dataset_corpus_v2.txt in the same folder
Generating PHEME Development Sets
from src.preprocessing.pheme_data_processor import generate_development_set
# Root of downloaded PHEME "all-rnr-annotated-threads" folder
pheme_root = "/path/to/Data/all-rnr-annotated-threads"
generate_development_set(pheme_root, output_dir="aug_rnr_training")
# Produces: <event>.csv files under aug_rnr_training/
Creating Balanced Train/Validation Splits
from src.preprocessing.pheme_data_processor import generate_combined_dev_set
train_dir = "/path/to/pheme_training" # per-event CSVs from previous step
test_dir = "/path/to/pheme_test" # optional separate test data
generate_combined_dev_set(train_dir, test_dir)
# Writes balanced CSVs to data/test-balance/<event>/
Summary
- Both datasets utilize the shared
preprocessing_tweet_textfunction insrc/preprocessing/CredbankProcessor.pyfor identical text cleaning, including lower-casing, URL removal, de-accenting, and token filtering. - CredBank processing emphasizes CSV batch loading, column 8 extraction, set-based deduplication, and plain-text corpus export for language model pre-training.
- PHEME processing focuses on JSON thread traversal, source tweet extraction, metadata preservation, binary label assignment, and class balancing through optional undersampling before CSV export.
- The pipelines diverge in data handling—CredBank prioritizes deduplication for monolithic corpus generation, while PHEME implements event-based aggregation with train/validation splitting for supervised classification.
Frequently Asked Questions
What is the minimum token requirement for tweets in the RPDNN preprocessing pipeline?
The preprocessing_tweet_text function enforces two strict filtering thresholds: tweets must contain at least 4 tokens and at least 2 unique tokens. Content failing either criterion is discarded to eliminate extremely short or repetitive noise from the training data.
How does the CredbankProcessor handle duplicate tweets across batches?
Before writing the final output file, the processor accumulates all tweets from a batch into a Python set named dedup_tweet_corpus_batch_i (lines 48-52 in src/preprocessing/CredbankProcessor.py). This set operation automatically removes identical sentences, ensuring the exported credbank_dataset_corpus_v2.txt contains only unique instances.
Why does the PHEME pipeline process only source tweets instead of full conversation threads?
The generate_development_set function specifically loads source_tweet[0] from each thread directory (lines 60-66 in pheme_data_processor.py), extracting only the initial claim that started the rumor. This design focuses the RPDNN model on detecting the root veracity of claims rather than classifying the sentiment or stance of subsequent replies in the conversation tree.
Can the preprocessing scripts handle class imbalance in the PHEME dataset?
Yes. The generate_combined_dev_set function includes an undersampling_neg parameter that reduces the majority class (non-rumors) to match the minority class count. This optional balancing step occurs during CSV aggregation (lines 29-33) before the ShuffleSplit operation creates the final training and validation folds.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →