How the RPDNN Model Processes the PHEME Dataset Structure for Rumor Detection
The RPDNN model processes the PHEME dataset structure for rumor detection by recursively scanning the hierarchical directory tree, extracting source-tweet metadata into structured CSV files, applying undersampling to balance classes, and feeding the final splits to an AllenNLP DatasetReader that enriches each instance with contextual reply and retweet features.
The RPDNN (Recurrent Posterior Deep Neural Network) repository provides an end-to-end pipeline for rumor detection on the PHEME corpus (dataset ID 6392078). This article explains exactly how the model processes the PHEME dataset structure for rumor detection, transforming raw JSON threads into balanced training data suitable for deep learning.
Understanding the PHEME Directory Layout
The PHEME corpus organizes rumor threads in a nested hierarchy where each event contains separate folders for rumors and non-rumors. The load_tweets_context_dataset_dir function in src/data_loader.py (lines 75-94) walks this structure and builds a lookup dictionary mapping source-tweet IDs to their absolute folder paths.
def load_tweets_context_dataset_dir(social_context_data_dir):
# ... (omitted for brevity) ...
all_subdirectories = [x[0] for x in os.walk(social_context_dataset_abs_path)]
return {os.path.basename(subdirectory): subdirectory
for subdirectory in all_subdirectories if os.path.basename(subdirectory).isdigit()}
This function returns a dictionary like {'552784600502915072': '/…/552784600502915072', …}, enabling constant-time lookup of any source tweet's context folder. The directory structure follows the pattern:
<dataset_root>/
<event_name>-all-rnr-threads/
rumours/
<tweet_id>/
source-tweets/ ← Source tweet JSON
reactions/ ← Reply JSON files
retweets/ ← Retweet JSON files
non-rumours/
<tweet_id>/ …
Extracting Source-Tweet Metadata
Generating Per-Event CSV Files
The generate_development_set function in src/preprocessing/pheme_data_processor.py (lines 36-82) iterates over every event folder, reads each source-tweet JSON, and extracts the fields required for training. It outputs per-event CSV files containing standardized columns.
def generate_development_set(my_path, output_dir="aug_rnr_training"):
for folders in glob(os.path.join(my_path, '*')):
event_name = os.path.basename(folders) # event name
# … collect JSON paths …
for id_f in idfiles: # each thread folder
source_tweet = glob(os.path.join(id_f, 'source-tweets/*'))
with open(source_tweet[0], 'r') as f:
tweet = json.load(f)
timestamp = pd.DatetimeIndex([tweet['created_at']])
if category == 'rumours':
label = 1
else:
label = 0
output_df.loc[len(output_df)] = [
tweet['id_str'], timestamp[0],
tweet.get('full_text', tweet["text"]),
label, tweet['user']['id_str'], tweet['user']['screen_name']
]
# write per‑event CSV
filename = os.path.join(os.path.dirname(my_path),
output_dir + '/{}.csv'.format(event_name))
output_df.to_csv(filename, index=None, encoding='utf-8')
The resulting CSVs contain exactly six columns:
- tweet_id: Original tweet ID string
- created_at: Parsed timestamp (via
pd.DatetimeIndex) - text: Full tweet text (prefers
full_textif available) - label: Binary label (
1= rumour,0= non-rumour) - user_id: Author's numeric user ID
- user_name: Author's screen name
Loading Contextual Replies and Retweets
For contextual feature extraction, the load_source_tweet_context generator in src/data_loader.py (lines 106-166) streams replies and retweets for any given source-tweet ID. It optionally filters context by a time window relative to the source tweet's creation time.
def load_source_tweet_context(source_tweet_id, event_end_timedelta=None,
disable_cxt_type=DISABLE_CXT_TYPE_RETWEET):
# … initialise dictionary on first call …
context_tweets_dataset_dir = context_tweets_dataset_dir_dict[source_tweet_id]
source_tweet_json = load_source_tweet_json(source_tweet_id)
event_end_time = None
event_start_time = datetime.strptime(source_tweet_json["created_at"],
'%a %b %d %H:%M:%S %z %Y')
if event_end_timedelta:
event_end_time = event_start_time + event_end_timedelta
# iterate over “reactions” and “retweets”
for c_type in ["reactions", "retweets"]:
if disable_cxt_type == DISABLE_CXT_TYPE_REPLY and c_type == "reactions":
continue
if disable_cxt_type == DISABLE_CXT_TYPE_RETWEET and c_type == "retweets":
continue
reaction_dir = os.path.join(context_tweets_dataset_dir, c_type)
if not os.path.isdir(reaction_dir):
continue
for fname in os.listdir(reaction_dir):
if fname.startswith('.'):
continue
ctx_json = load_tweet_json(os.path.join(reaction_dir, fname))
ctx_json['context_type'] = c_type
# optional time‑filtering
reaction_time = datetime.strptime(ctx_json["created_at"],
'%a %b %d %H:%M:%S %z %Y')
if event_end_time and (reaction_time < event_start_time or
reaction_time > event_end_time):
continue
yield ctx_json
This generator yields dictionaries that src/context_features_extractor.py transforms into numeric vectors—such as retweet counts, sentiment scores, and user credibility metrics—for the neural network.
Merging, Balancing, and Splitting the Dataset
After generating per-event CSVs, the generate_combined_dev_set function in src/preprocessing/pheme_data_processor.py concatenates the data, addresses class imbalance, and creates the final train/validation/test splits.
The pipeline executes the following steps:
- Load all CSV paths from the development-set directory using
load_files_from_dataset_dir - Select events for training (
dev_set_events) and a held-out test event (test_set_events) - Concatenate DataFrames using
load_matrix_from_csvinto a single NumPy matrixX - Check class balance via
check_dataset_balance - Undersample the negative class using
undersampling_negfor both training and validation sets - Shuffle and split using scikit-learn's
ShuffleSplit(5 random splits, selecting one) - Export final CSVs (
train_X,validation_X,test_X) todata/train/<event>/ordata/test/<event>/
Key implementation details from lines 136-165:
# concatenate all dev CSVs
X = None
for dev_set_file in dev_set_path:
df = load_matrix_from_csv(dev_set_file, header=0,
start_col_index=0, end_col_index=6)
X = df[:] if X is None else np.append(X, df[:], axis=0)
# undersample negative examples
train_X = undersampling_neg(train_X)
validation_X = undersampling_neg(validation_X)
test_X = undersampling_neg(test_X)
# split with ShuffleSplit
rs = ShuffleSplit(n_splits=5, random_state=random.randint(1,10),
test_size=0.10, train_size=None)
shuffling_options = list(rs.split(X))
The resulting combined CSVs maintain the six-column schema and serve as the exact input for the AllenNLP training pipeline.
Feeding Processed Data into the RPDNN Model
The allennlp_rumor_classifier.py module defines a custom DatasetReader that consumes the merged CSVs. For each row, the reader:
- Tokenizes the tweet text
- Loads contextual features via
load_source_tweet_context - Constructs an AllenNLP
Instancewith fields fortweet_id(metadata),text(tokenized),label(binary),user(metadata), and context features
During training, the RPDNN model defined in src/rumour_dnn_trainer.py and src/rumour_dnn_evaluator.py consumes these instances, applying the custom embedding layer from src/embeddings/embedding_layer.py (which supports pre-trained ELMo embeddings) to perform rumor classification.
Complete Pipeline Implementation
Generate Per-Event CSVs from Raw PHEME Data
from src.preprocessing.pheme_data_processor import generate_development_set
pheme_root = "/data/pheme/pheme_training" # Root containing event folders
generate_development_set(pheme_root, output_dir="aug_rnr_training")
This writes files such as charliehebdo-all-rnr-threads.csv in the sibling aug_rnr_training folder.
Build Balanced Training and Validation Splits
from src.preprocessing.pheme_data_processor import generate_combined_dev_set
dev_dir = "/data/social_context/aug_rnr_training" # Output from previous step
test_dir = "/data/social_context/all_rnr_training" # Held-out test structure
generate_combined_dev_set(dev_dir, test_dir)
This creates pheme_6392078_train_set_combined.csv and pheme_6392078_heldout_set_combined.csv under data/train/<event>/.
Stream Context Tweets for Feature Extraction
from src.data_loader import load_source_tweet_context
tweet_id = "552784600502915072"
for ctx in load_source_tweet_context(tweet_id, disable_cxt_type=0):
print(ctx["id_str"], ctx["context_type"], ctx["created_at"])
Setting disable_cxt_type=0 includes both replies and retweets; use DISABLE_CXT_TYPE_REPLY or DISABLE_CXT_TYPE_RETWEET to filter.
Train the Model with AllenNLP
allennlp train \
experiments/config.jsonnet \
-s output_dir \
--include-package src
The configuration references the CSVs generated above and the custom DatasetReader that invokes load_source_tweet_context for contextual enrichment.
Summary
- Directory Mapping:
load_tweets_context_dataset_dirinsrc/data_loader.pybuilds a fast lookup table from the hierarchical PHEME structure. - CSV Generation:
generate_development_setinsrc/preprocessing/pheme_data_processor.pyextracts source-tweet metadata into standardized six-column CSVs. - Context Loading:
load_source_tweet_contextstreams replies and retweets with optional time-window filtering. - Data Balancing:
generate_combined_dev_setconcatenates events, undersamples the majority class, and produces stratified train/validation/test splits. - Model Integration:
allennlp_rumor_classifier.pyreads the final CSVs and enriches instances with contextual features for the RPDNN classifier.
Frequently Asked Questions
What is the PHEME dataset structure for rumor detection?
The PHEME dataset organizes rumor threads hierarchically: each event (e.g., "charliehebdo") contains separate rumours and non-rumours directories. Within each category, individual folders named by tweet ID contain three subdirectories: source-tweets/ (the original claim), reactions/ (reply JSONs), and retweets/ (retweet JSONs). This structure allows the RPDNN model to isolate source claims from their propagation context.
How does the RPDNN model handle class imbalance in the PHEME dataset?
The model uses undersampling via the undersampling_neg function in src/preprocessing/pheme_data_processor.py. After concatenating all per-event CSVs into a single matrix, the training and validation sets are passed through this function to randomly remove negative (non-rumor) examples until the classes are balanced. This prevents the classifier from bias toward the majority class during training.
What contextual features are extracted from PHEME threads?
The load_source_tweet_context generator in src/data_loader.py yields raw JSON objects for every reply and retweet. These are processed by src/context_features_extractor.py into numeric vectors including retweet counts, temporal features, sentiment scores, and user credibility metrics. These vectors are appended to the source-tweet embeddings before classification.
How do I generate training CSVs from a raw PHEME download?
First, run generate_development_set pointing to your PHEME root directory to create per-event CSVs. Then execute generate_combined_dev_set to merge these into balanced training and test files. Ensure your directory structure matches the expected *-all-rnr-threads/ pattern, as the preprocessing scripts rely on glob patterns to locate source-tweets/, reactions/, and retweets/ subdirectories.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →