How the RPDNN Model Processes the PHEME Dataset Structure for Rumor Detection

The RPDNN model processes the PHEME dataset structure for rumor detection by recursively scanning the hierarchical directory tree, extracting source-tweet metadata into structured CSV files, applying undersampling to balance classes, and feeding the final splits to an AllenNLP DatasetReader that enriches each instance with contextual reply and retweet features.

The RPDNN (Recurrent Posterior Deep Neural Network) repository provides an end-to-end pipeline for rumor detection on the PHEME corpus (dataset ID 6392078). This article explains exactly how the model processes the PHEME dataset structure for rumor detection, transforming raw JSON threads into balanced training data suitable for deep learning.

Understanding the PHEME Directory Layout

The PHEME corpus organizes rumor threads in a nested hierarchy where each event contains separate folders for rumors and non-rumors. The load_tweets_context_dataset_dir function in src/data_loader.py (lines 75-94) walks this structure and builds a lookup dictionary mapping source-tweet IDs to their absolute folder paths.

def load_tweets_context_dataset_dir(social_context_data_dir):
    # ... (omitted for brevity) ...

    all_subdirectories = [x[0] for x in os.walk(social_context_dataset_abs_path)]
    return {os.path.basename(subdirectory): subdirectory
            for subdirectory in all_subdirectories if os.path.basename(subdirectory).isdigit()}

This function returns a dictionary like {'552784600502915072': '/…/552784600502915072', …}, enabling constant-time lookup of any source tweet's context folder. The directory structure follows the pattern:


<dataset_root>/
    <event_name>-all-rnr-threads/
        rumours/
            <tweet_id>/
                source-tweets/   ← Source tweet JSON
                reactions/       ← Reply JSON files
                retweets/        ← Retweet JSON files
        non-rumours/
            <tweet_id>/ …

Extracting Source-Tweet Metadata

Generating Per-Event CSV Files

The generate_development_set function in src/preprocessing/pheme_data_processor.py (lines 36-82) iterates over every event folder, reads each source-tweet JSON, and extracts the fields required for training. It outputs per-event CSV files containing standardized columns.

def generate_development_set(my_path, output_dir="aug_rnr_training"):
    for folders in glob(os.path.join(my_path, '*')):
        event_name = os.path.basename(folders)               # event name

        # … collect JSON paths …

        for id_f in idfiles:                                 # each thread folder

            source_tweet = glob(os.path.join(id_f, 'source-tweets/*'))
            with open(source_tweet[0], 'r') as f:
                tweet = json.load(f)
                timestamp = pd.DatetimeIndex([tweet['created_at']])
                if category == 'rumours':
                    label = 1
                else:
                    label = 0
                output_df.loc[len(output_df)] = [
                    tweet['id_str'], timestamp[0],
                    tweet.get('full_text', tweet["text"]),
                    label, tweet['user']['id_str'], tweet['user']['screen_name']
                ]
        # write per‑event CSV

        filename = os.path.join(os.path.dirname(my_path),
                                output_dir + '/{}.csv'.format(event_name))
        output_df.to_csv(filename, index=None, encoding='utf-8')

The resulting CSVs contain exactly six columns:

  • tweet_id: Original tweet ID string
  • created_at: Parsed timestamp (via pd.DatetimeIndex)
  • text: Full tweet text (prefers full_text if available)
  • label: Binary label (1 = rumour, 0 = non-rumour)
  • user_id: Author's numeric user ID
  • user_name: Author's screen name

Loading Contextual Replies and Retweets

For contextual feature extraction, the load_source_tweet_context generator in src/data_loader.py (lines 106-166) streams replies and retweets for any given source-tweet ID. It optionally filters context by a time window relative to the source tweet's creation time.

def load_source_tweet_context(source_tweet_id, event_end_timedelta=None,
                              disable_cxt_type=DISABLE_CXT_TYPE_RETWEET):
    # … initialise dictionary on first call …

    context_tweets_dataset_dir = context_tweets_dataset_dir_dict[source_tweet_id]
    source_tweet_json = load_source_tweet_json(source_tweet_id)

    event_end_time = None
    event_start_time = datetime.strptime(source_tweet_json["created_at"],
                                         '%a %b %d %H:%M:%S %z %Y')
    if event_end_timedelta:
        event_end_time = event_start_time + event_end_timedelta

    # iterate over “reactions” and “retweets”

    for c_type in ["reactions", "retweets"]:
        if disable_cxt_type == DISABLE_CXT_TYPE_REPLY and c_type == "reactions":
            continue
        if disable_cxt_type == DISABLE_CXT_TYPE_RETWEET and c_type == "retweets":
            continue
        reaction_dir = os.path.join(context_tweets_dataset_dir, c_type)
        if not os.path.isdir(reaction_dir):
            continue
        for fname in os.listdir(reaction_dir):
            if fname.startswith('.'):
                continue
            ctx_json = load_tweet_json(os.path.join(reaction_dir, fname))
            ctx_json['context_type'] = c_type
            # optional time‑filtering

            reaction_time = datetime.strptime(ctx_json["created_at"],
                                              '%a %b %d %H:%M:%S %z %Y')
            if event_end_time and (reaction_time < event_start_time or
                                   reaction_time > event_end_time):
                continue
            yield ctx_json

This generator yields dictionaries that src/context_features_extractor.py transforms into numeric vectors—such as retweet counts, sentiment scores, and user credibility metrics—for the neural network.

Merging, Balancing, and Splitting the Dataset

After generating per-event CSVs, the generate_combined_dev_set function in src/preprocessing/pheme_data_processor.py concatenates the data, addresses class imbalance, and creates the final train/validation/test splits.

The pipeline executes the following steps:

  1. Load all CSV paths from the development-set directory using load_files_from_dataset_dir
  2. Select events for training (dev_set_events) and a held-out test event (test_set_events)
  3. Concatenate DataFrames using load_matrix_from_csv into a single NumPy matrix X
  4. Check class balance via check_dataset_balance
  5. Undersample the negative class using undersampling_neg for both training and validation sets
  6. Shuffle and split using scikit-learn's ShuffleSplit (5 random splits, selecting one)
  7. Export final CSVs (train_X, validation_X, test_X) to data/train/<event>/ or data/test/<event>/

Key implementation details from lines 136-165:


# concatenate all dev CSVs

X = None
for dev_set_file in dev_set_path:
    df = load_matrix_from_csv(dev_set_file, header=0,
                              start_col_index=0, end_col_index=6)
    X = df[:] if X is None else np.append(X, df[:], axis=0)

# undersample negative examples

train_X = undersampling_neg(train_X)
validation_X = undersampling_neg(validation_X)
test_X = undersampling_neg(test_X)

# split with ShuffleSplit

rs = ShuffleSplit(n_splits=5, random_state=random.randint(1,10),
                  test_size=0.10, train_size=None)
shuffling_options = list(rs.split(X))

The resulting combined CSVs maintain the six-column schema and serve as the exact input for the AllenNLP training pipeline.

Feeding Processed Data into the RPDNN Model

The allennlp_rumor_classifier.py module defines a custom DatasetReader that consumes the merged CSVs. For each row, the reader:

  • Tokenizes the tweet text
  • Loads contextual features via load_source_tweet_context
  • Constructs an AllenNLP Instance with fields for tweet_id (metadata), text (tokenized), label (binary), user (metadata), and context features

During training, the RPDNN model defined in src/rumour_dnn_trainer.py and src/rumour_dnn_evaluator.py consumes these instances, applying the custom embedding layer from src/embeddings/embedding_layer.py (which supports pre-trained ELMo embeddings) to perform rumor classification.

Complete Pipeline Implementation

Generate Per-Event CSVs from Raw PHEME Data

from src.preprocessing.pheme_data_processor import generate_development_set

pheme_root = "/data/pheme/pheme_training"  # Root containing event folders

generate_development_set(pheme_root, output_dir="aug_rnr_training")

This writes files such as charliehebdo-all-rnr-threads.csv in the sibling aug_rnr_training folder.

Build Balanced Training and Validation Splits

from src.preprocessing.pheme_data_processor import generate_combined_dev_set

dev_dir = "/data/social_context/aug_rnr_training"   # Output from previous step

test_dir = "/data/social_context/all_rnr_training"  # Held-out test structure

generate_combined_dev_set(dev_dir, test_dir)

This creates pheme_6392078_train_set_combined.csv and pheme_6392078_heldout_set_combined.csv under data/train/<event>/.

Stream Context Tweets for Feature Extraction

from src.data_loader import load_source_tweet_context

tweet_id = "552784600502915072"
for ctx in load_source_tweet_context(tweet_id, disable_cxt_type=0):
    print(ctx["id_str"], ctx["context_type"], ctx["created_at"])

Setting disable_cxt_type=0 includes both replies and retweets; use DISABLE_CXT_TYPE_REPLY or DISABLE_CXT_TYPE_RETWEET to filter.

Train the Model with AllenNLP

allennlp train \
    experiments/config.jsonnet \
    -s output_dir \
    --include-package src

The configuration references the CSVs generated above and the custom DatasetReader that invokes load_source_tweet_context for contextual enrichment.

Summary

  • Directory Mapping: load_tweets_context_dataset_dir in src/data_loader.py builds a fast lookup table from the hierarchical PHEME structure.
  • CSV Generation: generate_development_set in src/preprocessing/pheme_data_processor.py extracts source-tweet metadata into standardized six-column CSVs.
  • Context Loading: load_source_tweet_context streams replies and retweets with optional time-window filtering.
  • Data Balancing: generate_combined_dev_set concatenates events, undersamples the majority class, and produces stratified train/validation/test splits.
  • Model Integration: allennlp_rumor_classifier.py reads the final CSVs and enriches instances with contextual features for the RPDNN classifier.

Frequently Asked Questions

What is the PHEME dataset structure for rumor detection?

The PHEME dataset organizes rumor threads hierarchically: each event (e.g., "charliehebdo") contains separate rumours and non-rumours directories. Within each category, individual folders named by tweet ID contain three subdirectories: source-tweets/ (the original claim), reactions/ (reply JSONs), and retweets/ (retweet JSONs). This structure allows the RPDNN model to isolate source claims from their propagation context.

How does the RPDNN model handle class imbalance in the PHEME dataset?

The model uses undersampling via the undersampling_neg function in src/preprocessing/pheme_data_processor.py. After concatenating all per-event CSVs into a single matrix, the training and validation sets are passed through this function to randomly remove negative (non-rumor) examples until the classes are balanced. This prevents the classifier from bias toward the majority class during training.

What contextual features are extracted from PHEME threads?

The load_source_tweet_context generator in src/data_loader.py yields raw JSON objects for every reply and retweet. These are processed by src/context_features_extractor.py into numeric vectors including retweet counts, temporal features, sentiment scores, and user credibility metrics. These vectors are appended to the source-tweet embeddings before classification.

How do I generate training CSVs from a raw PHEME download?

First, run generate_development_set pointing to your PHEME root directory to create per-event CSVs. Then execute generate_combined_dev_set to merge these into balanced training and test files. Ensure your directory structure matches the expected *-all-rnr-threads/ pattern, as the preprocessing scripts rely on glob patterns to locate source-tweets/, reactions/, and retweets/ subdirectories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →