# How the RPDNN Model Processes the PHEME Dataset Structure for Rumor Detection

> Learn how the RPDNN model processes the PHEME dataset structure for rumor detection. Discover steps like directory scanning, metadata extraction, class balancing, and feature enrichment.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: internals
- Published: 2026-03-04

---

**The RPDNN model processes the PHEME dataset structure for rumor detection by recursively scanning the hierarchical directory tree, extracting source-tweet metadata into structured CSV files, applying undersampling to balance classes, and feeding the final splits to an AllenNLP DatasetReader that enriches each instance with contextual reply and retweet features.**

The RPDNN (Recurrent Posterior Deep Neural Network) repository provides an end-to-end pipeline for rumor detection on the PHEME corpus (dataset ID 6392078). This article explains exactly how the model processes the PHEME dataset structure for rumor detection, transforming raw JSON threads into balanced training data suitable for deep learning.

## Understanding the PHEME Directory Layout

The PHEME corpus organizes rumor threads in a nested hierarchy where each event contains separate folders for rumors and non-rumors. The `load_tweets_context_dataset_dir` function in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) (lines 75-94) walks this structure and builds a lookup dictionary mapping source-tweet IDs to their absolute folder paths.

```python
def load_tweets_context_dataset_dir(social_context_data_dir):
    # ... (omitted for brevity) ...

    all_subdirectories = [x[0] for x in os.walk(social_context_dataset_abs_path)]
    return {os.path.basename(subdirectory): subdirectory
            for subdirectory in all_subdirectories if os.path.basename(subdirectory).isdigit()}

```

This function returns a dictionary like `{'552784600502915072': '/…/552784600502915072', …}`, enabling constant-time lookup of any source tweet's context folder. The directory structure follows the pattern:

```

<dataset_root>/
    <event_name>-all-rnr-threads/
        rumours/
            <tweet_id>/
                source-tweets/   ← Source tweet JSON
                reactions/       ← Reply JSON files
                retweets/        ← Retweet JSON files
        non-rumours/
            <tweet_id>/ …

```

## Extracting Source-Tweet Metadata

### Generating Per-Event CSV Files

The `generate_development_set` function in [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py) (lines 36-82) iterates over every event folder, reads each source-tweet JSON, and extracts the fields required for training. It outputs per-event CSV files containing standardized columns.

```python
def generate_development_set(my_path, output_dir="aug_rnr_training"):
    for folders in glob(os.path.join(my_path, '*')):
        event_name = os.path.basename(folders)               # event name

        # … collect JSON paths …

        for id_f in idfiles:                                 # each thread folder

            source_tweet = glob(os.path.join(id_f, 'source-tweets/*'))
            with open(source_tweet[0], 'r') as f:
                tweet = json.load(f)
                timestamp = pd.DatetimeIndex([tweet['created_at']])
                if category == 'rumours':
                    label = 1
                else:
                    label = 0
                output_df.loc[len(output_df)] = [
                    tweet['id_str'], timestamp[0],
                    tweet.get('full_text', tweet["text"]),
                    label, tweet['user']['id_str'], tweet['user']['screen_name']
                ]
        # write per‑event CSV

        filename = os.path.join(os.path.dirname(my_path),
                                output_dir + '/{}.csv'.format(event_name))
        output_df.to_csv(filename, index=None, encoding='utf-8')

```

The resulting CSVs contain exactly six columns:
- **tweet_id**: Original tweet ID string
- **created_at**: Parsed timestamp (via `pd.DatetimeIndex`)
- **text**: Full tweet text (prefers `full_text` if available)
- **label**: Binary label (`1` = rumour, `0` = non-rumour)
- **user_id**: Author's numeric user ID
- **user_name**: Author's screen name

### Loading Contextual Replies and Retweets

For contextual feature extraction, the `load_source_tweet_context` generator in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) (lines 106-166) streams replies and retweets for any given source-tweet ID. It optionally filters context by a time window relative to the source tweet's creation time.

```python
def load_source_tweet_context(source_tweet_id, event_end_timedelta=None,
                              disable_cxt_type=DISABLE_CXT_TYPE_RETWEET):
    # … initialise dictionary on first call …

    context_tweets_dataset_dir = context_tweets_dataset_dir_dict[source_tweet_id]
    source_tweet_json = load_source_tweet_json(source_tweet_id)

    event_end_time = None
    event_start_time = datetime.strptime(source_tweet_json["created_at"],
                                         '%a %b %d %H:%M:%S %z %Y')
    if event_end_timedelta:
        event_end_time = event_start_time + event_end_timedelta

    # iterate over “reactions” and “retweets”

    for c_type in ["reactions", "retweets"]:
        if disable_cxt_type == DISABLE_CXT_TYPE_REPLY and c_type == "reactions":
            continue
        if disable_cxt_type == DISABLE_CXT_TYPE_RETWEET and c_type == "retweets":
            continue
        reaction_dir = os.path.join(context_tweets_dataset_dir, c_type)
        if not os.path.isdir(reaction_dir):
            continue
        for fname in os.listdir(reaction_dir):
            if fname.startswith('.'):
                continue
            ctx_json = load_tweet_json(os.path.join(reaction_dir, fname))
            ctx_json['context_type'] = c_type
            # optional time‑filtering

            reaction_time = datetime.strptime(ctx_json["created_at"],
                                              '%a %b %d %H:%M:%S %z %Y')
            if event_end_time and (reaction_time < event_start_time or
                                   reaction_time > event_end_time):
                continue
            yield ctx_json

```

This generator yields dictionaries that [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) transforms into numeric vectors—such as retweet counts, sentiment scores, and user credibility metrics—for the neural network.

## Merging, Balancing, and Splitting the Dataset

After generating per-event CSVs, the `generate_combined_dev_set` function in [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py) concatenates the data, addresses class imbalance, and creates the final train/validation/test splits.

The pipeline executes the following steps:
1. **Load all CSV paths** from the development-set directory using `load_files_from_dataset_dir`
2. **Select events** for training (`dev_set_events`) and a held-out test event (`test_set_events`)
3. **Concatenate** DataFrames using `load_matrix_from_csv` into a single NumPy matrix `X`
4. **Check class balance** via `check_dataset_balance`
5. **Undersample the negative class** using `undersampling_neg` for both training and validation sets
6. **Shuffle and split** using scikit-learn's `ShuffleSplit` (5 random splits, selecting one)
7. **Export** final CSVs (`train_X`, `validation_X`, `test_X`) to `data/train/<event>/` or `data/test/<event>/`

Key implementation details from lines 136-165:

```python

# concatenate all dev CSVs

X = None
for dev_set_file in dev_set_path:
    df = load_matrix_from_csv(dev_set_file, header=0,
                              start_col_index=0, end_col_index=6)
    X = df[:] if X is None else np.append(X, df[:], axis=0)

# undersample negative examples

train_X = undersampling_neg(train_X)
validation_X = undersampling_neg(validation_X)
test_X = undersampling_neg(test_X)

# split with ShuffleSplit

rs = ShuffleSplit(n_splits=5, random_state=random.randint(1,10),
                  test_size=0.10, train_size=None)
shuffling_options = list(rs.split(X))

```

The resulting combined CSVs maintain the six-column schema and serve as the exact input for the AllenNLP training pipeline.

## Feeding Processed Data into the RPDNN Model

The [`allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/allennlp_rumor_classifier.py) module defines a custom `DatasetReader` that consumes the merged CSVs. For each row, the reader:
- Tokenizes the tweet text
- Loads contextual features via `load_source_tweet_context`
- Constructs an AllenNLP `Instance` with fields for `tweet_id` (metadata), `text` (tokenized), `label` (binary), `user` (metadata), and context features

During training, the RPDNN model defined in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py) and [`src/rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_evaluator.py) consumes these instances, applying the custom embedding layer from [`src/embeddings/embedding_layer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/embeddings/embedding_layer.py) (which supports pre-trained ELMo embeddings) to perform rumor classification.

## Complete Pipeline Implementation

### Generate Per-Event CSVs from Raw PHEME Data

```python
from src.preprocessing.pheme_data_processor import generate_development_set

pheme_root = "/data/pheme/pheme_training"  # Root containing event folders

generate_development_set(pheme_root, output_dir="aug_rnr_training")

```

*This writes files such as `charliehebdo-all-rnr-threads.csv` in the sibling `aug_rnr_training` folder.*

### Build Balanced Training and Validation Splits

```python
from src.preprocessing.pheme_data_processor import generate_combined_dev_set

dev_dir = "/data/social_context/aug_rnr_training"   # Output from previous step

test_dir = "/data/social_context/all_rnr_training"  # Held-out test structure

generate_combined_dev_set(dev_dir, test_dir)

```

*This creates `pheme_6392078_train_set_combined.csv` and `pheme_6392078_heldout_set_combined.csv` under `data/train/<event>/`.*

### Stream Context Tweets for Feature Extraction

```python
from src.data_loader import load_source_tweet_context

tweet_id = "552784600502915072"
for ctx in load_source_tweet_context(tweet_id, disable_cxt_type=0):
    print(ctx["id_str"], ctx["context_type"], ctx["created_at"])

```

*Setting `disable_cxt_type=0` includes both replies and retweets; use `DISABLE_CXT_TYPE_REPLY` or `DISABLE_CXT_TYPE_RETWEET` to filter.*

### Train the Model with AllenNLP

```bash
allennlp train \
    experiments/config.jsonnet \
    -s output_dir \
    --include-package src

```

*The configuration references the CSVs generated above and the custom DatasetReader that invokes `load_source_tweet_context` for contextual enrichment.*

## Summary

- **Directory Mapping**: `load_tweets_context_dataset_dir` in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) builds a fast lookup table from the hierarchical PHEME structure.
- **CSV Generation**: `generate_development_set` in [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py) extracts source-tweet metadata into standardized six-column CSVs.
- **Context Loading**: `load_source_tweet_context` streams replies and retweets with optional time-window filtering.
- **Data Balancing**: `generate_combined_dev_set` concatenates events, undersamples the majority class, and produces stratified train/validation/test splits.
- **Model Integration**: [`allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/allennlp_rumor_classifier.py) reads the final CSVs and enriches instances with contextual features for the RPDNN classifier.

## Frequently Asked Questions

### What is the PHEME dataset structure for rumor detection?

The PHEME dataset organizes rumor threads hierarchically: each event (e.g., "charliehebdo") contains separate `rumours` and `non-rumours` directories. Within each category, individual folders named by tweet ID contain three subdirectories: `source-tweets/` (the original claim), `reactions/` (reply JSONs), and `retweets/` (retweet JSONs). This structure allows the RPDNN model to isolate source claims from their propagation context.

### How does the RPDNN model handle class imbalance in the PHEME dataset?

The model uses undersampling via the `undersampling_neg` function in [`src/preprocessing/pheme_data_processor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/preprocessing/pheme_data_processor.py). After concatenating all per-event CSVs into a single matrix, the training and validation sets are passed through this function to randomly remove negative (non-rumor) examples until the classes are balanced. This prevents the classifier from bias toward the majority class during training.

### What contextual features are extracted from PHEME threads?

The `load_source_tweet_context` generator in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) yields raw JSON objects for every reply and retweet. These are processed by [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) into numeric vectors including retweet counts, temporal features, sentiment scores, and user credibility metrics. These vectors are appended to the source-tweet embeddings before classification.

### How do I generate training CSVs from a raw PHEME download?

First, run `generate_development_set` pointing to your PHEME root directory to create per-event CSVs. Then execute `generate_combined_dev_set` to merge these into balanced training and test files. Ensure your directory structure matches the expected `*-all-rnr-threads/` pattern, as the preprocessing scripts rely on glob patterns to locate `source-tweets/`, `reactions/`, and `retweets/` subdirectories.