How to Reproduce the Leave-One-Out Cross-Validation (LOO-CV) Experimental Setup in RPDNN

The RPDNN repository provides pre-computed LOO-CV splits under data/loocv_set_20191002/ where each event folder contains aggregated training data from all other events plus the left-out event’s test set; reproduce the full experimental setup by iterating over each event directory, training with rumour_dnn_trainer.py, and evaluating with rumour_dnn_evaluator.py.

The jerrygaolondon/rpdnn repository implements a strict Leave-One-Out Cross-Validation (LOO-CV) protocol for early rumor detection experiments. This setup ensures that models are trained on all events except one, then tested exclusively on the held-out event, preventing data leakage across temporal boundaries.

Dataset Structure and File Organization

The LOO-CV data is archived in data/cv_dataset/loocv_set_20191002.zip. When extracted to data/loocv_set_20191002/, each event directory (e.g., sydneysiege, charliehebdo, ferguson) contains three CSV files that collectively define a single fold:

  • all_rnr_train_set_combined.csv – Aggregated training tweets from all other events (every event except the folder name).
  • all_rnr_heldout_set_combined.csv – Aggregated validation tweets from all other events, used for early stopping during training.
  • all_rnr_test_set_combined.csv – Test tweets for the left-out event (the folder name), used only for final evaluation.

This structure means that training on the files inside the sydneysiege folder actually trains the model on charliehebdo, ferguson, and other events, while reserving sydneysiege for testing.

Global Feature Scaling Across Folds

Before training, compute global normalization statistics across the entire dataset to ensure consistent scaling in every fold. In src/context_features_extractor.py, the function compute_global_values_4_all_dataset_numercial_features() calculates means and standard deviations for all numerical features using the training sets of every event.

These global values are passed to the trainer via the global_means and global_stds parameters, ensuring that when you train on the aggregated files in one event folder, the feature scaling remains identical to other folds. The repository also hard-codes these statistics as NumPy arrays for reproducibility.

Step-by-Step CLI Reproduction

1. Extract the LOO-CV Archive

Navigate to the repository root and unzip the dataset:

cd /path/to/rpdnn
unzip data/cv_dataset/loocv_set_20191002.zip -d data/

This creates data/loocv_set_20191002/ with subdirectories for each event.

2. Train a Single Fold

Run src/rumour_dnn_trainer.py with the three CSV paths from one event folder. The -f -1 flag selects the full feature set, and --max_cxt_size 200 matches the paper’s experimental conditions:

python src/rumour_dnn_trainer.py \
    -t data/loocv_set_20191002/sydneysiege/all_rnr_train_set_combined.csv \
    --heldout data/loocv_set_20191002/sydneysiege/all_rnr_heldout_set_combined.csv \
    -e data/loocv_set_20191002/sydneysiege/all_rnr_test_set_combined.csv \
    -p "sydneysiege_model" \
    -g 0 \
    -f -1 \
    --max_cxt_size 200 \
    --epochs 10

The trainer invokes model_training() from src/allennlp_rumor_classifier.py, which handles the DNN architecture, attention mechanisms, and early stopping based on the held-out set.

3. Evaluate the Left-Out Event

After training completes, checkpoints are saved under output/<timestamp>/<prefix>/. Evaluate the model on the left-out test set using src/rumour_dnn_evaluator.py:

python src/rumour_dnn_evaluator.py \
    -t data/loocv_set_20191002/sydneysiege/all_rnr_test_set_combined.csv \
    -m output/2024-11-05_12-34-56/sydneysiege_model/ \
    -g 0 \
    -f -1 \
    --max_cxt_size 200

The evaluator loads the checkpoint and prints precision, recall, and F1 scores via timestamped_print in src/allennlp_rumor_classifier.py.

4. Automate the Full LOO-CV Loop

Since the repository does not ship with a top-level orchestration script, wrap the commands in a Bash loop to process every event:

#!/usr/bin/env bash
DATA_ROOT=data/loocv_set_20191002

for ev in "$DATA_ROOT"/*; do
    ev_name=$(basename "$ev")
    echo "=== LOO-CV fold: $ev_name ==="
    
    python src/rumour_dnn_trainer.py \
        -t "$ev/all_rnr_train_set_combined.csv" \
        --heldout "$ev/all_rnr_heldout_set_combined.csv" \
        -e "$ev/all_rnr_test_set_combined.csv" \
        -p "${ev_name}_model" \
        -g 0 -f -1 --max_cxt_size 200 --epochs 10
    
    MODEL_DIR=$(ls -td output/*/${ev_name}_model/ | head -1)
    
    python src/rumour_dnn_evaluator.py \
        -t "$ev/all_rnr_test_set_combined.csv" \
        -m "$MODEL_DIR" \
        -g 0 -f -1 --max_cxt_size 200
done

This executes the complete Leave-One-Out Cross-Validation experimental setup, generating metrics for every event in the dataset.

Programmatic API Usage

Training via Python

For custom pipelines, call model_training() directly from src/allennlp_rumor_classifier.py. Pass the pre-computed global statistics to maintain consistency:

import os
from src.allennlp_rumor_classifier import model_training, config_gpu_use
import numpy as np

# Global statistics computed across all events

means = np.array([33482.21, 113.50, 16110.72, 1480.63, 1984.98,
                  0.4458, 13076.61, 1144.73, 0.01814, 67.85,
                  4.62, 40.20, 0.468, 11.87, 11.46,
                  0.0273, 2.30, 1.52, 0.147, 0.00307,
                  0.887, 0.101, 0.104, 0.080, 12.34,
                  2485.25, 1.0, 0.8286])
stds = np.array([96641.68, 2193.39, 577886.50, 7192.67, 134946.70,
                 0.227, 37326.88, 741.21, 0.133, 632.38,
                 47.67, 612.61, 0.499, 9.10, 5.25,
                 0.163, 127.01, 58.64, 0.354, 0.0553,
                 0.316, 0.301, 0.315, 0.271, 6.87,
                 45822.32, 0.0, 0.3769])

event = "sydneysiege"
root = f"data/loocv_set_20191002/{event}"

config_gpu_use(0)
model_training(
    train_file=os.path.join(root, "all_rnr_train_set_combined.csv"),
    heldout_file=os.path.join(root, "all_rnr_heldout_set_combined.csv"),
    test_file=os.path.join(root, "all_rnr_test_set_combined.csv"),
    cuda_device=0,
    train_batch_size=128,
    model_file_prefix=f"{event}_model",
    global_means=means,
    global_stds=stds,
    feature_setting=-1,  # Full model

    num_epochs=10,
    social_encoder_option=1,
    disable_cxt_type_option=2,
    attention_option=1,
    max_cxt_size_option=200
)

Evaluating via Python

Load a saved checkpoint and compute metrics using the evaluate() function:

from src.allennlp_rumor_classifier import evaluate, config_gpu_use

config_gpu_use(0)

metrics = evaluate(
    model_dir="output/2024-11-05_12-34-56/sydneysiege_model/",
    test_path="data/loocv_set_20191002/sydneysiege/all_rnr_test_set_combined.csv",
    cuda_device=0,
    global_means=means,
    global_stds=stds,
    feature_setting=-1
)
print(metrics)  # {'accuracy': ..., 'f1': ...}

Key Source Files and Functions

File Purpose Key Functions/Components
src/context_features_extractor.py Builds LOO-CV file lists and computes global feature statistics compute_global_values_4_all_dataset_numercial_features()
src/rumour_dnn_trainer.py CLI entry point for training folds Parses -t, --heldout, -e arguments
src/rumour_dnn_evaluator.py CLI entry point for evaluation Loads checkpoints and runs inference
src/allennlp_rumor_classifier.py Core model implementation and metric logging model_training(), evaluate(), timestamped_print()
data/cv_dataset/readme.txt Documentation of CV dataset structure Describes file naming conventions

Summary

  • Extract the LOO-CV archive from data/cv_dataset/loocv_set_20191002.zip to access pre-split folds.
  • Understand the folder structure: Each event directory contains training data from other events and test data for the left-out event.
  • Train using rumour_dnn_trainer.py with global feature scaling parameters (-f -1, --max_cxt_size 200).
  • Evaluate using rumour_dnn_evaluator.py on the test set of the left-out event.
  • Automate the process with a shell loop to complete the full Leave-One-Out Cross-Validation experimental setup across all events.

Frequently Asked Questions

What is the difference between the train, held-out, and test sets in the LOO-CV folders?

In data/loocv_set_20191002/<event>/, the all_rnr_train_set_combined.csv and all_rnr_heldout_set_combined.csv files contain aggregated tweets from all events except the folder name, while all_rnr_test_set_combined.csv contains only tweets from the folder-named event. This ensures strict leave-one-out validation where the model has never seen the test event during training.

How does RPDNN handle feature scaling across different folds?

The repository computes global means and standard deviations across the entire dataset using compute_global_values_4_all_dataset_numercial_features() in src/context_features_extractor.py. These values are passed to every training run via the global_means and global_stds parameters, ensuring identical normalization regardless of which event is left out.

Can I run the LOO-CV experiment without using the command line?

Yes. Import model_training() and evaluate() from src/allennlp_rumor_classifier.py to programmatically train and evaluate folds. Provide the same CSV paths and global statistics arrays used by the CLI scripts to maintain experimental consistency.

Where are the model checkpoints saved during LOO-CV training?

Checkpoints are written to output/<timestamp>/<prefix>/, where <prefix> is the value supplied to the -p argument (e.g., sydneysiege_model). The evaluator requires this directory path to load the trained weights for testing on the left-out event.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →