How to Use rumour_dnn_evaluator.py with Custom Trained Models
The rumour_dnn_evaluator.py script evaluates custom-trained RP-DNN models by loading an AllenNLP archive directory containing weights_best.th and vocabulary/ files, then applies mirrored training configurations via command-line flags to ensure consistent inference.
The jerrygaolondon/rpdnn repository implements a deep neural network for rumor detection that combines ELMo text embeddings with social context encoders. The rumour_dnn_evaluator.py script serves as the command-line interface for testing custom models against new datasets, requiring strict parameter alignment between training and evaluation phases to produce valid accuracy and F1 metrics.
Required Model Archive Structure
Before running the evaluator, your custom trained model must be serialized in a directory containing two critical components:
weights_best.th– The serialized model parameters produced byrumour_dnn_trainer.pyduring the training loop.vocabulary/– The AllenNLP vocabulary directory built from your training data, containing token-to-index mappings.
The script invokes load_classifier_from_archive() (source lines 28‑35) to deserialize the RumorTweetsClassifer defined in src/allennlp_rumor_classifier.py. This function reconstructs the exact architecture—including ELMo text encoders, optional LSTM or Transformer context encoders, and hierarchical attention networks—before loading the saved weights.
Critical Configuration Alignment
All architectural options must match the values used during training. The script mirrors configuration parameters onto the model object at runtime (source lines 46‑48), including:
- Feature setting (
-for--feature_setting): Determines which input streams are concatenated (source tweet only, context content, metadata, or combinations) as defined inallennlp_rumor_classifier.py(lines 71‑78). - Maximum context size (
--max_cxt_size): The number of reply tweets processed per source. - GPU device (
-g): CUDA device index or-1for CPU.
Mismatched settings raise ValueError exceptions during model reconfiguration (source lines 84‑88) because the loaded weights will not align with the instantiated layer dimensions.
Step-by-Step Evaluation Commands
1. Basic Evaluation of a Full Model
Evaluate a model trained with the complete feature set (-f -1) against a held-out test set:
python src/rumour_dnn_evaluator.py \
-t data/cv_dataset/test.csv \
-m output/2026-03-04_12-30-00/my_custom_model \
-g 0 \
-f -1 \
--max_cxt_size 200
The -m flag points to the directory containing weights_best.th. The script builds a BucketIterator for the test CSV (source lines 42‑45) and outputs per-batch metrics followed by aggregate scores:
accuracy: 0.842
precision: 0.81
recall: 0.79
f1: 0.80
2. Evaluating Source-Only Models
If you trained a model using only source tweet content without social context (-f 1), you must specify the identical feature setting during evaluation:
python src/rumour_dnn_evaluator.py \
-t data/cv_dataset/test.csv \
-m path/to/source_only_model \
-g -1 \
-f 1
Omitting the -f flag or using a different value causes immediate termination because the model's feature_setting attribute is strictly validated against the command-line option.
3. Disabling Retweet Context Dynamically
Control which social context types influence predictions using the --disable_context_type flag. This parameter is forwarded to model.set_disable_cxt_type_option() (source lines 49‑51):
python src/rumour_dnn_evaluator.py \
-t data/cv_dataset/test.csv \
-m path/to/model \
--disable_context_type 2
Value 2 disables retweets while preserving other reply types, allowing ablation studies without retraining.
4. CPU-Only Evaluation
For environments without GPU access, specify device -1 to force CPU computation:
python src/rumour_dnn_evaluator.py \
-t data/cv_dataset/test.csv \
-m path/to/model \
-g -1 \
-f -1
Architecture Validation Flow
The evaluation pipeline follows this strict sequence to ensure architectural fidelity:
- Archive Loading:
load_classifier_from_archive()reconstructs theRumorTweetsClassiferwithtweet_text_embedder(ELMo-based) and optionalcxt_content_encoder/cxt_metadata_encoderlayers. - Configuration Mirroring: The script assigns command-line options to model attributes (e.g.,
model.feature_setting = feature_setting_option). - Iterator Construction: A
BucketIteratorbatches the test CSV data according to the model's expected input fields. - Forward Pass: The model executes
forward()through text encoders, social context processors, and the feed-forward classifier to produce logits. - Metric Computation: Accuracy, precision, recall, and F1 scores are calculated across all batches.
Because the evaluator re-instantiates the exact training architecture, custom models must be trained with the same code base version. Changes to layer definitions in allennlp_rumor_classifier.py between training and evaluation will cause tensor shape mismatches.
Summary
- Archive Requirements: The evaluator requires a model directory containing
weights_best.thand avocabulary/folder produced byrumour_dnn_trainer.py. - Configuration Lock: Feature settings (
-f), context sizes, and attention types must exactly match training parameters or the script raisesValueError. - CLI Interface: All options are passed via command-line flags that are mirrored onto the model object at runtime in
rumour_dnn_evaluator.py. - Social Context Control: Use
--disable_context_typeto ablate specific reply types without model modification. - Source Files: Core logic resides in
src/allennlp_rumor_classifier.py(model definition) andsrc/rumour_dnn_evaluator.py(evaluation wrapper).
Frequently Asked Questions
What files are required to run rumour_dnn_evaluator.py with a custom model?
You need a model archive directory containing weights_best.th (the serialized parameters) and a vocabulary/ subdirectory (AllenNLP token mappings). These files are generated automatically when training completes via rumour_dnn_trainer.py. The evaluator calls load_classifier_from_archive() to reconstruct the model architecture from these artifacts.
Why do I get a ValueError when evaluating my custom trained model?
This error occurs when the evaluation configuration does not match the training configuration. Specifically, the feature_setting value (-f flag) used during evaluation must be identical to the value used during training, as the script validates this alignment at runtime (source lines 84‑88). Mismatched context sizes or attention types will also trigger validation failures.
Can I evaluate a GPU-trained model on CPU using rumour_dnn_evaluator.py?
Yes. Use the -g -1 flag to specify CPU-only execution. The script maps this value to the model's device configuration during initialization. However, ensure you have sufficient RAM, as ELMo embeddings and social context encoders are memory-intensive when processed without GPU acceleration.
How does rumour_dnn_evaluator.py handle different social context features?
The script passes the feature_setting flag to the model's initialization logic, which determines which encoders are active: source tweet only (-f 1), context content (-f 2), metadata (-f 3), or combinations thereof. Additionally, the --disable_context_type option allows runtime filtering of specific reply types (e.g., retweets) via model.set_disable_cxt_type_option() without reloading the model weights.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →