How to Use rumour_dnn_evaluator.py with Custom Trained Models

The rumour_dnn_evaluator.py script evaluates custom-trained RP-DNN models by loading an AllenNLP archive directory containing weights_best.th and vocabulary/ files, then applies mirrored training configurations via command-line flags to ensure consistent inference.

The jerrygaolondon/rpdnn repository implements a deep neural network for rumor detection that combines ELMo text embeddings with social context encoders. The rumour_dnn_evaluator.py script serves as the command-line interface for testing custom models against new datasets, requiring strict parameter alignment between training and evaluation phases to produce valid accuracy and F1 metrics.

Required Model Archive Structure

Before running the evaluator, your custom trained model must be serialized in a directory containing two critical components:

  • weights_best.th – The serialized model parameters produced by rumour_dnn_trainer.py during the training loop.
  • vocabulary/ – The AllenNLP vocabulary directory built from your training data, containing token-to-index mappings.

The script invokes load_classifier_from_archive() (source lines 28‑35) to deserialize the RumorTweetsClassifer defined in src/allennlp_rumor_classifier.py. This function reconstructs the exact architecture—including ELMo text encoders, optional LSTM or Transformer context encoders, and hierarchical attention networks—before loading the saved weights.

Critical Configuration Alignment

All architectural options must match the values used during training. The script mirrors configuration parameters onto the model object at runtime (source lines 46‑48), including:

  • Feature setting (-f or --feature_setting): Determines which input streams are concatenated (source tweet only, context content, metadata, or combinations) as defined in allennlp_rumor_classifier.py (lines 71‑78).
  • Maximum context size (--max_cxt_size): The number of reply tweets processed per source.
  • GPU device (-g): CUDA device index or -1 for CPU.

Mismatched settings raise ValueError exceptions during model reconfiguration (source lines 84‑88) because the loaded weights will not align with the instantiated layer dimensions.

Step-by-Step Evaluation Commands

1. Basic Evaluation of a Full Model

Evaluate a model trained with the complete feature set (-f -1) against a held-out test set:

python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m output/2026-03-04_12-30-00/my_custom_model \
    -g 0 \
    -f -1 \
    --max_cxt_size 200

The -m flag points to the directory containing weights_best.th. The script builds a BucketIterator for the test CSV (source lines 42‑45) and outputs per-batch metrics followed by aggregate scores:


accuracy: 0.842
precision: 0.81
recall: 0.79
f1: 0.80

2. Evaluating Source-Only Models

If you trained a model using only source tweet content without social context (-f 1), you must specify the identical feature setting during evaluation:

python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m path/to/source_only_model \
    -g -1 \
    -f 1

Omitting the -f flag or using a different value causes immediate termination because the model's feature_setting attribute is strictly validated against the command-line option.

3. Disabling Retweet Context Dynamically

Control which social context types influence predictions using the --disable_context_type flag. This parameter is forwarded to model.set_disable_cxt_type_option() (source lines 49‑51):

python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m path/to/model \
    --disable_context_type 2

Value 2 disables retweets while preserving other reply types, allowing ablation studies without retraining.

4. CPU-Only Evaluation

For environments without GPU access, specify device -1 to force CPU computation:

python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m path/to/model \
    -g -1 \
    -f -1

Architecture Validation Flow

The evaluation pipeline follows this strict sequence to ensure architectural fidelity:

  1. Archive Loading: load_classifier_from_archive() reconstructs the RumorTweetsClassifer with tweet_text_embedder (ELMo-based) and optional cxt_content_encoder/cxt_metadata_encoder layers.
  2. Configuration Mirroring: The script assigns command-line options to model attributes (e.g., model.feature_setting = feature_setting_option).
  3. Iterator Construction: A BucketIterator batches the test CSV data according to the model's expected input fields.
  4. Forward Pass: The model executes forward() through text encoders, social context processors, and the feed-forward classifier to produce logits.
  5. Metric Computation: Accuracy, precision, recall, and F1 scores are calculated across all batches.

Because the evaluator re-instantiates the exact training architecture, custom models must be trained with the same code base version. Changes to layer definitions in allennlp_rumor_classifier.py between training and evaluation will cause tensor shape mismatches.

Summary

  • Archive Requirements: The evaluator requires a model directory containing weights_best.th and a vocabulary/ folder produced by rumour_dnn_trainer.py.
  • Configuration Lock: Feature settings (-f), context sizes, and attention types must exactly match training parameters or the script raises ValueError.
  • CLI Interface: All options are passed via command-line flags that are mirrored onto the model object at runtime in rumour_dnn_evaluator.py.
  • Social Context Control: Use --disable_context_type to ablate specific reply types without model modification.
  • Source Files: Core logic resides in src/allennlp_rumor_classifier.py (model definition) and src/rumour_dnn_evaluator.py (evaluation wrapper).

Frequently Asked Questions

What files are required to run rumour_dnn_evaluator.py with a custom model?

You need a model archive directory containing weights_best.th (the serialized parameters) and a vocabulary/ subdirectory (AllenNLP token mappings). These files are generated automatically when training completes via rumour_dnn_trainer.py. The evaluator calls load_classifier_from_archive() to reconstruct the model architecture from these artifacts.

Why do I get a ValueError when evaluating my custom trained model?

This error occurs when the evaluation configuration does not match the training configuration. Specifically, the feature_setting value (-f flag) used during evaluation must be identical to the value used during training, as the script validates this alignment at runtime (source lines 84‑88). Mismatched context sizes or attention types will also trigger validation failures.

Can I evaluate a GPU-trained model on CPU using rumour_dnn_evaluator.py?

Yes. Use the -g -1 flag to specify CPU-only execution. The script maps this value to the model's device configuration during initialization. However, ensure you have sufficient RAM, as ELMo embeddings and social context encoders are memory-intensive when processed without GPU acceleration.

How does rumour_dnn_evaluator.py handle different social context features?

The script passes the feature_setting flag to the model's initialization logic, which determines which encoders are active: source tweet only (-f 1), context content (-f 2), metadata (-f 3), or combinations thereof. Additionally, the --disable_context_type option allows runtime filtering of specific reply types (e.g., retweets) via model.set_disable_cxt_type_option() without reloading the model weights.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →