# How to Use rumour_dnn_evaluator.py with Custom Trained Models

> Evaluate your custom-trained rumour_dnn models using rumour_dnn_evaluator.py. Load your AllenNLP archive and apply mirrored training configurations for consistent inference.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: how-to-guide
- Published: 2026-03-04

---

**The [`rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_evaluator.py) script evaluates custom-trained RP-DNN models by loading an AllenNLP archive directory containing `weights_best.th` and `vocabulary/` files, then applies mirrored training configurations via command-line flags to ensure consistent inference.**

The `jerrygaolondon/rpdnn` repository implements a deep neural network for rumor detection that combines ELMo text embeddings with social context encoders. The [`rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_evaluator.py) script serves as the command-line interface for testing custom models against new datasets, requiring strict parameter alignment between training and evaluation phases to produce valid accuracy and F1 metrics.

## Required Model Archive Structure

Before running the evaluator, your custom trained model must be serialized in a directory containing two critical components:

- **`weights_best.th`** – The serialized model parameters produced by [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py) during the training loop.
- **`vocabulary/`** – The AllenNLP vocabulary directory built from your training data, containing token-to-index mappings.

The script invokes `load_classifier_from_archive()` ([source lines 28‑35](/blob/master/src/rumour_dnn_evaluator.py#L28-L35)) to deserialize the `RumorTweetsClassifer` defined in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py). This function reconstructs the exact architecture—including ELMo text encoders, optional LSTM or Transformer context encoders, and hierarchical attention networks—before loading the saved weights.

## Critical Configuration Alignment

All architectural options must match the values used during training. The script mirrors configuration parameters onto the model object at runtime ([source lines 46‑48](/blob/master/src/rumour_dnn_evaluator.py#L46-L48)), including:

- **Feature setting** (`-f` or `--feature_setting`): Determines which input streams are concatenated (source tweet only, context content, metadata, or combinations) as defined in [`allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/allennlp_rumor_classifier.py) ([lines 71‑78](/blob/master/src/allennlp_rumor_classifier.py#L71-L78)).
- **Maximum context size** (`--max_cxt_size`): The number of reply tweets processed per source.
- **GPU device** (`-g`): CUDA device index or `-1` for CPU.

Mismatched settings raise `ValueError` exceptions during model reconfiguration ([source lines 84‑88](/blob/master/src/rumour_dnn_evaluator.py#L84-L88)) because the loaded weights will not align with the instantiated layer dimensions.

## Step-by-Step Evaluation Commands

### 1. Basic Evaluation of a Full Model

Evaluate a model trained with the complete feature set (`-f -1`) against a held-out test set:

```bash
python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m output/2026-03-04_12-30-00/my_custom_model \
    -g 0 \
    -f -1 \
    --max_cxt_size 200

```

The `-m` flag points to the directory containing `weights_best.th`. The script builds a `BucketIterator` for the test CSV ([source lines 42‑45](/blob/master/src/rumour_dnn_evaluator.py#L42-L45)) and outputs per-batch metrics followed by aggregate scores:

```

accuracy: 0.842
precision: 0.81
recall: 0.79
f1: 0.80

```

### 2. Evaluating Source-Only Models

If you trained a model using only source tweet content without social context (`-f 1`), you must specify the identical feature setting during evaluation:

```bash
python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m path/to/source_only_model \
    -g -1 \
    -f 1

```

Omitting the `-f` flag or using a different value causes immediate termination because the model's `feature_setting` attribute is strictly validated against the command-line option.

### 3. Disabling Retweet Context Dynamically

Control which social context types influence predictions using the `--disable_context_type` flag. This parameter is forwarded to `model.set_disable_cxt_type_option()` ([source lines 49‑51](/blob/master/src/rumour_dnn_evaluator.py#L49-L51)):

```bash
python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m path/to/model \
    --disable_context_type 2

```

Value `2` disables retweets while preserving other reply types, allowing ablation studies without retraining.

### 4. CPU-Only Evaluation

For environments without GPU access, specify device `-1` to force CPU computation:

```bash
python src/rumour_dnn_evaluator.py \
    -t data/cv_dataset/test.csv \
    -m path/to/model \
    -g -1 \
    -f -1

```

## Architecture Validation Flow

The evaluation pipeline follows this strict sequence to ensure architectural fidelity:

1. **Archive Loading**: `load_classifier_from_archive()` reconstructs the `RumorTweetsClassifer` with `tweet_text_embedder` (ELMo-based) and optional `cxt_content_encoder`/`cxt_metadata_encoder` layers.
2. **Configuration Mirroring**: The script assigns command-line options to model attributes (e.g., `model.feature_setting = feature_setting_option`).
3. **Iterator Construction**: A `BucketIterator` batches the test CSV data according to the model's expected input fields.
4. **Forward Pass**: The model executes `forward()` through text encoders, social context processors, and the feed-forward classifier to produce logits.
5. **Metric Computation**: Accuracy, precision, recall, and F1 scores are calculated across all batches.

Because the evaluator re-instantiates the exact training architecture, **custom models must be trained with the same code base version**. Changes to layer definitions in [`allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/allennlp_rumor_classifier.py) between training and evaluation will cause tensor shape mismatches.

## Summary

- **Archive Requirements**: The evaluator requires a model directory containing `weights_best.th` and a `vocabulary/` folder produced by [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py).
- **Configuration Lock**: Feature settings (`-f`), context sizes, and attention types must exactly match training parameters or the script raises `ValueError`.
- **CLI Interface**: All options are passed via command-line flags that are mirrored onto the model object at runtime in [`rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_evaluator.py).
- **Social Context Control**: Use `--disable_context_type` to ablate specific reply types without model modification.
- **Source Files**: Core logic resides in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) (model definition) and [`src/rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_evaluator.py) (evaluation wrapper).

## Frequently Asked Questions

### What files are required to run rumour_dnn_evaluator.py with a custom model?

You need a model archive directory containing `weights_best.th` (the serialized parameters) and a `vocabulary/` subdirectory (AllenNLP token mappings). These files are generated automatically when training completes via [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py). The evaluator calls `load_classifier_from_archive()` to reconstruct the model architecture from these artifacts.

### Why do I get a ValueError when evaluating my custom trained model?

This error occurs when the evaluation configuration does not match the training configuration. Specifically, the `feature_setting` value (`-f` flag) used during evaluation must be identical to the value used during training, as the script validates this alignment at runtime ([source lines 84‑88](/blob/master/src/rumour_dnn_evaluator.py#L84-L88)). Mismatched context sizes or attention types will also trigger validation failures.

### Can I evaluate a GPU-trained model on CPU using rumour_dnn_evaluator.py?

Yes. Use the `-g -1` flag to specify CPU-only execution. The script maps this value to the model's device configuration during initialization. However, ensure you have sufficient RAM, as ELMo embeddings and social context encoders are memory-intensive when processed without GPU acceleration.

### How does rumour_dnn_evaluator.py handle different social context features?

The script passes the `feature_setting` flag to the model's initialization logic, which determines which encoders are active: source tweet only (`-f 1`), context content (`-f 2`), metadata (`-f 3`), or combinations thereof. Additionally, the `--disable_context_type` option allows runtime filtering of specific reply types (e.g., retweets) via `model.set_disable_cxt_type_option()` without reloading the model weights.