# RPDNN Evaluation Metrics: F1, Accuracy, Precision, and Recall Explained

> Understand RPDNN evaluation metrics like F1, accuracy, precision, and recall. Learn how the Rumour DNN model assesses its performance for better results.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: performance
- Published: 2026-03-04

---

**The RPDNN (Rumour DNN) model evaluates performance using Accuracy, Precision, Recall, F1-score, and cross-entropy Loss, calculated via the `get_metrics()` method in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) and aggregated by `training_util.evaluate` in [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py).**

The jerrygaolondon/rpdnn repository implements a deep neural network for rumour detection on Twitter. Understanding how this model assesses its predictions requires examining the specific evaluation metrics built into its architecture. The system tracks both classification accuracy and nuanced performance indicators like F1-score to provide a complete picture of model effectiveness on rumour classification tasks.

## Core Evaluation Metrics in RPDNN

The model produces five key evaluation metrics during the validation and testing phases. These metrics are generated by the `RumorTweetsClassifer.get_metrics()` method and reported through the evaluation utility.

### Accuracy, Precision, Recall, and F1-Score

The primary classification metrics include:

- **Accuracy**: The fraction of correctly classified tweets, representing both rumors and non-rumors identified properly.
- **Precision**: The ratio of true-positive rumor predictions to all tweets predicted as rumors, measuring the model's exactness.
- **Recall**: The ratio of true-positive rumor predictions to all actual rumors in the dataset, measuring the model's completeness.
- **F1-score**: The harmonic mean of precision and recall, summarizing the balance between false positives and false negatives.

These four metrics are implemented in the `get_metrics()` method of the `RumorTweetsClassifer` class at line 55 of [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py).

### Loss Calculation

Beyond classification metrics, the system tracks **Loss** as the average cross-entropy loss per batch. This metric is added by the `training_util.evaluate` function at line 91 of [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py), providing a measure of model confidence and calibration during inference.

## Source Code Implementation

### Metric Computation in allennlp_rumor_classifier.py

In [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py), the `RumorTweetsClassifer.get_metrics()` method computes and returns the four primary metrics (accuracy, precision, recall, f1) after each evaluation pass. This method aggregates counts from the validation or test set to produce the final scores reported to the user.

### Evaluation Aggregation in training_util.py

The `evaluate` function in [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py) serves as the orchestration layer that calls the model's `get_metrics()` method and appends the loss calculation. This utility handles batch iteration and GPU/CPU device management while collecting metric values across the entire dataset.

## Running Model Evaluation

### Command-Line Evaluation

The primary entry point for model assessment is [`src/rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_evaluator.py). Execute evaluation from the terminal using:

```bash
python src/rumour_dnn_evaluator.py \
    -m path/to/trained/model_dir \
    -t path/to/test_set.csv \
    -g 0   # use GPU 0 (set -1 for CPU)

```

This script loads the trained model, processes the test set, and prints a metric summary:

```

accuracy: 0.8421
precision: 0.8012
recall: 0.7778
f1: 0.7893
loss: 0.4215

```

### Programmatic Evaluation

You can also evaluate the model programmatically using the underlying utilities:

```python
from allennlp_rumor_classifier import load_classifier_from_archive
from training_util import evaluate
from allennlp.data.iterators import BucketIterator
from allennlp.data.token_indexers import ELMoTokenCharactersIndexer
from allennlp_rumor_classifier import RumorTweetsDataReader

# Load model components

model, predictor = load_classifier_from_archive(
    vocab_dir_path="model_dir/vocabulary",
    model_weight_file="model_dir/weights_best.th",
    n_gpu_use=-1,
    max_cxt_size=200,
    feature_setting=1,
    global_means=global_means,
    global_stds=global_stds,
    attention_option=1
)

# Prepare data

token_indexer = ELMoTokenCharactersIndexer()
reader = RumorTweetsDataReader(token_indexers={"elmo": token_indexer})
instances = reader.read("data/test/twitter16_test_set.csv")

# Iterate and evaluate

iterator = BucketIterator(batch_size=128, sorting_keys=[("sentence", "num_tokens")])
iterator.index_with(model.vocab)

metrics = evaluate(model, instances, iterator, cuda_device=-1, batch_weight_key="")
print(metrics)   # {'accuracy': ..., 'precision': ..., 'recall': ..., 'f1': ..., 'loss': ...}

```

## Summary

- **RPDNN** uses five evaluation metrics: **Accuracy**, **Precision**, **Recall**, **F1-score**, and **Loss**.
- The `RumorTweetsClassifer.get_metrics()` method in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) computes classification metrics at line 55.
- The `training_util.evaluate` function in [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py) aggregates these metrics and adds loss calculation at line 91.
- Command-line evaluation is available via [`src/rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_evaluator.py) for batch testing.
- All metrics print simultaneously to provide a complete performance assessment after model inference.

## Frequently Asked Questions

### Does RPDNN use F1-score for evaluation?

Yes, the model explicitly calculates **F1-score** as one of its core evaluation metrics. The `get_metrics()` method in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) returns the F1-score alongside accuracy, precision, and recall to balance the trade-off between false positives and false negatives in rumour detection.

### What is the difference between accuracy and F1 in RPDNN?

**Accuracy** measures the overall fraction of correctly classified tweets regardless of class, while **F1-score** specifically balances precision and recall for the rumor class. In imbalanced datasets where rumors are rare compared to non-rumors, F1 provides a more reliable performance indicator than accuracy alone, as implemented in the model's metric reporting.

### How is the loss metric calculated during evaluation?

The **Loss** metric represents the average cross-entropy loss per batch during the evaluation pass. According to [`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py), the `evaluate` function accumulates loss values across all batches and computes the mean, providing insight into model confidence and prediction calibration beyond simple classification correctness.

### Can I evaluate the model on custom datasets using these metrics?

Yes, the evaluation pipeline supports custom datasets through both the CLI and programmatic interfaces. Pass your CSV file path to [`src/rumour_dnn_evaluator.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_evaluator.py) using the `-t` flag, or load instances programmatically using `RumorTweetsDataReader` with your custom data file to generate the full suite of evaluation metrics.