RPDNN Evaluation Metrics: F1, Accuracy, Precision, and Recall Explained

The RPDNN (Rumour DNN) model evaluates performance using Accuracy, Precision, Recall, F1-score, and cross-entropy Loss, calculated via the get_metrics() method in src/allennlp_rumor_classifier.py and aggregated by training_util.evaluate in src/training_util.py.

The jerrygaolondon/rpdnn repository implements a deep neural network for rumour detection on Twitter. Understanding how this model assesses its predictions requires examining the specific evaluation metrics built into its architecture. The system tracks both classification accuracy and nuanced performance indicators like F1-score to provide a complete picture of model effectiveness on rumour classification tasks.

Core Evaluation Metrics in RPDNN

The model produces five key evaluation metrics during the validation and testing phases. These metrics are generated by the RumorTweetsClassifer.get_metrics() method and reported through the evaluation utility.

Accuracy, Precision, Recall, and F1-Score

The primary classification metrics include:

  • Accuracy: The fraction of correctly classified tweets, representing both rumors and non-rumors identified properly.
  • Precision: The ratio of true-positive rumor predictions to all tweets predicted as rumors, measuring the model's exactness.
  • Recall: The ratio of true-positive rumor predictions to all actual rumors in the dataset, measuring the model's completeness.
  • F1-score: The harmonic mean of precision and recall, summarizing the balance between false positives and false negatives.

These four metrics are implemented in the get_metrics() method of the RumorTweetsClassifer class at line 55 of src/allennlp_rumor_classifier.py.

Loss Calculation

Beyond classification metrics, the system tracks Loss as the average cross-entropy loss per batch. This metric is added by the training_util.evaluate function at line 91 of src/training_util.py, providing a measure of model confidence and calibration during inference.

Source Code Implementation

Metric Computation in allennlp_rumor_classifier.py

In src/allennlp_rumor_classifier.py, the RumorTweetsClassifer.get_metrics() method computes and returns the four primary metrics (accuracy, precision, recall, f1) after each evaluation pass. This method aggregates counts from the validation or test set to produce the final scores reported to the user.

Evaluation Aggregation in training_util.py

The evaluate function in src/training_util.py serves as the orchestration layer that calls the model's get_metrics() method and appends the loss calculation. This utility handles batch iteration and GPU/CPU device management while collecting metric values across the entire dataset.

Running Model Evaluation

Command-Line Evaluation

The primary entry point for model assessment is src/rumour_dnn_evaluator.py. Execute evaluation from the terminal using:

python src/rumour_dnn_evaluator.py \
    -m path/to/trained/model_dir \
    -t path/to/test_set.csv \
    -g 0   # use GPU 0 (set -1 for CPU)

This script loads the trained model, processes the test set, and prints a metric summary:


accuracy: 0.8421
precision: 0.8012
recall: 0.7778
f1: 0.7893
loss: 0.4215

Programmatic Evaluation

You can also evaluate the model programmatically using the underlying utilities:

from allennlp_rumor_classifier import load_classifier_from_archive
from training_util import evaluate
from allennlp.data.iterators import BucketIterator
from allennlp.data.token_indexers import ELMoTokenCharactersIndexer
from allennlp_rumor_classifier import RumorTweetsDataReader

# Load model components

model, predictor = load_classifier_from_archive(
    vocab_dir_path="model_dir/vocabulary",
    model_weight_file="model_dir/weights_best.th",
    n_gpu_use=-1,
    max_cxt_size=200,
    feature_setting=1,
    global_means=global_means,
    global_stds=global_stds,
    attention_option=1
)

# Prepare data

token_indexer = ELMoTokenCharactersIndexer()
reader = RumorTweetsDataReader(token_indexers={"elmo": token_indexer})
instances = reader.read("data/test/twitter16_test_set.csv")

# Iterate and evaluate

iterator = BucketIterator(batch_size=128, sorting_keys=[("sentence", "num_tokens")])
iterator.index_with(model.vocab)

metrics = evaluate(model, instances, iterator, cuda_device=-1, batch_weight_key="")
print(metrics)   # {'accuracy': ..., 'precision': ..., 'recall': ..., 'f1': ..., 'loss': ...}

Summary

  • RPDNN uses five evaluation metrics: Accuracy, Precision, Recall, F1-score, and Loss.
  • The RumorTweetsClassifer.get_metrics() method in src/allennlp_rumor_classifier.py computes classification metrics at line 55.
  • The training_util.evaluate function in src/training_util.py aggregates these metrics and adds loss calculation at line 91.
  • Command-line evaluation is available via src/rumour_dnn_evaluator.py for batch testing.
  • All metrics print simultaneously to provide a complete performance assessment after model inference.

Frequently Asked Questions

Does RPDNN use F1-score for evaluation?

Yes, the model explicitly calculates F1-score as one of its core evaluation metrics. The get_metrics() method in src/allennlp_rumor_classifier.py returns the F1-score alongside accuracy, precision, and recall to balance the trade-off between false positives and false negatives in rumour detection.

What is the difference between accuracy and F1 in RPDNN?

Accuracy measures the overall fraction of correctly classified tweets regardless of class, while F1-score specifically balances precision and recall for the rumor class. In imbalanced datasets where rumors are rare compared to non-rumors, F1 provides a more reliable performance indicator than accuracy alone, as implemented in the model's metric reporting.

How is the loss metric calculated during evaluation?

The Loss metric represents the average cross-entropy loss per batch during the evaluation pass. According to src/training_util.py, the evaluate function accumulates loss values across all batches and computes the mean, providing insight into model confidence and prediction calibration beyond simple classification correctness.

Can I evaluate the model on custom datasets using these metrics?

Yes, the evaluation pipeline supports custom datasets through both the CLI and programmatic interfaces. Pass your CSV file path to src/rumour_dnn_evaluator.py using the -t flag, or load instances programmatically using RumorTweetsDataReader with your custom data file to generate the full suite of evaluation metrics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →