RPDNN Source Tweet Encoding vs Social Context Encoding: Architecture and Implementation

In the RPDNN model, source tweet encoding captures the semantic meaning of the original claim using language models, while social context encoding aggregates propagation patterns from replies and retweets using sequential encoders; the architecture supports independent or combined usage via configurable feature settings.

The RPDNN (Rumor Prediction with Deep Neural Networks) repository implements a dual-encoder architecture that explicitly separates source tweet encoding versus social context encoding to detect rumors on social media. This design allows researchers to isolate the predictive power of the original tweet's text from the structural patterns of its propagation. Understanding how these two encoding pathways interact is essential for configuring the model for different rumor detection scenarios.

What is Source Tweet Encoding?

Source tweet encoding processes the textual content of the original ("source") tweet to generate a dense semantic representation. This pathway focuses exclusively on the claim itself, independent of how it spreads through the network.

Implementation in allennlp_rumor_classifier.py

In src/allennlp_rumor_classifier.py, the source tweet encoding is produced by passing tokenized text through the tweet_text_embedder and then through the language model encoder (self.lang_model_encoder). The output, stored in lang_model_encoder_out, represents the final source tweet vector.


# Lines 79-80 in allennlp_rumor_classifier.py

lang_model_encoder_out = self.lang_model_encoder(tweet_text_embedder_out, mask)

When the model operates in source-tweet-only mode (FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY), this vector becomes the sole input to the classifier:


# Lines 1012-1014

if self.feature_setting == FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY:
    rumor_representation = lang_model_encoder_out

What is Social Context Encoding?

Social context encoding aggregates information from the tweet's propagation history, including replies, retweets, and associated user metadata. This pathway captures how the claim spreads through the social network rather than what the claim states.

Context Feature Extraction

The raw propagation data is first processed by batch_compute_context_feature_encoding (called at lines 84-86), which extracts structured tensors from the surrounding social activity. This function interfaces with src/context_features_extractor.py to convert JSON propagation trees into numerical features.

Transformer and LSTM Variants

Depending on the configuration, the social context encoder uses either a Transformer architecture or an LSTM to process the sequential propagation data:

  • Transformer path: social_cxt_encoding_with_transformer (lines 887-891) applies self-attention mechanisms to the context features.
  • LSTM path: social_cxt_encoding_with_lstm (lines 98-102) processes the context sequence with recurrent layers.

Both paths output social_context_encoder_out, a fixed-size vector representing the propagation pattern.


# Lines 98-102 (LSTM variant example)

social_context_encoder_out = self.social_cxt_encoding_with_lstm(
    context_content_features_encoding,
    context_metadata_features_encoding,
    context_mask
)

How the Model Combines Both Encodings

The RPDNN architecture provides configurable feature settings that determine how source tweet and social context encodings interact. This modularity allows ablation studies to measure the independent contribution of each signal.

Feature Configuration Options

The self.feature_setting attribute controls the encoding combination strategy using constants defined at the top of allennlp_rumor_classifier.py:

  • FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY: Uses only lang_model_encoder_out
  • FEATURE_SETTING_OPTION_SOCIAL_CONTEXT_ONLY: Uses only social_context_encoder_out
  • FEATURE_SETTING_OPTION_FULL: Concatenates both vectors

Concatenation Logic

When running in full mode, the model concatenates the source tweet encoding and social context encoding along the feature dimension to create a comprehensive rumor representation:


# Lines 1015-1017

elif self.feature_setting == FEATURE_SETTING_OPTION_FULL:
    rumor_representation = torch.cat((lang_model_encoder_out, social_context_encoder_out), dim=1)

This concatenated vector is then passed to the classifier_feedforward network for final rumor classification.

Key Files and Their Roles

Understanding the separation of concerns in the RPDNN codebase requires familiarity with these specific files:

  • src/allennlp_rumor_classifier.py: Core model implementation containing both encoders, the feature-setting logic, and the concatenation mechanism (lines 79-102, 887-891, 1010-1017).

  • src/data_loader.py: Handles loading of source tweet JSON files via load_source_tweet_json and load_source_tweet_context, feeding raw text into the source tweet encoder.

  • src/context_features_extractor.py: Extracts numerical features from propagation trees and user metadata, creating the input tensors for the social context encoder.

  • src/training_util.py: Provides training utilities such as gradient clipping (global_norm) that support the optimization of both encoding pathways.

Summary

The RPDNN model explicitly separates source tweet encoding from social context encoding to enable granular analysis of rumor signals:

  • Source tweet encoding generates semantic representations of the original claim using language models (lang_model_encoder_out), suitable for text-only rumor detection.

  • Social context encoding aggregates propagation patterns from replies and retweets using Transformer or LSTM architectures (social_context_encoder_out), capturing how the rumor spreads through the network.

  • The architecture supports configurable combinations via feature_setting flags, allowing researchers to use either encoding alone or concatenate both for the full model (torch.cat at lines 1015-1017).

This modular design facilitates ablation studies and allows the model to adapt to scenarios where either the source text or the social context may be missing or unreliable.

Frequently Asked Questions

Can RPDNN work with only the source tweet text without any propagation data?

Yes. By setting model.feature_setting = FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY, the model uses only the lang_model_encoder_out vector derived from the source tweet text. In this configuration, the social context encoder is not utilized, making the model functional even when reply and retweet data is unavailable.

What types of social context features does the model encode?

The social context encoder processes features extracted from the tweet's propagation tree, including reply content, retweet patterns, and user metadata. These features are generated by batch_compute_context_feature_encoding and context_features_extractor.py, then processed through either social_cxt_encoding_with_transformer (self-attention) or social_cxt_encoding_with_lstm (recurrent layers) to produce the social_context_encoder_out vector.

How does the full model combine both encodings for classification?

In the full configuration (FEATURE_SETTING_OPTION_FULL), the model concatenates the source tweet encoding and social context encoding along dimension 1 using torch.cat((lang_model_encoder_out, social_context_encoder_out), dim=1) (lines 1015-1017 in allennlp_rumor_classifier.py). This combined vector is then passed to the classifier_feedforward network to produce the final rumor prediction logits.

Which encoder typically contributes more to rumor detection accuracy?

According to the modular design of RPDNN, the relative contribution depends on the dataset and rumor type. The source tweet encoder captures semantic suspiciousness of the claim itself, while the social context encoder captures anomalous propagation patterns known to correlate with misinformation. The feature_setting flags allow researchers to conduct ablation studies to quantify each encoder's contribution for specific tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →