# RPDNN Source Tweet Encoding vs Social Context Encoding: Architecture and Implementation

> Explore RPDNN source tweet encoding and social context encoding. Understand how semantic meaning and propagation patterns are captured and combined in this architecture.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: deep-dive
- Published: 2026-03-04

---

**In the RPDNN model, source tweet encoding captures the semantic meaning of the original claim using language models, while social context encoding aggregates propagation patterns from replies and retweets using sequential encoders; the architecture supports independent or combined usage via configurable feature settings.**

The RPDNN (Rumor Prediction with Deep Neural Networks) repository implements a dual-encoder architecture that explicitly separates **source tweet encoding versus social context encoding** to detect rumors on social media. This design allows researchers to isolate the predictive power of the original tweet's text from the structural patterns of its propagation. Understanding how these two encoding pathways interact is essential for configuring the model for different rumor detection scenarios.

## What is Source Tweet Encoding?

Source tweet encoding processes the textual content of the original ("source") tweet to generate a dense semantic representation. This pathway focuses exclusively on the claim itself, independent of how it spreads through the network.

### Implementation in allennlp_rumor_classifier.py

In [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py), the source tweet encoding is produced by passing tokenized text through the `tweet_text_embedder` and then through the language model encoder (`self.lang_model_encoder`). The output, stored in `lang_model_encoder_out`, represents the final source tweet vector.

```python

# Lines 79-80 in allennlp_rumor_classifier.py

lang_model_encoder_out = self.lang_model_encoder(tweet_text_embedder_out, mask)

```

When the model operates in **source-tweet-only** mode (`FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY`), this vector becomes the sole input to the classifier:

```python

# Lines 1012-1014

if self.feature_setting == FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY:
    rumor_representation = lang_model_encoder_out

```

## What is Social Context Encoding?

Social context encoding aggregates information from the tweet's propagation history, including replies, retweets, and associated user metadata. This pathway captures how the claim spreads through the social network rather than what the claim states.

### Context Feature Extraction

The raw propagation data is first processed by `batch_compute_context_feature_encoding` (called at lines 84-86), which extracts structured tensors from the surrounding social activity. This function interfaces with [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) to convert JSON propagation trees into numerical features.

### Transformer and LSTM Variants

Depending on the configuration, the social context encoder uses either a Transformer architecture or an LSTM to process the sequential propagation data:

- **Transformer path**: `social_cxt_encoding_with_transformer` (lines 887-891) applies self-attention mechanisms to the context features.
- **LSTM path**: `social_cxt_encoding_with_lstm` (lines 98-102) processes the context sequence with recurrent layers.

Both paths output `social_context_encoder_out`, a fixed-size vector representing the propagation pattern.

```python

# Lines 98-102 (LSTM variant example)

social_context_encoder_out = self.social_cxt_encoding_with_lstm(
    context_content_features_encoding,
    context_metadata_features_encoding,
    context_mask
)

```

## How the Model Combines Both Encodings

The RPDNN architecture provides configurable feature settings that determine how source tweet and social context encodings interact. This modularity allows ablation studies to measure the independent contribution of each signal.

### Feature Configuration Options

The `self.feature_setting` attribute controls the encoding combination strategy using constants defined at the top of [`allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/allennlp_rumor_classifier.py):

- `FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY`: Uses only `lang_model_encoder_out`
- `FEATURE_SETTING_OPTION_SOCIAL_CONTEXT_ONLY`: Uses only `social_context_encoder_out`
- `FEATURE_SETTING_OPTION_FULL`: Concatenates both vectors

### Concatenation Logic

When running in **full** mode, the model concatenates the source tweet encoding and social context encoding along the feature dimension to create a comprehensive rumor representation:

```python

# Lines 1015-1017

elif self.feature_setting == FEATURE_SETTING_OPTION_FULL:
    rumor_representation = torch.cat((lang_model_encoder_out, social_context_encoder_out), dim=1)

```

This concatenated vector is then passed to the `classifier_feedforward` network for final rumor classification.

## Key Files and Their Roles

Understanding the separation of concerns in the RPDNN codebase requires familiarity with these specific files:

- **[`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py)**: Core model implementation containing both encoders, the feature-setting logic, and the concatenation mechanism (lines 79-102, 887-891, 1010-1017).

- **[`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py)**: Handles loading of source tweet JSON files via `load_source_tweet_json` and `load_source_tweet_context`, feeding raw text into the source tweet encoder.

- **[`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py)**: Extracts numerical features from propagation trees and user metadata, creating the input tensors for the social context encoder.

- **[`src/training_util.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/training_util.py)**: Provides training utilities such as gradient clipping (`global_norm`) that support the optimization of both encoding pathways.

## Summary

The RPDNN model explicitly separates **source tweet encoding** from **social context encoding** to enable granular analysis of rumor signals:

- **Source tweet encoding** generates semantic representations of the original claim using language models (`lang_model_encoder_out`), suitable for text-only rumor detection.

- **Social context encoding** aggregates propagation patterns from replies and retweets using Transformer or LSTM architectures (`social_context_encoder_out`), capturing how the rumor spreads through the network.

- The architecture supports **configurable combinations** via `feature_setting` flags, allowing researchers to use either encoding alone or concatenate both for the full model (`torch.cat` at lines 1015-1017).

This modular design facilitates ablation studies and allows the model to adapt to scenarios where either the source text or the social context may be missing or unreliable.

## Frequently Asked Questions

### Can RPDNN work with only the source tweet text without any propagation data?

Yes. By setting `model.feature_setting = FEATURE_SETTING_OPTION_SOURCE_TWEET_CONTENT_ONLY`, the model uses only the `lang_model_encoder_out` vector derived from the source tweet text. In this configuration, the social context encoder is not utilized, making the model functional even when reply and retweet data is unavailable.

### What types of social context features does the model encode?

The social context encoder processes features extracted from the tweet's propagation tree, including reply content, retweet patterns, and user metadata. These features are generated by `batch_compute_context_feature_encoding` and [`context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/context_features_extractor.py), then processed through either `social_cxt_encoding_with_transformer` (self-attention) or `social_cxt_encoding_with_lstm` (recurrent layers) to produce the `social_context_encoder_out` vector.

### How does the full model combine both encodings for classification?

In the full configuration (`FEATURE_SETTING_OPTION_FULL`), the model concatenates the source tweet encoding and social context encoding along dimension 1 using `torch.cat((lang_model_encoder_out, social_context_encoder_out), dim=1)` (lines 1015-1017 in [`allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/allennlp_rumor_classifier.py)). This combined vector is then passed to the `classifier_feedforward` network to produce the final rumor prediction logits.

### Which encoder typically contributes more to rumor detection accuracy?

According to the modular design of RPDNN, the relative contribution depends on the dataset and rumor type. The source tweet encoder captures semantic suspiciousness of the claim itself, while the social context encoder captures anomalous propagation patterns known to correlate with misinformation. The `feature_setting` flags allow researchers to conduct ablation studies to quantify each encoder's contribution for specific tasks.