# What Is the Difference Between Context Metadata and Context Content in RP-DNN?

> Understand the distinction between context metadata and context content in RP-DNN. Learn how CM captures user reaction data and CC represents the text's semantic meaning.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: deep-dive
- Published: 2026-03-04

---

**In the RP-DNN (Rumour-Propagation Deep Neural Network) architecture, Context Metadata (CM) captures hand-crafted numerical features about who reacted and when, while Context Content (CC) captures dense semantic embeddings of what was actually said in reply or retweet text.**

The RP-DNN model developed in the [jerrygaolondon/rpdnn](https://github.com/jerrygaolondon/rpdnn) repository analyzes rumor propagation by splitting social context into two orthogonal signal types. Understanding the difference between context metadata and context content in RP-DNN is essential for configuring the model's feature extraction pipeline and interpreting its predictions.

## Context Metadata (CM) vs Context Content (CC)

### What Is Context Metadata (CM)?

**Context Metadata (CM)** is a hand-crafted numerical vector that encodes user-level, tweet-level, and temporal properties of reactions. Defined by the constant `NUMERICAL_FEATURE_DIM` in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py), this vector contains **28 dimensions** of engineered features:

- **User-profile attributes**: Follower count, friend count, verification status, and account age proxies
- **Tweet-level indicators**: Presence of URLs, hashtags, mentions, and sentiment-derived metrics
- **Temporal dynamics**: Time difference between the reaction and the source tweet
- **Interaction flags**: Binary indicators for "reply versus retweet" and "has profile description"

The function `extract_social_numerical_features` aggregates outputs from `user_features_main` (defined in [`src/context_features/user_features.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features/user_features.py)) and `tweet_features_main` (defined in [`src/context_features/tweet_features.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features/tweet_features.py)) to construct this vector. CM answers the question: **who** reacted and **when**?

### What Is Context Content (CC)?

**Context Content (CC)** represents the semantic meaning of reaction text through dense embeddings. Unlike the engineered features of CM, CC captures **what** was actually said using a **1024-dimensional** ELMo-based embedding (`ELMO_EMBEDDING_DIM`).

The extraction pipeline in `context_feature_extraction_from_context_status` tokenizes reply or retweet text and processes it through the fine-tuned ELMo model (`fine_tuned_elmo`) via `sentence_embedding_elmo`. Optionally, user profile descriptions can be concatenated to enrich the representation, creating a rich linguistic embedding of the conversational context.

## Neural Encoding and Model Integration

In [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py), these distinct feature types feed into specialized neural pathways:

**Metadata Encoder (`cxt_metadata_encoder`)**
The 28-dimensional CM vectors pass through an LSTM or transformer that learns compact representations of user behavior and temporal patterns.

**Content Encoder (`cxt_content_encoder`)**
The 1024-dimensional CC embeddings flow through a separate LSTM that captures sequential linguistic dependencies in the reaction text.

The model's `batch_compute_context_feature_encoding` method returns both tensors—`cxt_content` and `cxt_metadata`—along with their respective masks, allowing the architecture to process behavioral and semantic signals separately before fusion.

## Configuration Modes: CM-Only, CC-Only, or Combined

The RP-DNN architecture supports four operational modes controlled by the `feature_setting_option` parameter in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py):

- **`feature_setting_option == 2` (CM-only)**: Disables the ELMo text pipeline; relies solely on numerical metadata with content tensors zero-filled
- **`feature_setting_option == 3` (CC-only)**: Drops the 28-dimensional metadata vector; uses only semantic embeddings
- **`feature_setting_option == 4` (CC + CM)**: Concatenates both representations before classification, giving the model access to both "what was said" and "who said it"
- **`feature_setting_option == 1` (Baseline)**: Source tweet only, ignoring all context

## Code Examples

### Extracting CM and CC Features

```python
from context_features_extractor import context_feature_extraction

# source_tweet_context_dataset: list of reply/retweet JSON objects

# source_tweet: source tweet JSON object

cxt_tensor = context_feature_extraction(
    source_tweet_context_dataset,
    source_tweet,
    context_feature_input_dim=20 * 3,  # EXPECTED_CONTEXT_INPUT_SIZE

)

print("Shape of the context feature tensor:", cxt_tensor.shape)

# Output: (batch_size, 60) where 60 = 20 max reactions × 3 (combined features)

```

The function internally calls `context_feature_extraction_from_context_status`, which returns `user_profile_embedding` (part of CC), `reply_sent_embedding` (the main CC), and `numerical_features` (CM).

### Configuring CM-Only Mode

```python

# In src/rumour_dnn_trainer.py

feature_setting_option = 2  # CM-only mode

model = RumourDNN(
    feature_setting=feature_setting_option,
    # ... additional parameters

)

```

When configured for metadata-only operation, the trainer prints a diagnostic message (lines 84-86 in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py)) confirming that textual embeddings are disabled.

### Accessing Raw Tensors in the Forward Pass

```python

# Inside src/allennlp_rumor_classifier.py

cxt_content, cxt_metadata, cxt_content_mask, cxt_metadata_mask = \
    self.batch_compute_context_feature_encoding(
        tweet_id, 
        maximum_cxt_sequence_length=self.max_cxt_size
    )

# cxt_metadata shape: (batch, max_cxt, 28)

# cxt_content shape: (batch, max_cxt, 1024)

```

When `feature_setting == FEATURE_SETTING_OPTION_METADATA_ONLY`, the `cxt_content` tensor contains zeros, and predictions rely exclusively on the metadata encoder.

## Summary

- **Context Metadata (CM)** provides a 28-dimensional vector of engineered numerical features describing user profiles, tweet attributes, and temporal dynamics—telling the model *who* reacted and *when*.
- **Context Content (CC)** delivers 1024-dimensional ELMo embeddings of reaction text—telling the model *what* was actually said.
- Both feature types are extracted in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) and encoded via dedicated LSTM encoders (`cxt_metadata_encoder` and `cxt_content_encoder`) in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py).
- The `feature_setting_option` flag controls whether the model uses CM only, CC only, both combined, or neither during training and evaluation.

## Frequently Asked Questions

### How are CM and CC features combined in the RP-DNN model?

When `feature_setting_option` is set to `4` (combined mode), the architecture concatenates the 28-dimensional Context Metadata vector with the 1024-dimensional Context Content embedding before feeding them into the downstream classifier. This concatenation happens within the `batch_compute_context_feature_encoding` method, allowing the model to simultaneously leverage behavioral signals and semantic content for rumor detection.

### What is the dimensionality of CM and CC vectors in RP-DNN?

Context Metadata (CM) vectors have a fixed dimensionality of **28** as defined by the `NUMERICAL_FEATURE_DIM` constant. Context Content (CC) vectors have a dimensionality of **1024** (`ELMO_EMBEDDING_DIM`) resulting from the ELMo embedding pipeline in `context_feature_extraction_from_context_status`.

### Can RP-DNN run using only metadata without text content?

Yes. Setting `feature_setting_option = 2` enables CM-only mode, which disables the ELMo text processing pipeline and relies exclusively on the numerical metadata features extracted by `extract_social_numerical_features`. In this configuration, the content tensors are zero-filled and only the metadata encoder drives predictions, as implemented in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py).

### Where are the user and tweet features for CM defined?

The numerical features composing CM are sourced from [`src/context_features/user_features.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features/user_features.py) (handling user-level attributes like followers and verification status) and [`src/context_features/tweet_features.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features/tweet_features.py) (handling tweet-level attributes like URLs and hashtags). These modules are aggregated by the main extraction logic in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) to produce the final 28-dimensional vector.