What Is the Difference Between Context Metadata and Context Content in RP-DNN?

In the RP-DNN (Rumour-Propagation Deep Neural Network) architecture, Context Metadata (CM) captures hand-crafted numerical features about who reacted and when, while Context Content (CC) captures dense semantic embeddings of what was actually said in reply or retweet text.

The RP-DNN model developed in the jerrygaolondon/rpdnn repository analyzes rumor propagation by splitting social context into two orthogonal signal types. Understanding the difference between context metadata and context content in RP-DNN is essential for configuring the model's feature extraction pipeline and interpreting its predictions.

Context Metadata (CM) vs Context Content (CC)

What Is Context Metadata (CM)?

Context Metadata (CM) is a hand-crafted numerical vector that encodes user-level, tweet-level, and temporal properties of reactions. Defined by the constant NUMERICAL_FEATURE_DIM in src/context_features_extractor.py, this vector contains 28 dimensions of engineered features:

  • User-profile attributes: Follower count, friend count, verification status, and account age proxies
  • Tweet-level indicators: Presence of URLs, hashtags, mentions, and sentiment-derived metrics
  • Temporal dynamics: Time difference between the reaction and the source tweet
  • Interaction flags: Binary indicators for "reply versus retweet" and "has profile description"

The function extract_social_numerical_features aggregates outputs from user_features_main (defined in src/context_features/user_features.py) and tweet_features_main (defined in src/context_features/tweet_features.py) to construct this vector. CM answers the question: who reacted and when?

What Is Context Content (CC)?

Context Content (CC) represents the semantic meaning of reaction text through dense embeddings. Unlike the engineered features of CM, CC captures what was actually said using a 1024-dimensional ELMo-based embedding (ELMO_EMBEDDING_DIM).

The extraction pipeline in context_feature_extraction_from_context_status tokenizes reply or retweet text and processes it through the fine-tuned ELMo model (fine_tuned_elmo) via sentence_embedding_elmo. Optionally, user profile descriptions can be concatenated to enrich the representation, creating a rich linguistic embedding of the conversational context.

Neural Encoding and Model Integration

In src/allennlp_rumor_classifier.py, these distinct feature types feed into specialized neural pathways:

Metadata Encoder (cxt_metadata_encoder) The 28-dimensional CM vectors pass through an LSTM or transformer that learns compact representations of user behavior and temporal patterns.

Content Encoder (cxt_content_encoder) The 1024-dimensional CC embeddings flow through a separate LSTM that captures sequential linguistic dependencies in the reaction text.

The model's batch_compute_context_feature_encoding method returns both tensors—cxt_content and cxt_metadata—along with their respective masks, allowing the architecture to process behavioral and semantic signals separately before fusion.

Configuration Modes: CM-Only, CC-Only, or Combined

The RP-DNN architecture supports four operational modes controlled by the feature_setting_option parameter in src/rumour_dnn_trainer.py:

  • feature_setting_option == 2 (CM-only): Disables the ELMo text pipeline; relies solely on numerical metadata with content tensors zero-filled
  • feature_setting_option == 3 (CC-only): Drops the 28-dimensional metadata vector; uses only semantic embeddings
  • feature_setting_option == 4 (CC + CM): Concatenates both representations before classification, giving the model access to both "what was said" and "who said it"
  • feature_setting_option == 1 (Baseline): Source tweet only, ignoring all context

Code Examples

Extracting CM and CC Features

from context_features_extractor import context_feature_extraction

# source_tweet_context_dataset: list of reply/retweet JSON objects

# source_tweet: source tweet JSON object

cxt_tensor = context_feature_extraction(
    source_tweet_context_dataset,
    source_tweet,
    context_feature_input_dim=20 * 3,  # EXPECTED_CONTEXT_INPUT_SIZE

)

print("Shape of the context feature tensor:", cxt_tensor.shape)

# Output: (batch_size, 60) where 60 = 20 max reactions × 3 (combined features)

The function internally calls context_feature_extraction_from_context_status, which returns user_profile_embedding (part of CC), reply_sent_embedding (the main CC), and numerical_features (CM).

Configuring CM-Only Mode


# In src/rumour_dnn_trainer.py

feature_setting_option = 2  # CM-only mode

model = RumourDNN(
    feature_setting=feature_setting_option,
    # ... additional parameters

)

When configured for metadata-only operation, the trainer prints a diagnostic message (lines 84-86 in src/rumour_dnn_trainer.py) confirming that textual embeddings are disabled.

Accessing Raw Tensors in the Forward Pass


# Inside src/allennlp_rumor_classifier.py

cxt_content, cxt_metadata, cxt_content_mask, cxt_metadata_mask = \
    self.batch_compute_context_feature_encoding(
        tweet_id, 
        maximum_cxt_sequence_length=self.max_cxt_size
    )

# cxt_metadata shape: (batch, max_cxt, 28)

# cxt_content shape: (batch, max_cxt, 1024)

When feature_setting == FEATURE_SETTING_OPTION_METADATA_ONLY, the cxt_content tensor contains zeros, and predictions rely exclusively on the metadata encoder.

Summary

  • Context Metadata (CM) provides a 28-dimensional vector of engineered numerical features describing user profiles, tweet attributes, and temporal dynamics—telling the model who reacted and when.
  • Context Content (CC) delivers 1024-dimensional ELMo embeddings of reaction text—telling the model what was actually said.
  • Both feature types are extracted in src/context_features_extractor.py and encoded via dedicated LSTM encoders (cxt_metadata_encoder and cxt_content_encoder) in src/allennlp_rumor_classifier.py.
  • The feature_setting_option flag controls whether the model uses CM only, CC only, both combined, or neither during training and evaluation.

Frequently Asked Questions

How are CM and CC features combined in the RP-DNN model?

When feature_setting_option is set to 4 (combined mode), the architecture concatenates the 28-dimensional Context Metadata vector with the 1024-dimensional Context Content embedding before feeding them into the downstream classifier. This concatenation happens within the batch_compute_context_feature_encoding method, allowing the model to simultaneously leverage behavioral signals and semantic content for rumor detection.

What is the dimensionality of CM and CC vectors in RP-DNN?

Context Metadata (CM) vectors have a fixed dimensionality of 28 as defined by the NUMERICAL_FEATURE_DIM constant. Context Content (CC) vectors have a dimensionality of 1024 (ELMO_EMBEDDING_DIM) resulting from the ELMo embedding pipeline in context_feature_extraction_from_context_status.

Can RP-DNN run using only metadata without text content?

Yes. Setting feature_setting_option = 2 enables CM-only mode, which disables the ELMo text processing pipeline and relies exclusively on the numerical metadata features extracted by extract_social_numerical_features. In this configuration, the content tensors are zero-filled and only the metadata encoder drives predictions, as implemented in src/allennlp_rumor_classifier.py.

Where are the user and tweet features for CM defined?

The numerical features composing CM are sourced from src/context_features/user_features.py (handling user-level attributes like followers and verification status) and src/context_features/tweet_features.py (handling tweet-level attributes like URLs and hashtags). These modules are aggregated by the main extraction logic in src/context_features_extractor.py to produce the final 28-dimensional vector.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →