How to Add Custom Context Features to the context_features_extractor Module in RPDNN
To add custom context features to the context_features_extractor module, extend the extract_social_numerical_features function in src/context_features_extractor.py, increment NUMERICAL_FEATURE_DIM, and verify the output shape matches EXPECTED_CONTEXT_INPUT_SIZE.
The context_features_extractor module in the jerrygaolondon/rpdnn repository generates the numeric representation of reaction tweets—replies and retweets—that feeds into the Rumour Detection RNN. Adding custom context features allows you to enrich this representation with new measurable properties, such as hashtag counts or sentiment scores, by extending the core extraction pipeline. This guide identifies the exact source code locations, constants, and validation steps required to safely integrate new feature dimensions.
Understanding the Context Feature Extraction Pipeline
Before modifying code, understand the three-stage architecture that processes each reaction tweet:
- Numeric Feature Extraction: The
extract_social_numerical_featuresfunction insrc/context_features_extractor.pycomputes a flat numeric vector from user-profile, tweet-level, and temporal attributes. - Embedding Concatenation:
context_feature_extraction_from_context_statuscalls the numeric extractor and optionally appends ELMo embeddings for user profiles and reply content viasentence_embedding_elmo. - Batch Stacking:
context_feature_extractionloops over all context tweets and stacks per-reaction vectors into a matrix of shape(N, EXPECTED_CONTEXT_INPUT_SIZE)for the downstream RNN.
To add a custom feature, you must extend the numeric vector produced in stage one and ensure the dimensionality constants reflect the change.
Step-by-Step Implementation Guide
Step 1: Define Your Custom Feature
Select a measurable property of a reaction tweet that can be computed solely from the reaction_status_json object passed to the extractor. Examples include the number of hashtags, sentiment polarity, or domain-specific flags. Avoid features requiring external I/O during extraction.
Step 2: Implement the Feature in extract_social_numerical_features
Open src/context_features_extractor.py and locate the extract_social_numerical_features function (lines 45–61). Append your custom logic after the existing user_features and tweet_features calculations:
def extract_social_numerical_features(reaction_status_json, source_tweet_user_id,
source_tweet_user_screen_name, source_text):
# Existing feature extraction
user_features = user_features_main(reaction_status_json, source_tweet_user_id)
tweet_features = tweet_features_main(reaction_status_json,
source_tweet_user_screen_name,
source_text)
# Build base numerical vector
numerical_features = []
numerical_features.extend((user_features + tweet_features))
# 👇 Custom feature: count hashtags in the reaction tweet
hashtags = reaction_status_json.get('entities', {}).get('hashtags', [])
num_hashtags = len(hashtags)
numerical_features.append(num_hashtags)
# Append temporal and meta features (preserve existing order)
t_diff = calculate_time_diff(...) # existing variable
c_type = determine_context_type(...) # existing variable
has_description = check_description(...) # existing variable
numerical_features.extend([t_diff, c_type, has_description])
return np.array(numerical_features)
This ensures every reaction processed by context_feature_extraction includes your new numeric cue.
Step 3: Update Dimensionality Constants
Two module-level constants in src/context_features_extractor.py control the tensor shape:
NUMERICAL_FEATURE_DIM: Defines the length of the flat numeric vector (default is28)EXPECTED_CONTEXT_INPUT_SIZE: Defines the total width expected by the RNN (default is20 * 3)
If you add a single numeric field, increment NUMERICAL_FEATURE_DIM accordingly:
# src/context_features_extractor.py
NUMERICAL_FEATURE_DIM: int = 29 # Increased from 28 to include hashtag count
EXPECTED_CONTEXT_INPUT_SIZE: int = 20 * 3 # Adjust only if changing total embedding size
Search the repository for hard-coded slices or references to the old dimension (e.g., [..., :28]) and update them to prevent shape mismatches downstream.
Step 4: Validate with the Test Harness
Run the built-in validation to confirm the new feature integrates without breaking the expected tensor shape:
from src.context_features_extractor import test_context_features_shape_by_tweet_id
# Verify shape matches EXPECTED_CONTEXT_INPUT_SIZE
test_context_features_shape_by_tweet_id("500350777844457473")
If the assertion passes, your custom feature is correctly integrated. If it fails, check that NUMERICAL_FEATURE_DIM matches the actual length of your modified numerical_features array.
Optional: Extending User and Tweet Feature Modules
For features that fit naturally into user-profile or tweet-level categorizations, consider extending the helper modules instead of inline code:
src/context_features/user_features.py: Add user-specific custom cues (e.g., account age buckets) by modifyinguser_features_main.src/context_features/tweet_features.py: Add tweet-specific cues (e.g., media attachment flags) by modifyingtweet_features_main.
Return the new values as additional list elements, then handle the concatenation in extract_social_numerical_features as shown in Step 2.
Summary
- Inject new numeric cues exclusively in
extract_social_numerical_featureswithinsrc/context_features_extractor.py. - Synchronize dimensionality constants by incrementing
NUMERICAL_FEATURE_DIMto match your new vector length. - Avoid shape errors by grepping for hard-coded dimension references (e.g.,
NUMERICAL_FEATURE_DIMor28) across the codebase. - Validate changes using
test_context_features_shape_by_tweet_idto ensure the RNN receives tensors of the expected width. - Document additions with inline comments explaining feature semantics for future maintainers.
Frequently Asked Questions
Where exactly do I modify the code to add a custom numeric feature?
Add your computation logic inside the extract_social_numerical_features function in src/context_features_extractor.py, specifically after the existing user_features and tweet_features lists are combined but before the final conversion to a NumPy array (around line 55). This is the only location where the flat numeric vector is constructed.
What happens if I forget to update NUMERICAL_FEATURE_DIM?
If you add features to the vector but leave NUMERICAL_FEATURE_DIM at its default value (28), downstream components that rely on this constant for padding, slicing, or shape validation will encounter dimension mismatches. The RNN may raise tensor shape errors during training or inference because EXPECTED_CONTEXT_INPUT_SIZE will no longer align with the actual data width.
Can I add features that require external API calls or database lookups?
No. The context_features_extractor is designed to operate solely on the reaction_status_json dictionary passed to it. External I/O would significantly slow down the batch processing loop in context_feature_extraction. Pre-compute external data and embed it into the JSON before extraction, or add lookup logic upstream in your data pipeline.
How do I test that my custom feature is working correctly?
Use the test_context_features_shape_by_tweet_id function provided in src/context_features_extractor.py. Pass a valid tweet ID string to this function; it will extract context features for that tweet and assert that the output matrix shape matches EXPECTED_CONTEXT_INPUT_SIZE. If the assertion passes, your dimensionality updates are correct.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →