How to Add Custom Context Features to the context_features_extractor Module in RPDNN

To add custom context features to the context_features_extractor module, extend the extract_social_numerical_features function in src/context_features_extractor.py, increment NUMERICAL_FEATURE_DIM, and verify the output shape matches EXPECTED_CONTEXT_INPUT_SIZE.

The context_features_extractor module in the jerrygaolondon/rpdnn repository generates the numeric representation of reaction tweets—replies and retweets—that feeds into the Rumour Detection RNN. Adding custom context features allows you to enrich this representation with new measurable properties, such as hashtag counts or sentiment scores, by extending the core extraction pipeline. This guide identifies the exact source code locations, constants, and validation steps required to safely integrate new feature dimensions.

Understanding the Context Feature Extraction Pipeline

Before modifying code, understand the three-stage architecture that processes each reaction tweet:

  1. Numeric Feature Extraction: The extract_social_numerical_features function in src/context_features_extractor.py computes a flat numeric vector from user-profile, tweet-level, and temporal attributes.
  2. Embedding Concatenation: context_feature_extraction_from_context_status calls the numeric extractor and optionally appends ELMo embeddings for user profiles and reply content via sentence_embedding_elmo.
  3. Batch Stacking: context_feature_extraction loops over all context tweets and stacks per-reaction vectors into a matrix of shape (N, EXPECTED_CONTEXT_INPUT_SIZE) for the downstream RNN.

To add a custom feature, you must extend the numeric vector produced in stage one and ensure the dimensionality constants reflect the change.

Step-by-Step Implementation Guide

Step 1: Define Your Custom Feature

Select a measurable property of a reaction tweet that can be computed solely from the reaction_status_json object passed to the extractor. Examples include the number of hashtags, sentiment polarity, or domain-specific flags. Avoid features requiring external I/O during extraction.

Step 2: Implement the Feature in extract_social_numerical_features

Open src/context_features_extractor.py and locate the extract_social_numerical_features function (lines 45–61). Append your custom logic after the existing user_features and tweet_features calculations:

def extract_social_numerical_features(reaction_status_json, source_tweet_user_id, 
                                      source_tweet_user_screen_name, source_text):
    # Existing feature extraction

    user_features = user_features_main(reaction_status_json, source_tweet_user_id)
    tweet_features = tweet_features_main(reaction_status_json,
                                         source_tweet_user_screen_name,
                                         source_text)
    
    # Build base numerical vector

    numerical_features = []
    numerical_features.extend((user_features + tweet_features))
    
    # 👇 Custom feature: count hashtags in the reaction tweet

    hashtags = reaction_status_json.get('entities', {}).get('hashtags', [])
    num_hashtags = len(hashtags)
    numerical_features.append(num_hashtags)
    
    # Append temporal and meta features (preserve existing order)

    t_diff = calculate_time_diff(...)  # existing variable

    c_type = determine_context_type(...)  # existing variable

    has_description = check_description(...)  # existing variable

    numerical_features.extend([t_diff, c_type, has_description])
    
    return np.array(numerical_features)

This ensures every reaction processed by context_feature_extraction includes your new numeric cue.

Step 3: Update Dimensionality Constants

Two module-level constants in src/context_features_extractor.py control the tensor shape:

  • NUMERICAL_FEATURE_DIM: Defines the length of the flat numeric vector (default is 28)
  • EXPECTED_CONTEXT_INPUT_SIZE: Defines the total width expected by the RNN (default is 20 * 3)

If you add a single numeric field, increment NUMERICAL_FEATURE_DIM accordingly:


# src/context_features_extractor.py

NUMERICAL_FEATURE_DIM: int = 29   # Increased from 28 to include hashtag count

EXPECTED_CONTEXT_INPUT_SIZE: int = 20 * 3  # Adjust only if changing total embedding size

Search the repository for hard-coded slices or references to the old dimension (e.g., [..., :28]) and update them to prevent shape mismatches downstream.

Step 4: Validate with the Test Harness

Run the built-in validation to confirm the new feature integrates without breaking the expected tensor shape:

from src.context_features_extractor import test_context_features_shape_by_tweet_id

# Verify shape matches EXPECTED_CONTEXT_INPUT_SIZE

test_context_features_shape_by_tweet_id("500350777844457473")

If the assertion passes, your custom feature is correctly integrated. If it fails, check that NUMERICAL_FEATURE_DIM matches the actual length of your modified numerical_features array.

Optional: Extending User and Tweet Feature Modules

For features that fit naturally into user-profile or tweet-level categorizations, consider extending the helper modules instead of inline code:

Return the new values as additional list elements, then handle the concatenation in extract_social_numerical_features as shown in Step 2.

Summary

  • Inject new numeric cues exclusively in extract_social_numerical_features within src/context_features_extractor.py.
  • Synchronize dimensionality constants by incrementing NUMERICAL_FEATURE_DIM to match your new vector length.
  • Avoid shape errors by grepping for hard-coded dimension references (e.g., NUMERICAL_FEATURE_DIM or 28) across the codebase.
  • Validate changes using test_context_features_shape_by_tweet_id to ensure the RNN receives tensors of the expected width.
  • Document additions with inline comments explaining feature semantics for future maintainers.

Frequently Asked Questions

Where exactly do I modify the code to add a custom numeric feature?

Add your computation logic inside the extract_social_numerical_features function in src/context_features_extractor.py, specifically after the existing user_features and tweet_features lists are combined but before the final conversion to a NumPy array (around line 55). This is the only location where the flat numeric vector is constructed.

What happens if I forget to update NUMERICAL_FEATURE_DIM?

If you add features to the vector but leave NUMERICAL_FEATURE_DIM at its default value (28), downstream components that rely on this constant for padding, slicing, or shape validation will encounter dimension mismatches. The RNN may raise tensor shape errors during training or inference because EXPECTED_CONTEXT_INPUT_SIZE will no longer align with the actual data width.

Can I add features that require external API calls or database lookups?

No. The context_features_extractor is designed to operate solely on the reaction_status_json dictionary passed to it. External I/O would significantly slow down the batch processing loop in context_feature_extraction. Pre-compute external data and embed it into the JSON before extraction, or add lookup logic upstream in your data pipeline.

How do I test that my custom feature is working correctly?

Use the test_context_features_shape_by_tweet_id function provided in src/context_features_extractor.py. Pass a valid tweet ID string to this function; it will extract context features for that tweet and assert that the output matrix shape matches EXPECTED_CONTEXT_INPUT_SIZE. If the assertion passes, your dimensionality updates are correct.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →