# How the Hierarchical Encoder Processes Multi-Level Propagation Trees in RPDNN

> Learn how the hierarchical encoder processes multi-level propagation trees by converting JSON to tensor sequences, encoding tweets, and using attention layers for classification.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: deep-dive
- Published: 2026-03-04

---

**The hierarchical encoder converts nested JSON propagation trees into time-ordered tensor sequences, applies modality-specific neural encoders to individual tweets, and uses stacked attention layers to aggregate the entire tree into a single classification vector.**

The RPDNN (Rumor Propagation Deep Neural Network) repository implements a sophisticated hierarchical encoder that transforms complex, tree-structured social media propagation histories into compact feature representations suitable for rumor classification. This architecture processes multi-level propagation trees through four tightly coupled stages, preserving both temporal ordering and hierarchical depth information while learning the relative importance of each node in the propagation structure.

## Stage 1: Flattening the Propagation Tree into Temporal Sequences

The encoder begins by transforming the nested tree structure of replies and retweets into ordered sequences that neural networks can process.

### Recursive Tree Traversal

Each source tweet’s propagation history arrives as a nested JSON dictionary stored in the dataset. The function `looping_nested_dict()` in [`src/context_features/propagation_features.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features/propagation_features.py) recursively walks this structure, yielding each tweet identifier together with its **depth** in the propagation tree. This preserves hierarchical position information even as the tree is flattened.

### Feature Extraction and Temporal Ordering

For every context tweet discovered during traversal, `context_feature_extraction_from_context_status()` generates two distinct representations:

- **Content embedding**: An ELMo-based sentence vector produced by the embedding layer in [`src/embeddings/embedding_layer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/embeddings/embedding_layer.py)
- **Metadata vector**: A numeric feature vector encoding temporal and structural attributes

The method `_context_sequence_encoding()` in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) collects these per-tweet embeddings, sorts them strictly by timestamp, and returns two ordered tensors: `cxt_content_seq_tensor` (content embeddings) and `cxt_metadata_seq_tensor` (metadata features). This transformation converts the multi-level propagation tree into parallel time-ordered sequences while maintaining the depth information for potential downstream use.

## Stage 2: Low-Level Per-Tweet Encoding

Once sequences are constructed, the encoder processes individual tweets through modality-specific neural networks before applying hierarchical aggregation.

**Content encoding** flows through `self.cxt_content_encoder`, typically implemented as an LSTM or feed-forward layer that projects raw ELMo embeddings into a unified hidden dimension. **Metadata encoding** passes through `self.cxt_metadata_encoder`, a dedicated feed-forward network that processes the numeric feature vectors. Both encoders output sequences of hidden states with shape `(seq_len, hidden_dim)`, creating dense representations for every node in the propagation tree.

## Stage 3: Hierarchical Attention Mechanism

The core of the hierarchical encoder resides in the `HierarchicalAttentionNet` class defined in [`src/attention.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/attention.py), which aggregates variable-length sequences into fixed-size vectors through learned attention weights.

### The HierarchicalAttentionNet Architecture

During model initialization in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py), the classifier instantiates three separate attention modules:

- `self.cxt_content_attention`: Aggregates content embeddings
- `self.cxt_metadata_attention`: Aggregates metadata features  
- `self.cxt_multimodal_attention`: Fuses the two modalities

Each module receives two critical parameters: `feature_dim` (the hidden dimension size) and `step_dim` (the maximum sequence length, referenced as `max_cxt_size`).

### Attention Computation

In the `forward(x, mask=None)` method, the layer computes attention through the following operations:

```python
eij = torch.tanh(torch.matmul(x, self.w) + self.b)  # (batch, steps, 1)

eij = eij.squeeze(-1).masked_fill(~mask, float('-inf'))  # Apply padding mask

a = F.softmax(eij, dim=1)  # Attention weights over time steps

weighted_input = x * a.unsqueeze(-1)  # Element-wise weighting

context_weighted_sum = weighted_input.sum(dim=1)  # Single output vector

```

The output `context_weighted_sum` provides a single vector summarizing the entire propagation sequence, while the intermediate tensor `a` represents the learned importance of each tweet in the tree.

### Multi-Modal Attention Stacking

By stacking three attention modules sequentially, the model implements a hierarchical decision process:

1. **Content attention** determines which specific tweets carry the most semantic information
2. **Metadata attention** identifies which structural or temporal features are most predictive  
3. **Multimodal attention** (when enabled) learns how content and metadata interact by concatenating the outputs of the first two layers and applying a final attention pass

This architecture allows the model to weigh the importance of different propagation depths and branches dynamically rather than treating all tree nodes equally.

## Stage 4: Integration with the Classifier

After hierarchical aggregation completes, the weighted vectors undergo layer normalization via `MyLayerNorm` before passing through `self.classifier_feedforward`, a task-specific feed-forward network that produces the final classification logits. The resulting representation carries information from the original propagation tree at multiple levels: the temporal order of diffusion, the hierarchical depth of replies, and the learned relevance of individual nodes.

## Code Implementation Examples

### Instantiating and Applying Hierarchical Attention

The following example demonstrates how to apply the attention mechanism to a batch of content sequences:

```python
from attention import HierarchicalAttentionNet
import torch

# Batch of 8 propagation trees, padded to max_len=100, encoded to 256-dim vectors

content_seq = torch.randn(8, 100, 256)  # (batch, steps, hidden_dim)

mask = torch.arange(100).unsqueeze(0) < torch.tensor([70, 45, 90, 30, 100, 55, 20, 80]).unsqueeze(1)

# Initialize attention (feature_dim=256, step_dim=100)

content_att = HierarchicalAttentionNet(feature_dim=256, step_dim=100)

# Get aggregated representation and attention weights

content_vec, weighted_seq, att_weights = content_att(content_seq, mask)
print(content_vec.shape)  # torch.Size([8, 256])

```

### Full Forward Pass Integration

This excerpt from `RumorTweetsClassifer.forward` illustrates the complete hierarchical encoding pipeline:

```python

# 1️⃣ Encode propagation history into ordered tensors

cxt_content_tensor, cxt_metadata_tensor = self._context_sequence_encoding(
    source_tweet_context_dataset, source_tweet
)

# 2️⃣ Apply hierarchical attention per modality

h_cc, cc_weighted, cc_att_w = self.cxt_content_attention(
    cxt_content_tensor, batch_cxt_content_mask
)

h_cm, cm_weighted, cm_att_w = self.cxt_metadata_attention(
    cxt_metadata_tensor, batch_cxt_metadata_mask
)

# 3️⃣ Fuse modalities with final attention layer

if self.cxt_multimodal_attention:
    multimodal = torch.cat([h_cc, h_cm], dim=-1)
    h_c, _, _ = self.cxt_multimodal_attention(multimodal, batch_cxt_content_mask)

```

## Summary

- The hierarchical encoder flattens nested propagation trees into timestamp-ordered sequences while preserving depth information via `looping_nested_dict()` in [`src/context_features/propagation_features.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features/propagation_features.py).
- Separate encoders in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py) process content (via LSTM) and metadata (via feed-forward networks) to produce hidden state sequences.
- The `HierarchicalAttentionNet` class in [`src/attention.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/attention.py) implements learnable attention that aggregates variable-length sequences into fixed-size vectors using masked softmax operations.
- Three stacked attention layers (content, metadata, multimodal) enable the model to learn which tweets, which feature types, and which modality interactions are most predictive for rumor detection.
- Final representations pass through layer normalization and classification layers, incorporating hierarchical information from all levels of the propagation tree.

## Frequently Asked Questions

### What is the input format for the hierarchical encoder?

The encoder expects propagation trees stored as nested JSON dictionaries where each node represents a tweet with associated metadata. The `context_feature_extraction_from_context_status()` function in the feature extraction pipeline transforms these raw structures into parallel tensors—`cxt_content_seq_tensor` for ELMo text embeddings and `cxt_metadata_seq_tensor` for numeric features—sorted chronologically by timestamp.

### How does the attention mechanism handle variable-length propagation trees?

The `HierarchicalAttentionNet` class accepts a boolean `mask` tensor during its forward pass. This mask, typically generated from `batch_cxt_content_mask` or `batch_cxt_metadata_mask`, identifies valid time steps versus padding tokens. The implementation uses `masked_fill(~mask, float('-inf'))` to ensure that padded positions receive zero attention weight after the softmax operation, allowing batches to contain trees of varying depths without information leakage from padding.

### Why use separate encoders for content and metadata?

Content embeddings (derived from ELMo) and metadata vectors (containing temporal and structural statistics) reside in different semantic spaces with distinct dimensionalities and scales. The architecture uses `self.cxt_content_encoder` (typically an LSTM) to capture sequential dependencies in text, while `self.cxt_metadata_encoder` (a feed-forward network) processes fixed-size numeric feature vectors. This separation allows each modality to learn appropriate transformations before the hierarchical attention layers fuse them.

### Where does the depth information from the propagation tree get used?

During the initial traversal, `looping_nested_dict()` in [`src/context_features/propagation_features.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features/propagation_features.py) yields each tweet's depth in the tree alongside its identifier. While the primary encoding pipeline focuses on temporal ordering via `_context_sequence_encoding()`, the depth values are preserved in the extracted features and can be incorporated into the metadata vectors. This allows the hierarchical attention mechanism to potentially learn that replies at certain depths (e.g., immediate replies versus deep nested threads) carry different predictive weights for rumor classification.