# How to Integrate RP-DNN with External NLP Pipelines for Production Deployment

> Deploy RP-DNN in production NLP pipelines. Package your AllenNLP model as a Predictor, expose it via API, and inject external features using the context tensor interface.

- Repository: [jerrygao/rpdnn](https://github.com/jerrygaolondon/rpdnn)
- Tags: how-to-guide
- Published: 2026-03-04

---

**You can integrate RP-DNN into production NLP workflows by packaging the AllenNLP-based model as a Predictor, exposing it via a REST API, and injecting external features through the context tensor interface defined in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py).**

RP-DNN (Propagation-Context Deep Neural Network) is an AllenNLP-based rumor detection system developed in the `jerrygaolondon/rpdnn` repository. To integrate RP-DNN with external NLP pipelines for production deployment, you must understand its three-stream architecture and leverage the native AllenNLP Predictor interface for standardized inference.

## Understanding the RP-DNN Architecture for Integration

The RP-DNN model processes three distinct information streams that must be maintained when integrating with external pipelines.

### The Three-Stream Input Design

| Stream | Source | Key Implementation |
|--------|--------|-------------------|
| **Source tweet text** | Raw tweet string | `RumorTweetsClassifer` uses an ELMo-based `tweet_text_embedder` and `lang_model_encoder` (BiLSTM) for dense representation |
| **Social-context content** | Reply/retweet text from PHEME corpus | `context_feature_extraction_from_context_status` in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) tokenizes replies through the same ELMo pipeline (`sentence_embedding_elmo`) |
| **Social-context metadata** | User-level, temporal, and structural features | `extract_social_numerical_features` generates a **28-dimensional** numeric vector (defined by `NUMERICAL_FEATURE_DIM`) |

These streams merge inside `RumorTweetsClassifer.forward` in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py), which applies hierarchical attention (`HierarchicalAttentionNet`) or structured self-attention (`StructuredSelfAttention` from [`src/attention.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/attention.py)), followed by custom layer normalization (`MyLayerNorm` in [`src/my_layer_norm.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/my_layer_norm.py)).

### Core Inference Components

The model's `forward` method signature determines how external pipelines must format inputs:

```python
def forward(self,
            sentence: Dict[str, torch.Tensor],
            tweet_id: list,
            label: torch.LongTensor = None) -> Dict[str, torch.Tensor]:

```

- `sentence` = AllenNLP `TextField` output (tokenized by the embedder)
- `tweet_id` = list of source tweet IDs used by `load_source_tweet_context` in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) to retrieve social context

## Packaging RP-DNN as an AllenNLP Predictor

The standard method for production integration uses AllenNLP's `Predictor` class to load serialized checkpoints produced by [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py).

```python
from allennlp.predictors import Predictor
import os

# Configure GPU visibility before loading (matches lines 97-100 in rumour_dnn_trainer.py)

os.environ["CUDA_VISIBLE_DEVICES"] = "0"

# Load archived model produced by the training script

predictor = Predictor.from_path(
    "/path/to/trained/model.tar.gz",
    predictor_name="rumor_tweets_classifier"
)

# Single prediction

result = predictor.predict_json({
    "sentence": "Breaking: major incident reported downtown.",
    "tweet_id": "524963572083085313"
})
print(result["label"], result["class_probabilities"])

```

The predictor expects a JSON payload containing both the `sentence` (source tweet text) and `tweet_id` for context retrieval from the PHEME corpus directory structure.

## Bridging External Preprocessing Pipelines

When your upstream NLP pipeline already extracts entities, sentiment, or linguistic features, you can inject them by bypassing the default feature extractors and passing pre-computed tensors directly to the model's `forward` method.

### Method 1: Extending the Feature Extractor

Modify `context_feature_extraction_from_context_status` in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) to accept external annotations alongside the raw reply JSON.

### Method 2: Direct Tensor Injection

For pipelines that pre-compute embeddings, construct the `cxt_content_tensor` and `cxt_metadata_tensor` externally and feed them through the model's `forward` method, bypassing `load_source_tweet_context`.

## Deploying as a Production Service

Wrap the Predictor in a FastAPI application for containerized deployment behind load balancers.

```python
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from allennlp.predictors import Predictor
import os

app = FastAPI()

# Initialize once at startup

os.environ["CUDA_VISIBLE_DEVICES"] = "0"
predictor = Predictor.from_path(
    "model.tar.gz", 
    predictor_name="rumor_tweets_classifier"
)

class PredictionRequest(BaseModel):
    tweet_id: str
    sentence: str

@app.post("/predict")
async def predict(req: PredictionRequest):
    try:
        result = predictor.predict_json({
            "sentence": req.sentence, 
            "tweet_id": req.tweet_id
        })
        return {
            "label": result["label"],
            "probabilities": result["class_probabilities"].tolist()
        }
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))

```

The service automatically handles GPU tensor placement based on the `cuda_device` parameter (default `-1` for CPU).

## Optimizing Inference Performance

### Batch Processing for High Throughput

Use `predict_batch_json` to process multiple tweets in parallel, leveraging `batch_compute_context_feature_encoding` (lines 222-260 in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py)) for efficient context loading and padding.

```python
def batch_rp_dnn_predict(tweets, predictor):
    """
    tweets: list of dicts with keys "sentence" and "tweet_id"
    """
    batch_payload = {
        "sentence": [t["sentence"] for t in tweets],
        "tweet_id": [t["tweet_id"] for t in tweets]
    }
    results = predictor.predict_batch_json(batch_payload)
    return list(zip(results["label"], results["class_probabilities"]))

```

### GPU Resource Management

- Set `CUDA_VISIBLE_DEVICES` before importing AllenNLP to control device visibility
- The model respects AllenNLP's `cuda_device` configuration (-1 for CPU, 0+ for GPU)
- For multi-GPU deployment, launch separate container instances per GPU rather than using DataParallel

## Extending Features with External NLP Outputs

To incorporate external sentiment scores or entity features into the 28-dimensional social metadata vector:

1. **Modify the extractor** in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py):

```python
def extract_social_numerical_features(
    ..., 
    external_sentiment: float = 0.0,
    entity_count: int = 0
):
    numerical_features = [...]  # existing 28 features

    numerical_features.append(external_sentiment)
    numerical_features.append(entity_count)
    return np.array(numerical_features)

```

2. **Update the dimension constant** in the same file:

```python
NUMERICAL_FEATURE_DIM = 30  # increased from 28

```

3. **Retrain** using [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py) to align the classifier input layer with the new feature dimensions.

## Summary

- **RP-DNN is AllenNLP-native**: Integration relies on the `Predictor` interface and standard AllenNLP archive formats produced by [`rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/rumour_dnn_trainer.py).
- **Three-stream architecture**: Production pipelines must supply source text, reply content (via `tweet_id` lookup or direct tensor injection), and 28-dimensional metadata features.
- **Extension points**: Modify `extract_social_numerical_features` in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) to inject external NLP features (sentiment, entities), updating `NUMERICAL_FEATURE_DIM` accordingly.
- **Performance optimization**: Use `predict_batch_json` and `batch_compute_context_feature_encoding` for high-throughput scenarios, with explicit GPU device management via `CUDA_VISIBLE_DEVICES`.
- **Service deployment**: Wrap the AllenNLP Predictor in FastAPI/Flask, ensuring the `tweet_id` resolution logic in [`src/data_loader.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/data_loader.py) can access your PHEME-formatted context directory or override it with custom tensor injection.

## Frequently Asked Questions

### How do I handle high-throughput batch inference with RP-DNN?

Use the `predictor.predict_batch_json()` method with a payload containing lists of sentences and tweet IDs. The internal `batch_compute_context_feature_encoding` function (lines 222-260 in [`src/allennlp_rumor_classifier.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/allennlp_rumor_classifier.py)) parallelizes context loading and handles padding efficiently. For maximum throughput, ensure your context data is stored on fast SSD storage or cache frequently accessed tweet contexts in memory.

### Can I use RP-DNN with newer AllenNLP versions or Hugging Face transformers?

The current implementation relies on ELMo embedders and specific AllenNLP 0.x/1.x `Seq2VecEncoder` interfaces. While the core `RumorTweetsClassifer` architecture is modular, replacing ELMo with Hugging Face transformers requires modifying the `tweet_text_embedder` configuration and ensuring the `forward` method's tensor dimensions align with the new encoder outputs. The attention mechanisms in [`src/attention.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/attention.py) are encoder-agnostic and will function with any compatible tensor shapes.

### How do I inject custom entity extraction features into the model?

Extend `extract_social_numerical_features` in [`src/context_features_extractor.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/context_features_extractor.py) to accept entity counts or types as parameters, append them to the `numerical_features` list, and increment `NUMERICAL_FEATURE_DIM` to match the new vector length. When calling from your external pipeline, pass the entity features extracted by your NER system (e.g., spaCy or Stanza) as keyword arguments. You must retrain the model using [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py) after modifying the feature dimensions.

### What is the recommended GPU memory configuration for production?

RP-DNN loads ELMo embeddings and maintains LSTM or transformer encoders for context processing, requiring approximately 4-6 GB GPU memory for batch sizes of 32-64. Set `CUDA_VISIBLE_DEVICES` before initialization (as implemented in [`src/rumour_dnn_trainer.py`](https://github.com/jerrygaolondon/rpdnn/blob/main/src/rumour_dnn_trainer.py) lines 97-100) to isolate devices. For production stability, run inference on a single GPU per process rather than sharing GPUs across multiple model instances, as the ELMo embedder is memory-intensive.