How to Fine-Tune ELMo Embeddings for Domain-Specific Rumor Detection with RP-DNN
Fine-tune ELMo embeddings for domain-specific rumor detection by loading the CredBank-adapted weight file into AllenNLP's ElmoEmbedder, then integrating it into the RP-DNN pipeline to encode tweets, replies, and user profiles with contextualized 1024-dimensional vectors.
The RP-DNN repository implements a deep neural network for early rumor detection on Twitter, leveraging ELMo (Embeddings from Language Models) to capture contextual word representations. Fine-tuning ELMo embeddings for the rumor verification domain allows the model to understand credibility-specific language patterns, reducing out-of-vocabulary errors and improving detection accuracy on social media text.
Locating the Domain-Adapted Weight Files
Pre-trained Weights Location
The fine-tuned ELMo weights are stored in the repository under resource/embedding/elmo_model/ as the HDF5 file elmo_credbank_2x4096_512_2048cnn_2xhighway_weights_10052019.hdf5. This file contains the domain-adapted parameters trained on the CredBank corpus, which is rich in credibility-related language.
Configuration Files
The matching options JSON for the original ELMo architecture is downloaded on-the-fly from the AllenNLP S3 bucket at https://s3-us-west-2.amazonaws.com/allennlp/models/elmo/2x4096_512_2048cnn_2xhighway/elmo_2x4096_512_2048cnn_2xhighway_options.json. This file defines the character-level CNN filters and BiLSTM layer dimensions that the weight file expects.
Loading the Fine-Tuned ELMo Embedder
Global Embedder in Context Feature Extractor
The embedder is instantiated in src/context_features_extractor.py at lines 24-27 to provide a global fine_tuned_elmo object used throughout the context-feature extraction pipeline:
from allennlp.commands.elmo import ElmoEmbedder
fine_tuned_elmo = ElmoEmbedder(
options_file="https://s3-us-west-2.amazonaws.com/allennlp/models/elmo/2x4096_512_2048cnn_2xhighway/elmo_2x4096_512_2048cnn_2xhighway_options.json",
weight_file=elmo_credbank_model_path
)
Model-Specific Embedder in Rumor Classifier
The classifier maintains its own ElmoEmbedder in src/allennlp_rumor_classifier.py at lines 258-262 for downstream inference:
self.elmo_model = ElmoEmbedder(
options_file="https://s3-us-west-2.amazonaws.com/allennlp/models/elmo/2x4096_512_2048cnn_2xhighway/elmo_2x4096_512_2048cnn_2xhighway_options.json",
weight_file=elmo_credbank_model_path,
cuda_device=cuda_device
)
Both snippets reference the same elmo_credbank_model_path, which resolves to the file loaded by load_abs_path in src/data_loader.py. This path indirection allows you to change the weight file location without modifying the source code.
Generating Sentence Embeddings with ELMo
The sentence_embedding_elmo Function
Utility functions in src/embeddings/embedding_layer.py at lines 28-87 wrap the raw embedder calls to provide a consistent API:
# src/embeddings/embedding_layer.py
def sentence_embedding_elmo(sentence: List[str],
elmo_model: ElmoEmbedder,
remove_stopwords=False,
avg_all_layers=False) -> np.ndarray:
"""
Returns the mean-pooled ELMo vector for a tokenised sentence.
"""
if remove_stopwords:
sentence = list(stop_words_filter(sentence))
sentence_vectors = elmo_model.embed_sentence(sentence) # <-- raw ELMo tensors
if not avg_all_layers:
sentence_word_embeddings = sentence_vectors[2][:] # top layer only
else:
avg_all_layer_sent_embedding = np.mean(sentence_vectors, axis=0, dtype='float32')
return np.mean(avg_all_layer_sent_embedding, axis=0, dtype='float32')
return np.mean(sentence_word_embeddings, axis=0).astype('float32')
Layer Selection and Pooling Strategies
The function supports two pooling strategies controlled by the avg_all_layers parameter:
- Top layer only (
avg_all_layers=False): Usessentence_vectors[2]to extract the final BiLSTM layer representation, yielding task-specific contextualized embeddings. - Average all layers (
avg_all_layers=True): Computes the mean across all three ELMo layers (character CNN, first BiLSTM, second BiLSTM), producing a more general semantic representation suitable for user profile descriptions and reply content.
Integrating ELMo into the Rumor Detection Pipeline
Encoding Source Tweets
In RumorTweetsClassifer.forward, after tokenizing the source tweet, the pipeline executes:
embeddings = self.tweet_text_embedder(sentence) # AllenNLP TextFieldEmbedder
lang_model_encoder_out = self.lang_model_encoder(embeddings, mask)
The tweet_text_embedder is a BasicTextFieldEmbedder that internally uses the ELMo embedder defined in the model configuration. By default, this points to the fine-tuned CredBank weights, ensuring domain-specific representations for the source claim.
Processing Reply Content and User Profiles
The contextual features are extracted in src/context_features_extractor.py. When processing reply text, the function calls:
reply_sent_embedding = sentence_embedding_elmo(tokens, elmo_model, avg_all_layers=True)
This yields a 1024-dimensional vector that is later concatenated with numerical features and passed through LSTM or attention layers. For user profile descriptions, the same sentence_embedding_elmo routine is applied (see lines 95-106 in context_features_extractor.py), ensuring consistent domain adaptation across all text inputs.
Training and Inference Configuration
JSON Configuration for AllenNLP
The trainer script (src/rumour_dnn_trainer.py) builds an AllenNLP Model from a JSON configuration. The configuration includes an "elmo" token indexer that points to the custom weight file:
{
"token_embedders": {
"elmo_emb": {
"type": "elmo_token_embedder",
"options_file": "https://s3-us-west-2.amazonaws.com/allennlp/models/elmo/2x4096_512_2048cnn_2xhighway/elmo_2x4096_512_2048cnn_2xhighway_options.json",
"weight_file": "<PATH>/elmo_credbank_2x4096_512_2048cnn_2xhighway_weights_10052019.hdf5"
}
}
}
Freezing vs. Fine-Tuning Parameters
During training, the embedder's parameters are frozen by default because AllenNLP's ElmoTokenEmbedder loads them as constants. To continue fine-tuning on your own rumor corpus, set "requires_grad": true for the ElmoTokenEmbedder in the config, or manually wrap the embedder in a torch.nn.Parameter and enable back-propagation.
Running the Trainer and Evaluator
Execute the training pipeline with the fine-tuned embedder:
python src/rumour_dnn_trainer.py \
-t data/train.csv \
-e data/val.csv \
-p my_rpdnn \
-g 0 \
-f -1 \
--params_file params.json
For inference, the evaluator automatically loads the serialized fine-tuned embedder:
python src/rumour_dnn_evaluator.py \
-t data/test.csv \
-m output/RPDNN_model_output_20221015/full/ferguson_full202210151230 \
-g 0 \
-f -1 \
--max_cxt_size 200
Summary
- Fine-tuned ELMo weights for rumor detection are stored in
resource/embedding/elmo_model/elmo_credbank_2x4096_512_2048cnn_2xhighway_weights_10052019.hdf5. - The AllenNLP
ElmoEmbedderloads these weights alongside the standard options file to provide 1024-dimensional contextualized vectors. - Utility functions in
src/embeddings/embedding_layer.pywrap the embedder to handle tokenization, stopword removal, and layer averaging. - The pipeline applies the same fine-tuned parameters to source tweets, reply content, and user profiles, ensuring domain consistency.
- Training configurations in
params.jsoncontrol whether ELMo parameters are frozen or fine-tuned via therequires_gradflag.
Frequently Asked Questions
What is the difference between pre-trained and fine-tuned ELMo in RP-DNN?
The pre-trained ELMo model provides general-purpose contextualized embeddings trained on large corpora like the 1 Billion Word Benchmark. The fine-tuned version in RP-DNN uses weights adapted on the CredBank corpus, which contains credibility-related language specific to rumor verification. This domain adaptation captures terms like "misinformation" and "unverified" more accurately than the generic model.
Can I replace the CredBank weights with my own domain-specific ELMo model?
Yes. Place your custom HDF5 weight file in resource/embedding/elmo_model/ or update the elmo_credbank_model_path variable in src/data_loader.py to point to your file. Ensure your model uses the same architecture (2x4096_512_2048cnn_2xhighway) so the options file remains compatible, or provide a matching options JSON for your architecture.
How do I enable gradient updates for the ELMo layers during training?
By default, AllenNLP's ElmoTokenEmbedder freezes ELMo parameters. To enable fine-tuning during training, set "requires_grad": true in the elmo_emb configuration within your params.json file. Alternatively, manually instantiate the ElmoEmbedder in Python and wrap the parameters with torch.nn.Parameter, setting requires_grad=True before passing them to the optimizer.
What hardware requirements are needed for training with fine-tuned ELMo?
The fine-tuned ELMo model requires approximately 4-5 GB of GPU memory for the embedding layer alone when processing batches of Twitter-sized texts. For the full RP-DNN pipeline including LSTM encoders and attention mechanisms, a GPU with at least 8-11 GB VRAM (such as an NVIDIA RTX 2080 Ti or V100) is recommended. CPU-only inference is possible but significantly slower for the embedding generation step.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →