How to Implement Topic Modeling for Financial News: A Complete LDA Pipeline
You can implement topic modeling for financial news by preprocessing JSON articles with spaCy, vectorizing the corpus using scikit-learn TF-IDF, training a gensim LDA model with auto-tuned hyperparameters, and evaluating results through coherence scores and interactive pyLDAvis visualizations.
This guide walks you through a production-ready implementation of topic modeling for financial news using Latent Dirichlet Allocation (LDA). Based on the stefan-jansen/machine-learning-for-trading repository, this workflow processes thousands of Reuters and CNBC articles to extract latent themes from market-moving text data using the notebooks found in 16_word_embeddings/ and 15_topic_modeling/.
Preprocessing Financial News Articles
The pipeline begins in 16_word_embeddings/03_financial_news_preprocessing.ipynb, which handles ingestion and cleaning of raw JSON articles from the data/us-financial-news/ directory. This stage ensures only relevant financial content enters the model.
Filtering by Financial Sections
To maintain domain relevance, the code filters articles using a curated whitelist of section titles including "Reuters: Company News", "Reuters: Business News", and "Press Releases - CNBC". The read_articles() function loads each JSON file, checks the section_title field, and accumulates articles while removing standard English stop words from an external list.
Text Cleaning with spaCy
The repository uses spaCy for sentence-level tokenization and lemmatization. The English pipeline (en) is loaded with the NER component disabled to improve processing speed on large volumes of text.
The clean_doc() function performs aggressive normalization:
- Converts text to lowercase and removes digits
- Strips punctuation, short tokens, and pronoun lemmas (
-PRON-) - Lemmatizes remaining tokens using spaCy's linguistic pipelines
Processing occurs in parallel via nlp.pipe() with n_process=8 and batch_size=100 for efficient throughput on the multi-gigabyte corpus.
N-gram Construction
To capture multi-word financial expressions like "interest rates" or "market volatility", the implementation uses gensim's Phrases and Phraser classes. These detect statistically significant bigrams and trigrams from the token stream using a threshold of 100 and minimum count of 10, smoothing over rare co-occurrences that could introduce noise.
Building the LDA Pipeline
The core modeling logic resides in 15_topic_modeling/07_lda_financial_news.ipynb, which transforms cleaned text into interpretable topics.
TF-IDF Vectorization
Rather than raw count vectors, the pipeline uses scikit-learn's TfidfVectorizer to weight terms by their importance within documents relative to the corpus. Key parameters include:
min_df=0.005andmax_df=0.1to prune extremely rare and overly common termsngram_range=(1, 1)to use unigrams (following the preprocessing n-gram stage)stop_words='english'as a secondary filter
This produces a sparse document-term matrix (dtm) optimized for financial vocabulary density.
Converting to Gensim Format
Since gensim's LDA implementation requires specific data structures, the code converts the scikit-learn matrix using Sparse2Corpus with documents_columns=False. A mapping dictionary is created via Dictionary.from_corpus() using the feature names from vectorizer.get_feature_names_out() to maintain token-to-ID alignment.
Training the LDA Model
The LdaModel is configured for robust convergence on financial text with these parameters:
num_topics(typically explored between 5 and 25)chunksizeset to the full document count for batch processingalpha='auto'andeta='auto'for data-driven Dirichlet priorspasses=10anditerations=50to ensure sufficient traversal of the corpusminimum_probability=0.01to filter negligible topic weights
Setting random_state=42 ensures reproducible results across training runs.
Evaluation Metrics
The implementation computes two standard metrics to assess model quality:
- Perplexity: Calculated as
2 ** (-lda.log_perplexity(corpus))to measure how well the model predicts a sample; lower values indicate better generalization - Coherence: Using the
u_massmetric to score topic interpretability based on word co-occurrence statistics
Helper functions show_word_list() and show_coherence() generate heatmaps plotting the top N words per topic alongside their probability distributions.
Complete Implementation Example
The following script consolidates the entire workflow into a single executable pipeline. Ensure you have downloaded the US Financial News dataset to data/us-financial-news/ and installed the required dependencies (spacy, gensim, scikit-learn, pyLDAvis).
import json
import warnings
from pathlib import Path
from collections import Counter
import spacy
import numpy as np
import pandas as pd
import pyLDAvis
from gensim.models import LdaModel, Phrases, Phraser
from gensim.corpora import Dictionary
from gensim.matutils import Sparse2Corpus
from pyLDAvis.gensim_models import prepare
from sklearn.feature_extraction.text import TfidfVectorizer
warnings.filterwarnings("ignore")
# -------------------------------------------------
# 1. Load & filter raw articles
# -------------------------------------------------
data_path = Path("data/us-financial-news")
section_titles = [
"Press Releases - CNBC", "Reuters: Company News", "Reuters: World News",
"Reuters: Business News", "Reuters: Financial Services and Real Estate",
"Top News and Analysis (pro)", "Reuters: Top News",
"The Wall Street Journal & Breaking News, Business, Financial and Economic News, World News and Video",
"Business & Financial News, U.S & International Breaking News | Reuters",
"Reuters: Money News", "Reuters: Technology News",
]
stop_words = set(
pd.read_csv(
"http://ir.dcs.gla.ac.uk/resources/linguistic_utils/stop_words",
header=None,
squeeze=True,
)
)
def read_articles():
articles, cnt = [], Counter()
for f in data_path.rglob("*.json"):
article = json.load(f.open())
if article["thread"]["section_title"] in set(section_titles):
tokens = article["text"].lower().split()
cnt.update(tokens)
articles.append(" ".join(t for t in tokens if t not in stop_words))
return articles, cnt
articles, _ = read_articles()
# -------------------------------------------------
# 2. SpaCy cleaning
# -------------------------------------------------
nlp = spacy.load("en", disable=["ner"])
nlp.max_length = 6_000_000
def clean_doc(doc):
return " ".join(
t.lemma_
for t in doc
if not (
t.is_stop
or t.is_digit
or not t.is_alpha
or t.is_punct
or t.is_space
or t.lemma_ == "-PRON-"
)
)
cleaned = [
clean_doc(doc) for doc in nlp.pipe(articles, batch_size=100, n_process=8)
]
# -------------------------------------------------
# 3. Build n-grams
# -------------------------------------------------
sentences = [sentence.split() for sentence in cleaned]
phrases = Phrases(sentences, threshold=100, min_count=10)
bigram = Phraser(phrases)
tokenised = [bigram[sentence] for sentence in sentences]
# -------------------------------------------------
# 4. TF-IDF vectorisation
# -------------------------------------------------
vectorizer = TfidfVectorizer(
stop_words="english", min_df=0.005, max_df=0.1, ngram_range=(1, 1)
)
dtm = vectorizer.fit_transform([" ".join(t) for t in tokenised])
tokens = vectorizer.get_feature_names_out()
# -------------------------------------------------
# 5. Convert to gensim corpus & dictionary
# -------------------------------------------------
corpus = Sparse2Corpus(dtm, documents_columns=False)
id2word = dict(enumerate(tokens))
dictionary = Dictionary.from_corpus(corpus, id2word=id2word)
# -------------------------------------------------
# 6. Train LDA (example: 15 topics)
# -------------------------------------------------
lda = LdaModel(
corpus=corpus,
id2word=id2word,
num_topics=15,
chunksize=len(tokenised),
passes=10,
iterations=50,
alpha="auto",
eta="auto",
random_state=42,
)
# -------------------------------------------------
# 7. Evaluate & visualise
# -------------------------------------------------
perplexity = 2 ** (-lda.log_perplexity(corpus))
print(f"Perplexity: {perplexity:.2f}")
vis = prepare(lda, corpus, dictionary, mds="tsne")
pyLDAvis.display(vis) # Inline in notebooks
# pyLDAvis.save_html(vis, "financial_news_lda.html") # Export static HTML
Running this script executes the full pipeline from raw JSON to trained model, outputting perplexity metrics and launching an interactive browser visualization where you can inspect topic-term relevance and document-topic distributions.
Summary
- Preprocessing matters: Filter articles by financial section titles and use spaCy with disabled NER for efficient lemmatization of market-specific language.
- Hybrid vectorization works best: Combine scikit-learn's
TfidfVectorizerwith gensim'sSparse2Corpusto leverage TF-IDF weighting within the LDA framework. - Auto-tune your priors: Setting
alpha='auto'andeta='auto'allows the model to learn optimal Dirichlet priors from the financial corpus rather than using fixed defaults. - Evaluate rigorously: Use both perplexity (statistical fit) and coherence (human interpretability) to select the optimal
num_topicsbetween 5 and 25. - Visualize interactively: Export results using
pyLDAvis.prepare()withmds="tsne"to explore how financial themes cluster in semantic space.
Frequently Asked Questions
What is the optimal number of topics for financial news LDA?
According to the 15_topic_modeling/07_lda_financial_news.ipynb implementation, you should train models with num_topics ranging from 5 to 25 and select the value that maximizes coherence scores while minimizing perplexity. Financial news typically supports 10-20 distinct themes before topic overlap becomes significant.
Why use TF-IDF before LDA instead of raw bag-of-words?
The repository uses TfidfVectorizer with min_df=0.005 and max_df=0.1 to suppress corpus-specific noise—such as ticker symbols appearing in every article or rare technical jargon—that would otherwise dominate the Dirichlet distributions. TF-IDF weighting improves topic distinctiveness by focusing on discriminative financial terms rather than high-frequency market boilerplate.
How does the alpha='auto' parameter affect topic modeling results?
Setting alpha='auto' in gensim.models.LdaModel enables the model to learn asymmetric document-topic priors from the data rather than assuming a uniform 1/K distribution. For financial news, this allows the algorithm to automatically identify that certain articles (e.g., earnings reports) concentrate heavily in specific topics while others (e.g., general market summaries) distribute evenly across themes.
Can this pipeline handle real-time streaming financial news?
While the notebooks in stefan-jansen/machine-learning-for-trading process static JSON archives, the underlying functions are stateless and can be adapted for streaming. You would maintain the fitted TfidfVectorizer and trained LdaModel in memory, then call lda.get_document_topics() on the fly after applying the same clean_doc() preprocessing to incoming articles.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →