How to Implement Topic Modeling for Financial News: A Complete LDA Pipeline

You can implement topic modeling for financial news by preprocessing JSON articles with spaCy, vectorizing the corpus using scikit-learn TF-IDF, training a gensim LDA model with auto-tuned hyperparameters, and evaluating results through coherence scores and interactive pyLDAvis visualizations.

This guide walks you through a production-ready implementation of topic modeling for financial news using Latent Dirichlet Allocation (LDA). Based on the stefan-jansen/machine-learning-for-trading repository, this workflow processes thousands of Reuters and CNBC articles to extract latent themes from market-moving text data using the notebooks found in 16_word_embeddings/ and 15_topic_modeling/.

Preprocessing Financial News Articles

The pipeline begins in 16_word_embeddings/03_financial_news_preprocessing.ipynb, which handles ingestion and cleaning of raw JSON articles from the data/us-financial-news/ directory. This stage ensures only relevant financial content enters the model.

Filtering by Financial Sections

To maintain domain relevance, the code filters articles using a curated whitelist of section titles including "Reuters: Company News", "Reuters: Business News", and "Press Releases - CNBC". The read_articles() function loads each JSON file, checks the section_title field, and accumulates articles while removing standard English stop words from an external list.

Text Cleaning with spaCy

The repository uses spaCy for sentence-level tokenization and lemmatization. The English pipeline (en) is loaded with the NER component disabled to improve processing speed on large volumes of text.

The clean_doc() function performs aggressive normalization:

  • Converts text to lowercase and removes digits
  • Strips punctuation, short tokens, and pronoun lemmas (-PRON-)
  • Lemmatizes remaining tokens using spaCy's linguistic pipelines

Processing occurs in parallel via nlp.pipe() with n_process=8 and batch_size=100 for efficient throughput on the multi-gigabyte corpus.

N-gram Construction

To capture multi-word financial expressions like "interest rates" or "market volatility", the implementation uses gensim's Phrases and Phraser classes. These detect statistically significant bigrams and trigrams from the token stream using a threshold of 100 and minimum count of 10, smoothing over rare co-occurrences that could introduce noise.

Building the LDA Pipeline

The core modeling logic resides in 15_topic_modeling/07_lda_financial_news.ipynb, which transforms cleaned text into interpretable topics.

TF-IDF Vectorization

Rather than raw count vectors, the pipeline uses scikit-learn's TfidfVectorizer to weight terms by their importance within documents relative to the corpus. Key parameters include:

  • min_df=0.005 and max_df=0.1 to prune extremely rare and overly common terms
  • ngram_range=(1, 1) to use unigrams (following the preprocessing n-gram stage)
  • stop_words='english' as a secondary filter

This produces a sparse document-term matrix (dtm) optimized for financial vocabulary density.

Converting to Gensim Format

Since gensim's LDA implementation requires specific data structures, the code converts the scikit-learn matrix using Sparse2Corpus with documents_columns=False. A mapping dictionary is created via Dictionary.from_corpus() using the feature names from vectorizer.get_feature_names_out() to maintain token-to-ID alignment.

Training the LDA Model

The LdaModel is configured for robust convergence on financial text with these parameters:

  • num_topics (typically explored between 5 and 25)
  • chunksize set to the full document count for batch processing
  • alpha='auto' and eta='auto' for data-driven Dirichlet priors
  • passes=10 and iterations=50 to ensure sufficient traversal of the corpus
  • minimum_probability=0.01 to filter negligible topic weights

Setting random_state=42 ensures reproducible results across training runs.

Evaluation Metrics

The implementation computes two standard metrics to assess model quality:

  • Perplexity: Calculated as 2 ** (-lda.log_perplexity(corpus)) to measure how well the model predicts a sample; lower values indicate better generalization
  • Coherence: Using the u_mass metric to score topic interpretability based on word co-occurrence statistics

Helper functions show_word_list() and show_coherence() generate heatmaps plotting the top N words per topic alongside their probability distributions.

Complete Implementation Example

The following script consolidates the entire workflow into a single executable pipeline. Ensure you have downloaded the US Financial News dataset to data/us-financial-news/ and installed the required dependencies (spacy, gensim, scikit-learn, pyLDAvis).

import json
import warnings
from pathlib import Path
from collections import Counter

import spacy
import numpy as np
import pandas as pd
import pyLDAvis
from gensim.models import LdaModel, Phrases, Phraser
from gensim.corpora import Dictionary
from gensim.matutils import Sparse2Corpus
from pyLDAvis.gensim_models import prepare
from sklearn.feature_extraction.text import TfidfVectorizer

warnings.filterwarnings("ignore")

# -------------------------------------------------

# 1. Load & filter raw articles

# -------------------------------------------------

data_path = Path("data/us-financial-news")
section_titles = [
    "Press Releases - CNBC", "Reuters: Company News", "Reuters: World News",
    "Reuters: Business News", "Reuters: Financial Services and Real Estate",
    "Top News and Analysis (pro)", "Reuters: Top News",
    "The Wall Street Journal & Breaking News, Business, Financial and Economic News, World News and Video",
    "Business & Financial News, U.S & International Breaking News | Reuters",
    "Reuters: Money News", "Reuters: Technology News",
]

stop_words = set(
    pd.read_csv(
        "http://ir.dcs.gla.ac.uk/resources/linguistic_utils/stop_words",
        header=None,
        squeeze=True,
    )
)

def read_articles():
    articles, cnt = [], Counter()
    for f in data_path.rglob("*.json"):
        article = json.load(f.open())
        if article["thread"]["section_title"] in set(section_titles):
            tokens = article["text"].lower().split()
            cnt.update(tokens)
            articles.append(" ".join(t for t in tokens if t not in stop_words))
    return articles, cnt

articles, _ = read_articles()

# -------------------------------------------------

# 2. SpaCy cleaning

# -------------------------------------------------

nlp = spacy.load("en", disable=["ner"])
nlp.max_length = 6_000_000

def clean_doc(doc):
    return " ".join(
        t.lemma_
        for t in doc
        if not (
            t.is_stop
            or t.is_digit
            or not t.is_alpha
            or t.is_punct
            or t.is_space
            or t.lemma_ == "-PRON-"
        )
    )

cleaned = [
    clean_doc(doc) for doc in nlp.pipe(articles, batch_size=100, n_process=8)
]

# -------------------------------------------------

# 3. Build n-grams

# -------------------------------------------------

sentences = [sentence.split() for sentence in cleaned]
phrases = Phrases(sentences, threshold=100, min_count=10)
bigram = Phraser(phrases)
tokenised = [bigram[sentence] for sentence in sentences]

# -------------------------------------------------

# 4. TF-IDF vectorisation

# -------------------------------------------------

vectorizer = TfidfVectorizer(
    stop_words="english", min_df=0.005, max_df=0.1, ngram_range=(1, 1)
)
dtm = vectorizer.fit_transform([" ".join(t) for t in tokenised])
tokens = vectorizer.get_feature_names_out()

# -------------------------------------------------

# 5. Convert to gensim corpus & dictionary

# -------------------------------------------------

corpus = Sparse2Corpus(dtm, documents_columns=False)
id2word = dict(enumerate(tokens))
dictionary = Dictionary.from_corpus(corpus, id2word=id2word)

# -------------------------------------------------

# 6. Train LDA (example: 15 topics)

# -------------------------------------------------

lda = LdaModel(
    corpus=corpus,
    id2word=id2word,
    num_topics=15,
    chunksize=len(tokenised),
    passes=10,
    iterations=50,
    alpha="auto",
    eta="auto",
    random_state=42,
)

# -------------------------------------------------

# 7. Evaluate & visualise

# -------------------------------------------------

perplexity = 2 ** (-lda.log_perplexity(corpus))
print(f"Perplexity: {perplexity:.2f}")

vis = prepare(lda, corpus, dictionary, mds="tsne")
pyLDAvis.display(vis)  # Inline in notebooks

# pyLDAvis.save_html(vis, "financial_news_lda.html")  # Export static HTML

Running this script executes the full pipeline from raw JSON to trained model, outputting perplexity metrics and launching an interactive browser visualization where you can inspect topic-term relevance and document-topic distributions.

Summary

  • Preprocessing matters: Filter articles by financial section titles and use spaCy with disabled NER for efficient lemmatization of market-specific language.
  • Hybrid vectorization works best: Combine scikit-learn's TfidfVectorizer with gensim's Sparse2Corpus to leverage TF-IDF weighting within the LDA framework.
  • Auto-tune your priors: Setting alpha='auto' and eta='auto' allows the model to learn optimal Dirichlet priors from the financial corpus rather than using fixed defaults.
  • Evaluate rigorously: Use both perplexity (statistical fit) and coherence (human interpretability) to select the optimal num_topics between 5 and 25.
  • Visualize interactively: Export results using pyLDAvis.prepare() with mds="tsne" to explore how financial themes cluster in semantic space.

Frequently Asked Questions

What is the optimal number of topics for financial news LDA?

According to the 15_topic_modeling/07_lda_financial_news.ipynb implementation, you should train models with num_topics ranging from 5 to 25 and select the value that maximizes coherence scores while minimizing perplexity. Financial news typically supports 10-20 distinct themes before topic overlap becomes significant.

Why use TF-IDF before LDA instead of raw bag-of-words?

The repository uses TfidfVectorizer with min_df=0.005 and max_df=0.1 to suppress corpus-specific noise—such as ticker symbols appearing in every article or rare technical jargon—that would otherwise dominate the Dirichlet distributions. TF-IDF weighting improves topic distinctiveness by focusing on discriminative financial terms rather than high-frequency market boilerplate.

How does the alpha='auto' parameter affect topic modeling results?

Setting alpha='auto' in gensim.models.LdaModel enables the model to learn asymmetric document-topic priors from the data rather than assuming a uniform 1/K distribution. For financial news, this allows the algorithm to automatically identify that certain articles (e.g., earnings reports) concentrate heavily in specific topics while others (e.g., general market summaries) distribute evenly across themes.

Can this pipeline handle real-time streaming financial news?

While the notebooks in stefan-jansen/machine-learning-for-trading process static JSON archives, the underlying functions are stateless and can be adapted for streaming. You would maintain the fitted TfidfVectorizer and trained LdaModel in memory, then call lda.get_document_topics() on the fly after applying the same clean_doc() preprocessing to incoming articles.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →