# How to Implement Topic Modeling for Financial News: A Complete LDA Pipeline

> Learn to implement topic modeling for financial news using LDA. Preprocess, vectorize, train, and visualize your results for actionable insights.

- Repository: [Stefan Jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading)
- Tags: tutorial
- Published: 2026-06-02

---

**You can implement topic modeling for financial news by preprocessing JSON articles with spaCy, vectorizing the corpus using scikit-learn TF-IDF, training a gensim LDA model with auto-tuned hyperparameters, and evaluating results through coherence scores and interactive pyLDAvis visualizations.**

This guide walks you through a production-ready implementation of topic modeling for financial news using Latent Dirichlet Allocation (LDA). Based on the **stefan-jansen/machine-learning-for-trading** repository, this workflow processes thousands of Reuters and CNBC articles to extract latent themes from market-moving text data using the notebooks found in `16_word_embeddings/` and `15_topic_modeling/`.

## Preprocessing Financial News Articles

The pipeline begins in `16_word_embeddings/03_financial_news_preprocessing.ipynb`, which handles ingestion and cleaning of raw JSON articles from the `data/us-financial-news/` directory. This stage ensures only relevant financial content enters the model.

### Filtering by Financial Sections

To maintain domain relevance, the code filters articles using a curated whitelist of section titles including "Reuters: Company News", "Reuters: Business News", and "Press Releases - CNBC". The `read_articles()` function loads each JSON file, checks the `section_title` field, and accumulates articles while removing standard English stop words from an external list.

### Text Cleaning with spaCy

The repository uses **spaCy** for sentence-level tokenization and lemmatization. The English pipeline (`en`) is loaded with the **NER** component disabled to improve processing speed on large volumes of text.

The `clean_doc()` function performs aggressive normalization:
- Converts text to lowercase and removes digits
- Strips punctuation, short tokens, and pronoun lemmas (`-PRON-`)
- Lemmatizes remaining tokens using spaCy's linguistic pipelines

Processing occurs in parallel via `nlp.pipe()` with `n_process=8` and `batch_size=100` for efficient throughput on the multi-gigabyte corpus.

### N-gram Construction

To capture multi-word financial expressions like "interest rates" or "market volatility", the implementation uses **gensim's** `Phrases` and `Phraser` classes. These detect statistically significant bigrams and trigrams from the token stream using a threshold of 100 and minimum count of 10, smoothing over rare co-occurrences that could introduce noise.

## Building the LDA Pipeline

The core modeling logic resides in `15_topic_modeling/07_lda_financial_news.ipynb`, which transforms cleaned text into interpretable topics.

### TF-IDF Vectorization

Rather than raw count vectors, the pipeline uses **scikit-learn's** `TfidfVectorizer` to weight terms by their importance within documents relative to the corpus. Key parameters include:
- `min_df=0.005` and `max_df=0.1` to prune extremely rare and overly common terms
- `ngram_range=(1, 1)` to use unigrams (following the preprocessing n-gram stage)
- `stop_words='english'` as a secondary filter

This produces a sparse document-term matrix (`dtm`) optimized for financial vocabulary density.

### Converting to Gensim Format

Since **gensim's** LDA implementation requires specific data structures, the code converts the scikit-learn matrix using `Sparse2Corpus` with `documents_columns=False`. A mapping dictionary is created via `Dictionary.from_corpus()` using the feature names from `vectorizer.get_feature_names_out()` to maintain token-to-ID alignment.

### Training the LDA Model

The `LdaModel` is configured for robust convergence on financial text with these parameters:
- `num_topics` (typically explored between 5 and 25)
- `chunksize` set to the full document count for batch processing
- `alpha='auto'` and `eta='auto'` for data-driven Dirichlet priors
- `passes=10` and `iterations=50` to ensure sufficient traversal of the corpus
- `minimum_probability=0.01` to filter negligible topic weights

Setting `random_state=42` ensures reproducible results across training runs.

### Evaluation Metrics

The implementation computes two standard metrics to assess model quality:
- **Perplexity**: Calculated as `2 ** (-lda.log_perplexity(corpus))` to measure how well the model predicts a sample; lower values indicate better generalization
- **Coherence**: Using the `u_mass` metric to score topic interpretability based on word co-occurrence statistics

Helper functions `show_word_list()` and `show_coherence()` generate heatmaps plotting the top N words per topic alongside their probability distributions.

## Complete Implementation Example

The following script consolidates the entire workflow into a single executable pipeline. Ensure you have downloaded the US Financial News dataset to `data/us-financial-news/` and installed the required dependencies (`spacy`, `gensim`, `scikit-learn`, `pyLDAvis`).

```python
import json
import warnings
from pathlib import Path
from collections import Counter

import spacy
import numpy as np
import pandas as pd
import pyLDAvis
from gensim.models import LdaModel, Phrases, Phraser
from gensim.corpora import Dictionary
from gensim.matutils import Sparse2Corpus
from pyLDAvis.gensim_models import prepare
from sklearn.feature_extraction.text import TfidfVectorizer

warnings.filterwarnings("ignore")

# -------------------------------------------------

# 1. Load & filter raw articles

# -------------------------------------------------

data_path = Path("data/us-financial-news")
section_titles = [
    "Press Releases - CNBC", "Reuters: Company News", "Reuters: World News",
    "Reuters: Business News", "Reuters: Financial Services and Real Estate",
    "Top News and Analysis (pro)", "Reuters: Top News",
    "The Wall Street Journal & Breaking News, Business, Financial and Economic News, World News and Video",
    "Business & Financial News, U.S & International Breaking News | Reuters",
    "Reuters: Money News", "Reuters: Technology News",
]

stop_words = set(
    pd.read_csv(
        "http://ir.dcs.gla.ac.uk/resources/linguistic_utils/stop_words",
        header=None,
        squeeze=True,
    )
)

def read_articles():
    articles, cnt = [], Counter()
    for f in data_path.rglob("*.json"):
        article = json.load(f.open())
        if article["thread"]["section_title"] in set(section_titles):
            tokens = article["text"].lower().split()
            cnt.update(tokens)
            articles.append(" ".join(t for t in tokens if t not in stop_words))
    return articles, cnt

articles, _ = read_articles()

# -------------------------------------------------

# 2. SpaCy cleaning

# -------------------------------------------------

nlp = spacy.load("en", disable=["ner"])
nlp.max_length = 6_000_000

def clean_doc(doc):
    return " ".join(
        t.lemma_
        for t in doc
        if not (
            t.is_stop
            or t.is_digit
            or not t.is_alpha
            or t.is_punct
            or t.is_space
            or t.lemma_ == "-PRON-"
        )
    )

cleaned = [
    clean_doc(doc) for doc in nlp.pipe(articles, batch_size=100, n_process=8)
]

# -------------------------------------------------

# 3. Build n-grams

# -------------------------------------------------

sentences = [sentence.split() for sentence in cleaned]
phrases = Phrases(sentences, threshold=100, min_count=10)
bigram = Phraser(phrases)
tokenised = [bigram[sentence] for sentence in sentences]

# -------------------------------------------------

# 4. TF-IDF vectorisation

# -------------------------------------------------

vectorizer = TfidfVectorizer(
    stop_words="english", min_df=0.005, max_df=0.1, ngram_range=(1, 1)
)
dtm = vectorizer.fit_transform([" ".join(t) for t in tokenised])
tokens = vectorizer.get_feature_names_out()

# -------------------------------------------------

# 5. Convert to gensim corpus & dictionary

# -------------------------------------------------

corpus = Sparse2Corpus(dtm, documents_columns=False)
id2word = dict(enumerate(tokens))
dictionary = Dictionary.from_corpus(corpus, id2word=id2word)

# -------------------------------------------------

# 6. Train LDA (example: 15 topics)

# -------------------------------------------------

lda = LdaModel(
    corpus=corpus,
    id2word=id2word,
    num_topics=15,
    chunksize=len(tokenised),
    passes=10,
    iterations=50,
    alpha="auto",
    eta="auto",
    random_state=42,
)

# -------------------------------------------------

# 7. Evaluate & visualise

# -------------------------------------------------

perplexity = 2 ** (-lda.log_perplexity(corpus))
print(f"Perplexity: {perplexity:.2f}")

vis = prepare(lda, corpus, dictionary, mds="tsne")
pyLDAvis.display(vis)  # Inline in notebooks

# pyLDAvis.save_html(vis, "financial_news_lda.html")  # Export static HTML

```

Running this script executes the full pipeline from raw JSON to trained model, outputting perplexity metrics and launching an interactive browser visualization where you can inspect topic-term relevance and document-topic distributions.

## Summary

- **Preprocessing matters**: Filter articles by financial section titles and use spaCy with disabled NER for efficient lemmatization of market-specific language.
- **Hybrid vectorization works best**: Combine scikit-learn's `TfidfVectorizer` with gensim's `Sparse2Corpus` to leverage TF-IDF weighting within the LDA framework.
- **Auto-tune your priors**: Setting `alpha='auto'` and `eta='auto'` allows the model to learn optimal Dirichlet priors from the financial corpus rather than using fixed defaults.
- **Evaluate rigorously**: Use both perplexity (statistical fit) and coherence (human interpretability) to select the optimal `num_topics` between 5 and 25.
- **Visualize interactively**: Export results using `pyLDAvis.prepare()` with `mds="tsne"` to explore how financial themes cluster in semantic space.

## Frequently Asked Questions

### What is the optimal number of topics for financial news LDA?

According to the `15_topic_modeling/07_lda_financial_news.ipynb` implementation, you should train models with `num_topics` ranging from 5 to 25 and select the value that maximizes coherence scores while minimizing perplexity. Financial news typically supports 10-20 distinct themes before topic overlap becomes significant.

### Why use TF-IDF before LDA instead of raw bag-of-words?

The repository uses `TfidfVectorizer` with `min_df=0.005` and `max_df=0.1` to suppress corpus-specific noise—such as ticker symbols appearing in every article or rare technical jargon—that would otherwise dominate the Dirichlet distributions. TF-IDF weighting improves topic distinctiveness by focusing on discriminative financial terms rather than high-frequency market boilerplate.

### How does the alpha='auto' parameter affect topic modeling results?

Setting `alpha='auto'` in `gensim.models.LdaModel` enables the model to learn asymmetric document-topic priors from the data rather than assuming a uniform 1/K distribution. For financial news, this allows the algorithm to automatically identify that certain articles (e.g., earnings reports) concentrate heavily in specific topics while others (e.g., general market summaries) distribute evenly across themes.

### Can this pipeline handle real-time streaming financial news?

While the notebooks in **stefan-jansen/machine-learning-for-trading** process static JSON archives, the underlying functions are stateless and can be adapted for streaming. You would maintain the fitted `TfidfVectorizer` and trained `LdaModel` in memory, then call `lda.get_document_topics()` on the fly after applying the same `clean_doc()` preprocessing to incoming articles.