How to Extract Alpha Factors Using TA-Lib in Python: A Complete Guide

You extract alpha factors using TA-Lib by loading OHLCV data into a DataSource class that computes vectorized technical indicators—RSI, MACD, ATR, Stochastic, and Ultimate Oscillator—via TA-Lib functions, cleans the results, and standardizes them using sklearn preprocessing before downstream model consumption.

The stefan-jansen/machine-learning-for-trading repository demonstrates production-grade alpha factor extraction for algorithmic trading strategies. This guide explains how to leverage TA-Lib within the repository's architecture to transform raw price data into normalized technical indicators ready for reinforcement learning and supervised models.

The Alpha Factor Pipeline in trading_env.py

The core implementation resides in 22_deep_reinforcement_learning/trading_env.py within the DataSource class. The preprocess_data() method orchestrates the transformation of raw market data into a standardized feature matrix.

Loading Raw Price Data

The pipeline begins by reading daily adjusted close prices, volume, and high/low values from an HDF5 store named assets.h5. These raw columns provide the foundation for all subsequent technical calculations.

Computing Multi-Horizon Returns

Before applying TA-Lib indicators, the system calculates simple percentage changes across multiple lookback periods. The resulting columns (ret_2, ret_5, ret_10, ret_21) capture momentum at different time scales alongside the immediate returns target variable.

Adding TA-Lib Technical Indicators

Between lines 92–99, the repository implements five core TA-Lib indicators as new DataFrame columns:

  • RSI: talib.STOCHRSI(close)[1] — extracts the signal line from the Stochastic RSI
  • MACD: talib.MACD(close)[1] — captures the MACD signal line
  • ATR: talib.ATR(high, low, close) — measures volatility via Average True Range
  • Stochastic: slowd - slowk difference derived from talib.STOCH(high, low, close)
  • Ultimate Oscillator: talib.ULTOSC(high, low, close) — combines momentum across multiple periods

Intermediate outputs that are not needed for the final feature set are discarded to conserve memory.

Data Cleaning and Normalization

After factor calculation, the temporary raw price columns (high, low, close, volume) are dropped. Infinite values are replaced with NaN and any remaining missing rows are removed. Finally, sklearn.preprocessing.scale standardizes all alpha factors to zero-mean and unit-variance, while preserving the returns column in its original scale for supervised learning targets.

Extracting Built-In Alpha Factors

To generate the complete feature matrix using the repository's default configuration:

import pandas as pd
from trading_env import DataSource

# Initialise the data source (uses AAPL by default)

source = DataSource(ticker='AAPL')

# The DataFrame now contains the alpha factors

alpha_factors = source.data.copy()
print(alpha_factors.head())

The resulting DataFrame columns include: returns, ret_2, ret_5, ret_10, ret_21, rsi, macd, atr, stoch, ultosc, and additional scaled features (unless normalize=False).

Extending with Custom TA-Lib Indicators

The architecture cleanly separates data ingestion from feature engineering, allowing you to subclass DataSource and append additional TA-Lib calculations. For example, to include Bollinger Band width:

import talib
import pandas as pd
from trading_env import DataSource

class CustomDataSource(DataSource):
    def preprocess_data(self):
        # Run the original preprocessing

        super().preprocess_data()

        # Compute Bollinger Band width

        upper, middle, lower = talib.BBANDS(
            self.data['close'],
            timeperiod=20,
            nbdevup=2,
            nbdevdn=2,
            matype=0
        )
        self.data['bb_width'] = (upper - lower) / middle

        # Re‑scale the new column if normalisation is requested

        if self.normalize:
            self.data['bb_width'] = (self.data['bb_width'] - self.data['bb_width'].mean()) / self.data['bb_width'].std()

# Use the extended source

source = CustomDataSource(ticker='AAPL')
print(source.data[['bb_width']].head())

The extra column bb_width becomes part of the observation space automatically.

Preparing Features for Machine Learning Models

Once extracted, separate the feature matrix from the target variable:


# Get the numpy array of factors (excluding the target return)

X = source.data.drop(columns='returns').values
y = source.data['returns'].values

Now X can be fed to any scikit-learn model, TensorFlow network, or reinforcement-learning policy.

Summary

  • The DataSource class in 22_deep_reinforcement_learning/trading_env.py provides a complete alpha factor extraction pipeline using TA-Lib
  • Five technical indicators (RSI, MACD, ATR, Stochastic, Ultimate Oscillator) are computed in preprocess_data() around lines 92–99
  • Features are standardized using sklearn.preprocessing.scale while preserving raw returns for targets
  • The modular architecture supports easy extension via subclassing for custom technical indicators

Frequently Asked Questions

What TA-Lib indicators does the repository use by default?

The repository calculates five core indicators: RSI (Stochastic RSI signal), MACD (signal line), ATR (Average True Range), Stochastic (slow difference), and Ultimate Oscillator. These are implemented in the preprocess_data() method of the DataSource class.

How does the repository handle missing or infinite values?

After computing TA-Lib indicators, the code replaces infinite values with NaN using standard pandas operations, then drops any rows containing remaining missing values. This ensures the final feature matrix contains clean, numeric data for model training.

Can I use this DataSource class for non-reinforcement learning models?

Yes. While the class resides in the deep reinforcement learning module, it functions as a general-purpose feature engineering pipeline. You can extract the source.data DataFrame and use it for linear models, tree-based classifiers, or neural networks by separating features (X) from targets (y).

Where is the raw price data stored in the repository structure?

The DataSource constructor reads from an HDF5 file named assets.h5, which contains daily adjusted close prices along with high, low, and volume data. This file is typically generated by upstream data preparation scripts in the repository's workflow modules.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →