How to Extract Alpha Factors Using TA-Lib in Python: A Complete Guide
You extract alpha factors using TA-Lib by loading OHLCV data into a DataSource class that computes vectorized technical indicators—RSI, MACD, ATR, Stochastic, and Ultimate Oscillator—via TA-Lib functions, cleans the results, and standardizes them using sklearn preprocessing before downstream model consumption.
The stefan-jansen/machine-learning-for-trading repository demonstrates production-grade alpha factor extraction for algorithmic trading strategies. This guide explains how to leverage TA-Lib within the repository's architecture to transform raw price data into normalized technical indicators ready for reinforcement learning and supervised models.
The Alpha Factor Pipeline in trading_env.py
The core implementation resides in 22_deep_reinforcement_learning/trading_env.py within the DataSource class. The preprocess_data() method orchestrates the transformation of raw market data into a standardized feature matrix.
Loading Raw Price Data
The pipeline begins by reading daily adjusted close prices, volume, and high/low values from an HDF5 store named assets.h5. These raw columns provide the foundation for all subsequent technical calculations.
Computing Multi-Horizon Returns
Before applying TA-Lib indicators, the system calculates simple percentage changes across multiple lookback periods. The resulting columns (ret_2, ret_5, ret_10, ret_21) capture momentum at different time scales alongside the immediate returns target variable.
Adding TA-Lib Technical Indicators
Between lines 92–99, the repository implements five core TA-Lib indicators as new DataFrame columns:
- RSI:
talib.STOCHRSI(close)[1]— extracts the signal line from the Stochastic RSI - MACD:
talib.MACD(close)[1]— captures the MACD signal line - ATR:
talib.ATR(high, low, close)— measures volatility via Average True Range - Stochastic:
slowd - slowkdifference derived fromtalib.STOCH(high, low, close) - Ultimate Oscillator:
talib.ULTOSC(high, low, close)— combines momentum across multiple periods
Intermediate outputs that are not needed for the final feature set are discarded to conserve memory.
Data Cleaning and Normalization
After factor calculation, the temporary raw price columns (high, low, close, volume) are dropped. Infinite values are replaced with NaN and any remaining missing rows are removed. Finally, sklearn.preprocessing.scale standardizes all alpha factors to zero-mean and unit-variance, while preserving the returns column in its original scale for supervised learning targets.
Extracting Built-In Alpha Factors
To generate the complete feature matrix using the repository's default configuration:
import pandas as pd
from trading_env import DataSource
# Initialise the data source (uses AAPL by default)
source = DataSource(ticker='AAPL')
# The DataFrame now contains the alpha factors
alpha_factors = source.data.copy()
print(alpha_factors.head())
The resulting DataFrame columns include: returns, ret_2, ret_5, ret_10, ret_21, rsi, macd, atr, stoch, ultosc, and additional scaled features (unless normalize=False).
Extending with Custom TA-Lib Indicators
The architecture cleanly separates data ingestion from feature engineering, allowing you to subclass DataSource and append additional TA-Lib calculations. For example, to include Bollinger Band width:
import talib
import pandas as pd
from trading_env import DataSource
class CustomDataSource(DataSource):
def preprocess_data(self):
# Run the original preprocessing
super().preprocess_data()
# Compute Bollinger Band width
upper, middle, lower = talib.BBANDS(
self.data['close'],
timeperiod=20,
nbdevup=2,
nbdevdn=2,
matype=0
)
self.data['bb_width'] = (upper - lower) / middle
# Re‑scale the new column if normalisation is requested
if self.normalize:
self.data['bb_width'] = (self.data['bb_width'] - self.data['bb_width'].mean()) / self.data['bb_width'].std()
# Use the extended source
source = CustomDataSource(ticker='AAPL')
print(source.data[['bb_width']].head())
The extra column bb_width becomes part of the observation space automatically.
Preparing Features for Machine Learning Models
Once extracted, separate the feature matrix from the target variable:
# Get the numpy array of factors (excluding the target return)
X = source.data.drop(columns='returns').values
y = source.data['returns'].values
Now X can be fed to any scikit-learn model, TensorFlow network, or reinforcement-learning policy.
Summary
- The DataSource class in
22_deep_reinforcement_learning/trading_env.pyprovides a complete alpha factor extraction pipeline using TA-Lib - Five technical indicators (RSI, MACD, ATR, Stochastic, Ultimate Oscillator) are computed in
preprocess_data()around lines 92–99 - Features are standardized using sklearn.preprocessing.scale while preserving raw returns for targets
- The modular architecture supports easy extension via subclassing for custom technical indicators
Frequently Asked Questions
What TA-Lib indicators does the repository use by default?
The repository calculates five core indicators: RSI (Stochastic RSI signal), MACD (signal line), ATR (Average True Range), Stochastic (slow difference), and Ultimate Oscillator. These are implemented in the preprocess_data() method of the DataSource class.
How does the repository handle missing or infinite values?
After computing TA-Lib indicators, the code replaces infinite values with NaN using standard pandas operations, then drops any rows containing remaining missing values. This ensures the final feature matrix contains clean, numeric data for model training.
Can I use this DataSource class for non-reinforcement learning models?
Yes. While the class resides in the deep reinforcement learning module, it functions as a general-purpose feature engineering pipeline. You can extract the source.data DataFrame and use it for linear models, tree-based classifiers, or neural networks by separating features (X) from targets (y).
Where is the raw price data stored in the repository structure?
The DataSource constructor reads from an HDF5 file named assets.h5, which contains daily adjusted close prices along with high, low, and volume data. This file is typically generated by upstream data preparation scripts in the repository's workflow modules.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →