# How to Extract Alpha Factors Using TA-Lib in Python: A Complete Guide

> Easily extract alpha factors with TA-Lib in Python. Learn how to load OHLCV data, compute vectorized indicators like RSI and MACD, clean, and standardize them for your trading models.

- Repository: [Stefan Jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading)
- Tags: how-to-guide
- Published: 2026-06-02

---

**You extract alpha factors using TA-Lib by loading OHLCV data into a DataSource class that computes vectorized technical indicators—RSI, MACD, ATR, Stochastic, and Ultimate Oscillator—via TA-Lib functions, cleans the results, and standardizes them using sklearn preprocessing before downstream model consumption.**

The **stefan-jansen/machine-learning-for-trading** repository demonstrates production-grade alpha factor extraction for algorithmic trading strategies. This guide explains how to leverage TA-Lib within the repository's architecture to transform raw price data into normalized technical indicators ready for reinforcement learning and supervised models.

## The Alpha Factor Pipeline in trading_env.py

The core implementation resides in [`22_deep_reinforcement_learning/trading_env.py`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/22_deep_reinforcement_learning/trading_env.py) within the **DataSource** class. The `preprocess_data()` method orchestrates the transformation of raw market data into a standardized feature matrix.

### Loading Raw Price Data

The pipeline begins by reading daily adjusted close prices, volume, and high/low values from an HDF5 store named `assets.h5`. These raw columns provide the foundation for all subsequent technical calculations.

### Computing Multi-Horizon Returns

Before applying TA-Lib indicators, the system calculates simple percentage changes across multiple lookback periods. The resulting columns (`ret_2`, `ret_5`, `ret_10`, `ret_21`) capture momentum at different time scales alongside the immediate `returns` target variable.

### Adding TA-Lib Technical Indicators

Between lines 92–99, the repository implements five core TA-Lib indicators as new DataFrame columns:

- **RSI**: `talib.STOCHRSI(close)[1]` — extracts the signal line from the Stochastic RSI
- **MACD**: `talib.MACD(close)[1]` — captures the MACD signal line
- **ATR**: `talib.ATR(high, low, close)` — measures volatility via Average True Range
- **Stochastic**: `slowd - slowk` difference derived from `talib.STOCH(high, low, close)`
- **Ultimate Oscillator**: `talib.ULTOSC(high, low, close)` — combines momentum across multiple periods

Intermediate outputs that are not needed for the final feature set are discarded to conserve memory.

### Data Cleaning and Normalization

After factor calculation, the temporary raw price columns (`high`, `low`, `close`, `volume`) are dropped. Infinite values are replaced with `NaN` and any remaining missing rows are removed. Finally, **sklearn.preprocessing.scale** standardizes all alpha factors to zero-mean and unit-variance, while preserving the `returns` column in its original scale for supervised learning targets.

## Extracting Built-In Alpha Factors

To generate the complete feature matrix using the repository's default configuration:

```python
import pandas as pd
from trading_env import DataSource

# Initialise the data source (uses AAPL by default)

source = DataSource(ticker='AAPL')

# The DataFrame now contains the alpha factors

alpha_factors = source.data.copy()
print(alpha_factors.head())

```

The resulting DataFrame columns include: `returns`, `ret_2`, `ret_5`, `ret_10`, `ret_21`, `rsi`, `macd`, `atr`, `stoch`, `ultosc`, and additional scaled features (unless `normalize=False`).

## Extending with Custom TA-Lib Indicators

The architecture cleanly separates data ingestion from feature engineering, allowing you to subclass **DataSource** and append additional TA-Lib calculations. For example, to include Bollinger Band width:

```python
import talib
import pandas as pd
from trading_env import DataSource

class CustomDataSource(DataSource):
    def preprocess_data(self):
        # Run the original preprocessing

        super().preprocess_data()

        # Compute Bollinger Band width

        upper, middle, lower = talib.BBANDS(
            self.data['close'],
            timeperiod=20,
            nbdevup=2,
            nbdevdn=2,
            matype=0
        )
        self.data['bb_width'] = (upper - lower) / middle

        # Re‑scale the new column if normalisation is requested

        if self.normalize:
            self.data['bb_width'] = (self.data['bb_width'] - self.data['bb_width'].mean()) / self.data['bb_width'].std()

# Use the extended source

source = CustomDataSource(ticker='AAPL')
print(source.data[['bb_width']].head())

```

The extra column `bb_width` becomes part of the observation space automatically.

## Preparing Features for Machine Learning Models

Once extracted, separate the feature matrix from the target variable:

```python

# Get the numpy array of factors (excluding the target return)

X = source.data.drop(columns='returns').values
y = source.data['returns'].values

```

Now `X` can be fed to any scikit-learn model, TensorFlow network, or reinforcement-learning policy.

## Summary

- The **DataSource** class in [`22_deep_reinforcement_learning/trading_env.py`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/22_deep_reinforcement_learning/trading_env.py) provides a complete alpha factor extraction pipeline using TA-Lib
- Five technical indicators (RSI, MACD, ATR, Stochastic, Ultimate Oscillator) are computed in `preprocess_data()` around lines 92–99
- Features are standardized using **sklearn.preprocessing.scale** while preserving raw returns for targets
- The modular architecture supports easy extension via subclassing for custom technical indicators

## Frequently Asked Questions

### What TA-Lib indicators does the repository use by default?

The repository calculates five core indicators: **RSI** (Stochastic RSI signal), **MACD** (signal line), **ATR** (Average True Range), **Stochastic** (slow difference), and **Ultimate Oscillator**. These are implemented in the `preprocess_data()` method of the `DataSource` class.

### How does the repository handle missing or infinite values?

After computing TA-Lib indicators, the code replaces infinite values with `NaN` using standard pandas operations, then drops any rows containing remaining missing values. This ensures the final feature matrix contains clean, numeric data for model training.

### Can I use this DataSource class for non-reinforcement learning models?

Yes. While the class resides in the deep reinforcement learning module, it functions as a general-purpose feature engineering pipeline. You can extract the `source.data` DataFrame and use it for linear models, tree-based classifiers, or neural networks by separating features (`X`) from targets (`y`).

### Where is the raw price data stored in the repository structure?

The `DataSource` constructor reads from an HDF5 file named `assets.h5`, which contains daily adjusted close prices along with high, low, and volume data. This file is typically generated by upstream data preparation scripts in the repository's workflow modules.