How to Implement Autoencoders for Conditional Risk Factors: A Deep Learning Approach for Asset Pricing

Conditional autoencoders extract latent risk factors by learning a low-dimensional representation of asset returns that conditions on firm‑specific characteristics such as beta, momentum, and liquidity.

The stefan-jansen/machine-learning-for-trading repository provides a production‑ready implementation of conditional autoencoders for risk‑factor modeling. This architecture compresses high‑dimensional stock characteristics into latent factors while preserving the predictive relationship between firm attributes and forward‑looking returns.

Data Preparation and Feature Engineering

The modeling pipeline begins with weekly returns and 15 cross‑sectional characteristics for approximately 4,400 U.S. equities. Proper feature engineering ensures that the autoencoder learns stable, economically meaningful relationships rather than data artifacts.

Loading Panel Data from HDF5

Raw price data and factor metadata are stored in a hierarchical HDF5 file. The notebook 20_autoencoders_for_conditional_risk_factors/05_conditional_autoencoder_for_asset_pricing_data.ipynb demonstrates how to merge these sources into a single multi‑index DataFrame indexed by date and ticker.

import pandas as pd
from pathlib import Path

results_path = Path('results', 'asset_pricing')
data = pd.read_hdf(results_path / 'autoencoder.h5', 'model_data')

Rank‑Normalization of Characteristics

All characteristic columns undergo rank‑normalization on each cross‑sectional date. This applies a quantile transform that maps values to the interval [-1, 1], reducing the influence of outliers and ensuring that the autoencoder treats all attributes on a comparable scale. Returns are clipped to [-1, 1] and missing values are filled with a sentinel value of -2 to prevent look‑ahead bias during training.

See cells 93‑99 in the data preparation notebook for the quantile transform implementation that scales features to the target range.

Model Architecture: The Conditional Autoencoder

The architecture is implemented in 20_autoencoders_for_conditional_risk_factors/06_conditional_autoencoder_for_asset_pricing_model.ipynb using TensorFlow Keras functional API. Unlike standard autoencoders, this model accepts two inputs: firm characteristics and observed factor exposures.

Defining Inputs and Latent Factor Space

The model requires dual inputs to separate the information channels:

from tensorflow.keras.layers import Input, Dense, Dot, BatchNormalization
from tensorflow.keras.models import Model

n_tickers = len(data.index.unique('ticker'))
n_characteristics = len([c for c in data.columns if c not in ['returns','returns_fwd']])
n_factors = 3

input_beta = Input((n_tickers, n_characteristics), name='input_beta')
input_factor = Input((n_tickers,), name='input_factor')

Stock Characteristics Network

The characteristics pathway employs a dense hidden layer (default 8 units with ReLU activation) followed by batch normalization. This feeds into a projection layer that outputs n_factors (default 3) latent loadings representing the firm's exposure to each risk factor.

def make_model(hidden_units=8, n_factors=3):
    # Characteristics → latent factors

    h = Dense(units=hidden_units, activation='relu', name='hidden_layer')(input_beta)
    h = BatchNormalization(name='batch_norm')(h)
    beta_latent = Dense(units=n_factors, name='output_beta')(h)
    
    # Direct factor input pathway

    factor_latent = Dense(units=n_factors, name='output_factor')(input_factor)
    
    # Predicted return via dot product

    out = Dot(axes=(2,1), name='output_layer')([beta_latent, factor_latent])
    
    model = Model(inputs=[input_beta, input_factor], outputs=out)
    model.compile(loss='mse', optimizer='adam')
    return model

The make_model function (lines 16‑31 in the model notebook) constructs the full computation graph, utilizing a Dot product layer to combine the latent representations of characteristics and factor inputs into return predictions.

Cross‑Validation for Time Series Panel Data

Standard k‑fold cross‑validation violates temporal ordering in financial time series. The repository provides MultipleTimeSeriesCV in utils.py (lines 18‑31) to generate rolling train‑test splits that respect the causal structure of the data.

from utils import MultipleTimeSeriesCV

cv = MultipleTimeSeriesCV(
    n_splits=5,
    train_period_length=20*52,  # 20 years of weekly data

    test_period_length=1*52,    # 1-year test window

    lookahead=1
)

This splitter ensures that training data always precedes validation data, eliminating look‑ahead bias when evaluating out‑of‑sample factor loadings.

Training the Conditional Autoencoder

For each cross‑validation split, reshape the panel data to match the model's expected tensor dimensions: (-1, n_tickers, n_characteristics) for characteristics and (n_tickers,) for factor inputs.

characteristics = [c for c in data.columns if c not in ['returns','returns_fwd']]

def get_train_valid_data(df, train_idx, val_idx):
    train, val = df.iloc[train_idx], df.iloc[val_idx]
    
    X1_train = train[characteristics].values.reshape(-1, n_tickers, n_characteristics)
    X1_val = val[characteristics].values.reshape(-1, n_tickers, n_characteristics)
    
    X2_train = train['returns'].unstack('ticker').values
    X2_val = val['returns'].unstack('ticker').values
    
    y_train = train['returns_fwd'].unstack('ticker').values
    y_val = val['returns_fwd'].unstack('ticker').values
    
    return X1_train, X2_train, y_train, X1_val, X2_val, y_val

# Training loop

for fold, (train_idx, test_idx) in enumerate(cv.split(data)):
    X1_tr, X2_tr, y_tr, X1_va, X2_va, y_va = get_train_valid_data(
        data, train_idx, test_idx
    )
    
    model = make_model()
    model.fit([X1_tr, X2_tr], y_tr,
              validation_data=([X1_va, X2_va], y_va),
              epochs=30, batch_size=32, verbose=0)
    
    val_mse = model.evaluate([X1_va, X2_va], y_va, verbose=0)
    print(f'Fold {fold+1} – Validation MSE: {val_mse:.6f}')

After training, extract the latent factor loadings from the output_beta layer to analyze risk‑factor exposures or construct minimum‑variance portfolios.

Summary

  • Conditional autoencoders learn asset‑pricing relationships by jointly modeling firm characteristics and observed returns in a shared latent space.
  • Rank‑normalization of cross‑sectional features to [-1, 1] is essential for stable training across diverse characteristic scales.
  • The make_model function implements a dual‑pathway architecture with batch normalization and a dot‑product output layer for return prediction.
  • MultipleTimeSeriesCV in utils.py provides bias‑free train‑test splits for panel data by enforcing temporal ordering.
  • Extractable latent factors (output_beta) enable downstream risk management and portfolio construction tasks.

Frequently Asked Questions

What distinguishes conditional autoencoders from traditional factor models?

Traditional models like PCA extract latent factors from returns alone, whereas conditional autoencoders incorporate firm‑specific characteristics (beta, momentum, liquidity) as conditioning information. This allows the model to learn time‑varying factor loadings that adapt to changing firm attributes, capturing non‑linear relationships that linear regression or static factor models miss.

Why is rank‑normalization necessary for the input characteristics?

Raw financial characteristics (e.g., market capitalization, book‑to‑market ratios) exhibit extreme skewness and outliers. Rank‑normalization applies a quantile transform per cross‑section to map all features into the uniform range [-1, 1], preventing the autoencoder from fitting spurious scale effects and ensuring that the learning process focuses on relative ranking information rather than absolute magnitudes.

How does MultipleTimeSeriesCV prevent data leakage?

The MultipleTimeSeriesCV class (defined in utils.py, lines 18‑31) creates rolling windows where the training set strictly precedes the validation set by at least one period (controlled by the lookahead parameter). This structure ensures that no future return information contaminates the training process, which is critical for obtaining unbiased estimates of out‑of‑sample predictive performance in financial forecasting.

Can this architecture accommodate higher‑frequency data or alternative asset classes?

Yes. While the repository demonstrates weekly equity returns, you can adapt the pipeline by adjusting n_characteristics in the make_model function and modifying the rank‑normalization step to handle intraday returns or fixed‑income attributes. Ensure that the train_period_length and test_period_length parameters in MultipleTimeSeriesCV are scaled appropriately for your data frequency to maintain sufficient statistical power in each fold.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →