# How to Implement Autoencoders for Conditional Risk Factors: A Deep Learning Approach for Asset Pricing

> Learn to implement autoencoders for conditional risk factors using deep learning for asset pricing. Extract latent risk factors by conditioning on firm characteristics. Explore machine learning for trading.

- Repository: [Stefan Jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading)
- Tags: how-to-guide
- Published: 2026-06-02

---

**Conditional autoencoders extract latent risk factors by learning a low-dimensional representation of asset returns that conditions on firm‑specific characteristics such as beta, momentum, and liquidity.**

The `stefan-jansen/machine-learning-for-trading` repository provides a production‑ready implementation of conditional autoencoders for risk‑factor modeling. This architecture compresses high‑dimensional stock characteristics into latent factors while preserving the predictive relationship between firm attributes and forward‑looking returns.

## Data Preparation and Feature Engineering

The modeling pipeline begins with weekly returns and 15 cross‑sectional characteristics for approximately 4,400 U.S. equities. Proper feature engineering ensures that the autoencoder learns stable, economically meaningful relationships rather than data artifacts.

### Loading Panel Data from HDF5

Raw price data and factor metadata are stored in a hierarchical HDF5 file. The notebook `20_autoencoders_for_conditional_risk_factors/05_conditional_autoencoder_for_asset_pricing_data.ipynb` demonstrates how to merge these sources into a single multi‑index DataFrame indexed by `date` and `ticker`.

```python
import pandas as pd
from pathlib import Path

results_path = Path('results', 'asset_pricing')
data = pd.read_hdf(results_path / 'autoencoder.h5', 'model_data')

```

### Rank‑Normalization of Characteristics

All characteristic columns undergo **rank‑normalization** on each cross‑sectional date. This applies a quantile transform that maps values to the interval `[-1, 1]`, reducing the influence of outliers and ensuring that the autoencoder treats all attributes on a comparable scale. Returns are clipped to `[-1, 1]` and missing values are filled with a sentinel value of `-2` to prevent look‑ahead bias during training.

See cells 93‑99 in the data preparation notebook for the quantile transform implementation that scales features to the target range.

## Model Architecture: The Conditional Autoencoder

The architecture is implemented in `20_autoencoders_for_conditional_risk_factors/06_conditional_autoencoder_for_asset_pricing_model.ipynb` using TensorFlow Keras functional API. Unlike standard autoencoders, this model accepts two inputs: firm characteristics and observed factor exposures.

### Defining Inputs and Latent Factor Space

The model requires dual inputs to separate the information channels:

```python
from tensorflow.keras.layers import Input, Dense, Dot, BatchNormalization
from tensorflow.keras.models import Model

n_tickers = len(data.index.unique('ticker'))
n_characteristics = len([c for c in data.columns if c not in ['returns','returns_fwd']])
n_factors = 3

input_beta = Input((n_tickers, n_characteristics), name='input_beta')
input_factor = Input((n_tickers,), name='input_factor')

```

### Stock Characteristics Network

The characteristics pathway employs a dense hidden layer (default 8 units with ReLU activation) followed by batch normalization. This feeds into a projection layer that outputs `n_factors` (default 3) latent loadings representing the firm's exposure to each risk factor.

```python
def make_model(hidden_units=8, n_factors=3):
    # Characteristics → latent factors

    h = Dense(units=hidden_units, activation='relu', name='hidden_layer')(input_beta)
    h = BatchNormalization(name='batch_norm')(h)
    beta_latent = Dense(units=n_factors, name='output_beta')(h)
    
    # Direct factor input pathway

    factor_latent = Dense(units=n_factors, name='output_factor')(input_factor)
    
    # Predicted return via dot product

    out = Dot(axes=(2,1), name='output_layer')([beta_latent, factor_latent])
    
    model = Model(inputs=[input_beta, input_factor], outputs=out)
    model.compile(loss='mse', optimizer='adam')
    return model

```

The `make_model` function (lines 16‑31 in the model notebook) constructs the full computation graph, utilizing a **Dot product** layer to combine the latent representations of characteristics and factor inputs into return predictions.

## Cross‑Validation for Time Series Panel Data

Standard k‑fold cross‑validation violates temporal ordering in financial time series. The repository provides `MultipleTimeSeriesCV` in [`utils.py`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/utils.py) (lines 18‑31) to generate rolling train‑test splits that respect the causal structure of the data.

```python
from utils import MultipleTimeSeriesCV

cv = MultipleTimeSeriesCV(
    n_splits=5,
    train_period_length=20*52,  # 20 years of weekly data

    test_period_length=1*52,    # 1-year test window

    lookahead=1
)

```

This splitter ensures that training data always precedes validation data, eliminating look‑ahead bias when evaluating out‑of‑sample factor loadings.

## Training the Conditional Autoencoder

For each cross‑validation split, reshape the panel data to match the model's expected tensor dimensions: `(-1, n_tickers, n_characteristics)` for characteristics and `(n_tickers,)` for factor inputs.

```python
characteristics = [c for c in data.columns if c not in ['returns','returns_fwd']]

def get_train_valid_data(df, train_idx, val_idx):
    train, val = df.iloc[train_idx], df.iloc[val_idx]
    
    X1_train = train[characteristics].values.reshape(-1, n_tickers, n_characteristics)
    X1_val = val[characteristics].values.reshape(-1, n_tickers, n_characteristics)
    
    X2_train = train['returns'].unstack('ticker').values
    X2_val = val['returns'].unstack('ticker').values
    
    y_train = train['returns_fwd'].unstack('ticker').values
    y_val = val['returns_fwd'].unstack('ticker').values
    
    return X1_train, X2_train, y_train, X1_val, X2_val, y_val

# Training loop

for fold, (train_idx, test_idx) in enumerate(cv.split(data)):
    X1_tr, X2_tr, y_tr, X1_va, X2_va, y_va = get_train_valid_data(
        data, train_idx, test_idx
    )
    
    model = make_model()
    model.fit([X1_tr, X2_tr], y_tr,
              validation_data=([X1_va, X2_va], y_va),
              epochs=30, batch_size=32, verbose=0)
    
    val_mse = model.evaluate([X1_va, X2_va], y_va, verbose=0)
    print(f'Fold {fold+1} – Validation MSE: {val_mse:.6f}')

```

After training, extract the latent factor loadings from the `output_beta` layer to analyze risk‑factor exposures or construct minimum‑variance portfolios.

## Summary

- **Conditional autoencoders** learn asset‑pricing relationships by jointly modeling firm characteristics and observed returns in a shared latent space.
- **Rank‑normalization** of cross‑sectional features to `[-1, 1]` is essential for stable training across diverse characteristic scales.
- The `make_model` function implements a dual‑pathway architecture with batch normalization and a dot‑product output layer for return prediction.
- **MultipleTimeSeriesCV** in [`utils.py`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/utils.py) provides bias‑free train‑test splits for panel data by enforcing temporal ordering.
- Extractable latent factors (`output_beta`) enable downstream risk management and portfolio construction tasks.

## Frequently Asked Questions

### What distinguishes conditional autoencoders from traditional factor models?

Traditional models like PCA extract latent factors from returns alone, whereas **conditional autoencoders** incorporate firm‑specific characteristics (beta, momentum, liquidity) as conditioning information. This allows the model to learn time‑varying factor loadings that adapt to changing firm attributes, capturing non‑linear relationships that linear regression or static factor models miss.

### Why is rank‑normalization necessary for the input characteristics?

Raw financial characteristics (e.g., market capitalization, book‑to‑market ratios) exhibit extreme skewness and outliers. **Rank‑normalization** applies a quantile transform per cross‑section to map all features into the uniform range `[-1, 1]`, preventing the autoencoder from fitting spurious scale effects and ensuring that the learning process focuses on relative ranking information rather than absolute magnitudes.

### How does MultipleTimeSeriesCV prevent data leakage?

The `MultipleTimeSeriesCV` class (defined in [`utils.py`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/utils.py), lines 18‑31) creates rolling windows where the training set strictly precedes the validation set by at least one period (controlled by the `lookahead` parameter). This structure ensures that no future return information contaminates the training process, which is critical for obtaining unbiased estimates of out‑of‑sample predictive performance in financial forecasting.

### Can this architecture accommodate higher‑frequency data or alternative asset classes?

Yes. While the repository demonstrates weekly equity returns, you can adapt the pipeline by adjusting `n_characteristics` in the `make_model` function and modifying the rank‑normalization step to handle intraday returns or fixed‑income attributes. Ensure that the `train_period_length` and `test_period_length` parameters in `MultipleTimeSeriesCV` are scaled appropriately for your data frequency to maintain sufficient statistical power in each fold.