# How to Implement Autoencoders for Asset Pricing: A Complete Guide to Conditional Risk Factors

> Implement autoencoders for asset pricing to extract conditional risk factors. Discover hidden data-driven factors by learning compressed latent representations of asset returns.

- Repository: [Stefan Jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading)
- Tags: tutorial
- Published: 2026-06-02

---

**Autoencoders extract conditional risk factors by learning compressed latent representations of asset returns while conditioning on macroeconomic variables, enabling you to discover hidden data-driven factors that traditional linear models miss.**

This guide walks you through implementing autoencoders for asset pricing using the **machine-learning-for-trading** repository by Stefan Jansen. You will learn how to build conditional autoencoders that combine cross-sectional equity returns with macro features to generate latent risk factors suitable for portfolio construction and cross-sectional return prediction.

## Step 1: Prepare the Return Panel and Macro Features

Before building the neural network, construct a clean panel of asset returns aligned with macroeconomic covariates. In `20_autoencoders_for_conditional_risk_factors/04_build_us_stock_dataset.ipynb`, the pipeline aggregates daily returns for a broad equity universe and merges them with macro features such as inflation, interest rates, and volatility indices.

The data structure follows a 3-dimensional tensor convention:

- **Axis 0**: Time (dates)
- **Axis 1**: Assets (cross-section)
- **Axis 2**: Features (returns + conditional variables)

```python
import pandas as pd
import numpy as np

# Load prepared datasets from the repository

returns = pd.read_pickle('data/returns_us_stock.pkl')    # Shape: (dates, assets)

macro = pd.read_pickle('data/macro_features.pkl')        # Shape: (dates, macro_features)

# Align and split chronologically (no shuffle to preserve time structure)

from sklearn.model_selection import train_test_split
X_ret, X_cond = returns.values, macro.values
X_ret_train, X_ret_val, X_cond_train, X_cond_val = train_test_split(
    X_ret, X_cond, test_size=0.2, shuffle=False)

```

## Step 2: Build the Base Encoder-Decoder Architecture

The foundational model in `20_autoencoders_for_conditional_risk_factors/01_deep_autoencoders.ipynb` implements a dense autoencoder that compresses the high-dimensional return vector into a low-dimensional latent space. The encoder flattens the input and passes it through two hidden layers before producing the latent code `z`.

```python
import tensorflow as tf

latent_dim = 5
n_assets = X_ret.shape[1]

# Encoder: (batch, n_assets) -> (batch, latent_dim)

encoder_input = tf.keras.Input(shape=(n_assets,), name='returns')
x = tf.keras.layers.Dense(128, activation='relu')(encoder_input)
x = tf.keras.layers.Dense(64, activation='relu')(x)
z = tf.keras.layers.Dense(latent_dim, name='latent')(x)
encoder = tf.keras.Model(encoder_input, z, name='encoder')

```

The decoder mirrors this architecture in reverse, reconstructing the original return vector from the latent representation:

```python
decoder_input = tf.keras.Input(shape=(latent_dim,), name='decoder_input')
x = tf.keras.layers.Dense(64, activation='relu')(decoder_input)
x = tf.keras.layers.Dense(128, activation='relu')(x)
recon = tf.keras.layers.Dense(n_assets, activation='linear', name='reconstruction')(x)
decoder = tf.keras.Model(decoder_input, recon, name='decoder')

```

## Step 3: Implement the Conditional Autoencoder

The critical innovation for asset pricing appears in `20_autoencoders_for_conditional_risk_factors/06_conditional_autoencoder_for_asset_pricing_model.ipynb`. Here, the **conditional autoencoder** feeds macroeconomic variables to both the encoder and decoder, forcing the latent representation to capture information orthogonal to known economic conditions.

The architecture concatenates asset returns with broadcasted macro features:

```python

# Expand dimensions for broadcasting: (dates, 1, n_assets) and (dates, macro_dim, 1)

X_ret_expanded = returns.values[..., None]           # Shape: (batch, n_assets, 1)

X_cond_expanded = macro.values[:, None, ...]         # Broadcast macro across assets

# Concatenate along feature dimension

X = np.concatenate([X_ret_expanded, X_cond_expanded], axis=-1)

```

Build the conditional model by injecting macro variables into the decoder input:

```python
cond_dim = X_cond.shape[1]

# Encoder input: raw returns only

returns_input = tf.keras.Input(shape=(n_assets, 1), name='returns_input')
x = tf.keras.layers.Flatten()(returns_input)
x = tf.keras.layers.Dense(128, activation='relu')(x)
latent = tf.keras.layers.Dense(latent_dim, name='latent_vector')(x)

# Decoder input: latent code + conditional variables

cond_input = tf.keras.Input(shape=(cond_dim,), name='macro_input')
decoder_concat = tf.keras.layers.Concatenate()([latent, cond_input])
x = tf.keras.layers.Dense(128, activation='relu')(decoder_concat)
reconstruction = tf.keras.layers.Dense(n_assets, activation='linear')(x)

# Full conditional autoencoder

conditional_ae = tf.keras.Model(
    inputs=[returns_input, cond_input], 
    outputs=reconstruction,
    name='conditional_autoencoder'
)
conditional_ae.compile(optimizer='adam', loss='mse')

```

By reconstructing returns conditioned on macro states, the latent factors `z` represent **conditional risk factors** that vary with the economic environment.

## Step 4: Train with Regularization and Early Stopping

Train the model using mean squared error (MSE) loss with early stopping to prevent overfitting on noisy return data. The repository demonstrates this pattern across `01_deep_autoencoders.ipynb` through `06_conditional_autoencoder_for_asset_pricing_model.ipynb`.

```python

# Early stopping configuration

early_stop = tf.keras.callbacks.EarlyStopping(
    monitor='val_loss',
    patience=10,
    restore_best_weights=True,
    verbose=1
)

# Training with validation holdout

history = conditional_ae.fit(
    [X_ret_train[..., None], X_cond_train],
    X_ret_train,
    epochs=200,
    batch_size=256,
    validation_data=([X_ret_val[..., None], X_cond_val], X_ret_val),
    callbacks=[early_stop],
    verbose=1
)

```

For enhanced generalization, consider the advanced variants in `20_autoencoders_for_conditional_risk_factors/02_convolutional_denoising_autoencoders.ipynb` (which adds Conv1D layers and input noise) and `03_variational_autoencoder.ipynb` (which imposes a KL-divergence penalty for stochastic latent spaces).

## Step 5: Extract Factors and Evaluate with Alphalens

Once trained, extract the latent factors by passing the full dataset through the encoder. These factors serve as inputs to standard asset-pricing evaluations such as cross-sectional regressions or portfolio sorts.

```python

# Extract conditional risk factors

latent_factors = encoder.predict(X_ret[..., None])
factor_df = pd.DataFrame(
    latent_factors,
    index=returns.index,
    columns=[f'latent_{i}' for i in range(latent_dim)]
)

# Save for Alphalens analysis

factor_df.to_pickle('data/conditional_latent_factors.pkl')

```

Evaluate the predictive power of these learned factors using `20_autoencoders_for_conditional_risk_factors/07_alphalens_analysis.ipynb`. This notebook maps the latent series into the Alphalens framework to compute:
- **Factor returns** (quantile-wise performance)
- **Information coefficient (IC)** (rank correlation with forward returns)
- **Turnover analysis** (stability of factor exposures)

The latent factors integrate directly into existing pipelines—treat them exactly like traditional Fama-French factors or momentum signals when running OLS regressions or constructing mean-variance efficient portfolios.

## Summary

- **Conditional autoencoders** combine asset returns with macroeconomic variables to extract risk factors that adapt to changing economic regimes.
- The implementation in `stefan-jansen/machine-learning-for-trading` provides a complete pipeline from raw price data to factor evaluation via seven interconnected notebooks.
- **Key architectural choice**: Feed macro variables to the decoder (or concatenate with latent codes) to enforce that latent representations capture idiosyncratic risk orthogonal to known conditions.
- **Training best practices**: Use early stopping, batch normalization, and consider denoising or variational variants to improve out-of-sample generalization.
- **Evaluation**: Extracted latent factors from `06_conditional_autoencoder_for_asset_pricing_model.ipynb` integrate seamlessly with Alphalens analysis via `07_alphalens_analysis.ipynb` for standard factor performance metrics.

## Frequently Asked Questions

### What is a conditional autoencoder in asset pricing?

A conditional autoencoder is a neural network that learns compressed representations of asset returns while explicitly accounting for macroeconomic or firm-specific covariates. According to the repository's implementation, the encoder processes raw returns while the decoder reconstructs returns conditioned on macro variables, producing latent factors that capture time-varying, conditional risk premia.

### How do autoencoders differ from PCA for factor extraction?

While PCA finds linear combinations of returns that maximize variance, autoencoders discover **non-linear** manifolds through deep neural network transformations. The conditional variant further allows the factor loadings to vary with macroeconomic states, whereas PCA produces static weights. As implemented in `01_deep_autoencoders.ipynb`, the dense layers can capture complex interactions that linear decomposition methods miss.

### Why concatenate macro variables to the decoder rather than just the encoder?

Concatenating macro variables to the decoder (as shown in `06_conditional_autoencoder_for_asset_pricing_model.ipynb`) forces the latent representation to be **orthogonal** to known economic conditions. If macro variables only fed the encoder, the latent space could simply pass through macro information. By requiring the decoder to reconstruct returns using both latent codes *and* macro variables, the encoder must compress only the residual, idiosyncratic variation that represents true conditional risk.

### How do you evaluate whether the learned factors are actually predictive?

The repository uses Alphalens analysis in `07_alphalens_analysis.ipynb` to compute standard quantitative metrics including information coefficient (IC), factor return spreads between top and bottom deciles, and turnover rates. You can also project the latent factors onto standard factor models via cross-sectional OLS regressions to test for incremental explanatory power beyond traditional factors like momentum and value.