How to Implement Autoencoders for Asset Pricing: A Complete Guide to Conditional Risk Factors

Autoencoders extract conditional risk factors by learning compressed latent representations of asset returns while conditioning on macroeconomic variables, enabling you to discover hidden data-driven factors that traditional linear models miss.

This guide walks you through implementing autoencoders for asset pricing using the machine-learning-for-trading repository by Stefan Jansen. You will learn how to build conditional autoencoders that combine cross-sectional equity returns with macro features to generate latent risk factors suitable for portfolio construction and cross-sectional return prediction.

Step 1: Prepare the Return Panel and Macro Features

Before building the neural network, construct a clean panel of asset returns aligned with macroeconomic covariates. In 20_autoencoders_for_conditional_risk_factors/04_build_us_stock_dataset.ipynb, the pipeline aggregates daily returns for a broad equity universe and merges them with macro features such as inflation, interest rates, and volatility indices.

The data structure follows a 3-dimensional tensor convention:

  • Axis 0: Time (dates)
  • Axis 1: Assets (cross-section)
  • Axis 2: Features (returns + conditional variables)
import pandas as pd
import numpy as np

# Load prepared datasets from the repository

returns = pd.read_pickle('data/returns_us_stock.pkl')    # Shape: (dates, assets)

macro = pd.read_pickle('data/macro_features.pkl')        # Shape: (dates, macro_features)

# Align and split chronologically (no shuffle to preserve time structure)

from sklearn.model_selection import train_test_split
X_ret, X_cond = returns.values, macro.values
X_ret_train, X_ret_val, X_cond_train, X_cond_val = train_test_split(
    X_ret, X_cond, test_size=0.2, shuffle=False)

Step 2: Build the Base Encoder-Decoder Architecture

The foundational model in 20_autoencoders_for_conditional_risk_factors/01_deep_autoencoders.ipynb implements a dense autoencoder that compresses the high-dimensional return vector into a low-dimensional latent space. The encoder flattens the input and passes it through two hidden layers before producing the latent code z.

import tensorflow as tf

latent_dim = 5
n_assets = X_ret.shape[1]

# Encoder: (batch, n_assets) -> (batch, latent_dim)

encoder_input = tf.keras.Input(shape=(n_assets,), name='returns')
x = tf.keras.layers.Dense(128, activation='relu')(encoder_input)
x = tf.keras.layers.Dense(64, activation='relu')(x)
z = tf.keras.layers.Dense(latent_dim, name='latent')(x)
encoder = tf.keras.Model(encoder_input, z, name='encoder')

The decoder mirrors this architecture in reverse, reconstructing the original return vector from the latent representation:

decoder_input = tf.keras.Input(shape=(latent_dim,), name='decoder_input')
x = tf.keras.layers.Dense(64, activation='relu')(decoder_input)
x = tf.keras.layers.Dense(128, activation='relu')(x)
recon = tf.keras.layers.Dense(n_assets, activation='linear', name='reconstruction')(x)
decoder = tf.keras.Model(decoder_input, recon, name='decoder')

Step 3: Implement the Conditional Autoencoder

The critical innovation for asset pricing appears in 20_autoencoders_for_conditional_risk_factors/06_conditional_autoencoder_for_asset_pricing_model.ipynb. Here, the conditional autoencoder feeds macroeconomic variables to both the encoder and decoder, forcing the latent representation to capture information orthogonal to known economic conditions.

The architecture concatenates asset returns with broadcasted macro features:


# Expand dimensions for broadcasting: (dates, 1, n_assets) and (dates, macro_dim, 1)

X_ret_expanded = returns.values[..., None]           # Shape: (batch, n_assets, 1)

X_cond_expanded = macro.values[:, None, ...]         # Broadcast macro across assets

# Concatenate along feature dimension

X = np.concatenate([X_ret_expanded, X_cond_expanded], axis=-1)

Build the conditional model by injecting macro variables into the decoder input:

cond_dim = X_cond.shape[1]

# Encoder input: raw returns only

returns_input = tf.keras.Input(shape=(n_assets, 1), name='returns_input')
x = tf.keras.layers.Flatten()(returns_input)
x = tf.keras.layers.Dense(128, activation='relu')(x)
latent = tf.keras.layers.Dense(latent_dim, name='latent_vector')(x)

# Decoder input: latent code + conditional variables

cond_input = tf.keras.Input(shape=(cond_dim,), name='macro_input')
decoder_concat = tf.keras.layers.Concatenate()([latent, cond_input])
x = tf.keras.layers.Dense(128, activation='relu')(decoder_concat)
reconstruction = tf.keras.layers.Dense(n_assets, activation='linear')(x)

# Full conditional autoencoder

conditional_ae = tf.keras.Model(
    inputs=[returns_input, cond_input], 
    outputs=reconstruction,
    name='conditional_autoencoder'
)
conditional_ae.compile(optimizer='adam', loss='mse')

By reconstructing returns conditioned on macro states, the latent factors z represent conditional risk factors that vary with the economic environment.

Step 4: Train with Regularization and Early Stopping

Train the model using mean squared error (MSE) loss with early stopping to prevent overfitting on noisy return data. The repository demonstrates this pattern across 01_deep_autoencoders.ipynb through 06_conditional_autoencoder_for_asset_pricing_model.ipynb.


# Early stopping configuration

early_stop = tf.keras.callbacks.EarlyStopping(
    monitor='val_loss',
    patience=10,
    restore_best_weights=True,
    verbose=1
)

# Training with validation holdout

history = conditional_ae.fit(
    [X_ret_train[..., None], X_cond_train],
    X_ret_train,
    epochs=200,
    batch_size=256,
    validation_data=([X_ret_val[..., None], X_cond_val], X_ret_val),
    callbacks=[early_stop],
    verbose=1
)

For enhanced generalization, consider the advanced variants in 20_autoencoders_for_conditional_risk_factors/02_convolutional_denoising_autoencoders.ipynb (which adds Conv1D layers and input noise) and 03_variational_autoencoder.ipynb (which imposes a KL-divergence penalty for stochastic latent spaces).

Step 5: Extract Factors and Evaluate with Alphalens

Once trained, extract the latent factors by passing the full dataset through the encoder. These factors serve as inputs to standard asset-pricing evaluations such as cross-sectional regressions or portfolio sorts.


# Extract conditional risk factors

latent_factors = encoder.predict(X_ret[..., None])
factor_df = pd.DataFrame(
    latent_factors,
    index=returns.index,
    columns=[f'latent_{i}' for i in range(latent_dim)]
)

# Save for Alphalens analysis

factor_df.to_pickle('data/conditional_latent_factors.pkl')

Evaluate the predictive power of these learned factors using 20_autoencoders_for_conditional_risk_factors/07_alphalens_analysis.ipynb. This notebook maps the latent series into the Alphalens framework to compute:

  • Factor returns (quantile-wise performance)
  • Information coefficient (IC) (rank correlation with forward returns)
  • Turnover analysis (stability of factor exposures)

The latent factors integrate directly into existing pipelines—treat them exactly like traditional Fama-French factors or momentum signals when running OLS regressions or constructing mean-variance efficient portfolios.

Summary

  • Conditional autoencoders combine asset returns with macroeconomic variables to extract risk factors that adapt to changing economic regimes.
  • The implementation in stefan-jansen/machine-learning-for-trading provides a complete pipeline from raw price data to factor evaluation via seven interconnected notebooks.
  • Key architectural choice: Feed macro variables to the decoder (or concatenate with latent codes) to enforce that latent representations capture idiosyncratic risk orthogonal to known conditions.
  • Training best practices: Use early stopping, batch normalization, and consider denoising or variational variants to improve out-of-sample generalization.
  • Evaluation: Extracted latent factors from 06_conditional_autoencoder_for_asset_pricing_model.ipynb integrate seamlessly with Alphalens analysis via 07_alphalens_analysis.ipynb for standard factor performance metrics.

Frequently Asked Questions

What is a conditional autoencoder in asset pricing?

A conditional autoencoder is a neural network that learns compressed representations of asset returns while explicitly accounting for macroeconomic or firm-specific covariates. According to the repository's implementation, the encoder processes raw returns while the decoder reconstructs returns conditioned on macro variables, producing latent factors that capture time-varying, conditional risk premia.

How do autoencoders differ from PCA for factor extraction?

While PCA finds linear combinations of returns that maximize variance, autoencoders discover non-linear manifolds through deep neural network transformations. The conditional variant further allows the factor loadings to vary with macroeconomic states, whereas PCA produces static weights. As implemented in 01_deep_autoencoders.ipynb, the dense layers can capture complex interactions that linear decomposition methods miss.

Why concatenate macro variables to the decoder rather than just the encoder?

Concatenating macro variables to the decoder (as shown in 06_conditional_autoencoder_for_asset_pricing_model.ipynb) forces the latent representation to be orthogonal to known economic conditions. If macro variables only fed the encoder, the latent space could simply pass through macro information. By requiring the decoder to reconstruct returns using both latent codes and macro variables, the encoder must compress only the residual, idiosyncratic variation that represents true conditional risk.

How do you evaluate whether the learned factors are actually predictive?

The repository uses Alphalens analysis in 07_alphalens_analysis.ipynb to compute standard quantitative metrics including information coefficient (IC), factor return spreads between top and bottom deciles, and turnover rates. You can also project the latent factors onto standard factor models via cross-sectional OLS regressions to test for incremental explanatory power beyond traditional factors like momentum and value.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →