How to Use RNNs for Multivariate Time Series Prediction in Trading

Recurrent Neural Networks (RNNs) with LSTM or GRU layers capture non-linear dependencies across multiple financial time series by reshaping data into (samples, window_size, n_features) tensors and training sequence-to-vector models that outperform traditional VAR baselines.

Multivariate time series prediction is essential for portfolio management and algorithmic trading, where asset prices move in complex interdependent patterns. The machine-learning-for-trading repository by Stefan Jansen provides a complete implementation using TensorFlow/Keras to forecast multiple economic indicators simultaneously. This guide explains how to prepare FRED macroeconomic data and architect RNN models that exploit temporal dependencies superior to vector autoregression.

Why RNNs Outperform VAR for Trading Data

Vector Autoregression (VAR) models assume linear relationships and require stationarity, often failing to capture regime shifts in financial markets. In contrast, RNNs—specifically LSTM and GRU variants—maintain internal state vectors that remember long-term dependencies across variables.

As implemented in 19_recurrent_neural_nets/04_multivariate_timeseries.ipynb, the architecture accepts input tensors of shape (n_samples, window_size, n_series), where window_size represents the lookback period and n_series the number of macroeconomic indicators (e.g., industrial production, unemployment, consumer sentiment). The hidden layers use tanh activations with recurrent dropout for regularization, while the output Dense layer projects to the forecast horizon. The model compiles with the RMSprop optimizer and Mean Squared Error loss, converging faster than VAR maximum likelihood estimation on non-stationary FRED data loaded from 00_build_dataset.ipynb.

Preparing Multivariate Sequences

Financial time series require careful windowing to preserve temporal causality. The custom helper function create_multivariate_rnn_data transforms a 2D DataFrame into supervised learning samples by generating rolling windows.

import numpy as np
import pandas as pd
from sklearn.preprocessing import MinMaxScaler

def create_multivariate_rnn_data(data, window_size):
    """
    Reshape time series into (samples, timesteps, features) format.
    Source: 19_recurrent_neural_nets/04_multivariate_timeseries.ipynb#L45
    """
    y = data[window_size:]
    n_features = data.shape[1]
    X = np.zeros((len(y), window_size, n_features))
    
    for i in range(len(y)):
        X[i] = data[i:i+window_size]
    
    return X, y

# Load FRED monthly data from helper notebook

df = pd.read_csv('fred_macro_data.csv', index_col=0, parse_dates=True)
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(df)

# Create sequences with 24-month lookback for business cycle capture

X, y = create_multivariate_rnn_data(data_scaled, window_size=24)
print(f"Input shape: {X.shape}")  # (n_samples, 24, n_series)

Architecting the LSTM/GRU Model

The repository demonstrates a stacked LSTM configuration suitable for capturing hierarchical temporal patterns. The first LSTM layer returns sequences to feed subsequent recurrent layers, while the final LSTM feeds a Dense output predicting the next time step for all series simultaneously.

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import LSTM, Dense, Dropout

# Model definition as per 19_recurrent_neural_nets/04_multivariate_timeseries.ipynb#L88

model = Sequential([
    LSTM(50, activation='tanh', return_sequences=True, 
         input_shape=(X.shape[1], X.shape[2])),
    Dropout(0.2),
    LSTM(50, activation='tanh'),
    Dense(y.shape[1])  # Predicts all series simultaneously

])

model.compile(optimizer='rmsprop', loss='mse', metrics=['mae'])
model.summary()

Training and Rolling Forecast Generation

Training uses early stopping to prevent overfitting on the training window. For trading applications, the repository implements a rolling forecast where the model recursively predicts the next period, then appends the prediction to the input window for multi-step ahead forecasting.

from tensorflow.keras.callbacks import EarlyStopping

# Training configuration from 19_recurrent_neural_nets/04_multivariate_timeseries.ipynb#L120

early_stop = EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True)

history = model.fit(X_train, y_train, 
                    epochs=100, 
                    batch_size=32, 
                    validation_split=0.2,
                    callbacks=[early_stop],
                    verbose=1)

# Rolling forecast for trading signals (multi-step ahead)

predictions = []
current_batch = X_test[0:1].copy()  # Start with first test window

for i in range(len(y_test)):
    pred = model.predict(current_batch, verbose=0)
    predictions.append(pred[0])
    # Update window: drop oldest timestep, append new prediction

    current_batch = np.roll(current_batch, -1, axis=1)
    current_batch[0, -1, :] = pred

predictions = scaler.inverse_transform(predictions)

Summary

  • RNNs handle non-linear multivariate dependencies better than VAR models for financial forecasting, particularly during market regime shifts.
  • The create_multivariate_rnn_data function reshapes FRED data into (samples, window_size, features) tensors required by Keras LSTM layers.
  • LSTM or GRU architectures with recurrent dropout prevent overfitting on noisy trading data while maintaining temporal memory.
  • RMSprop optimizer is preferred over Adam for RNN stability on stationary-transformed economic series from the FRED database.
  • Rolling forecasts enable realistic multi-step prediction for portfolio rebalancing strategies without future data leakage.

Frequently Asked Questions

Can I use GRU instead of LSTM for trading prediction?

Yes. GRU layers are computationally efficient and perform comparably on smaller macroeconomic datasets with fewer parameters. In 04_multivariate_timeseries.ipynb, replacing LSTM with GRU reduces training time by approximately 25% while maintaining similar MSE on the FRED test set.

How do I select the optimal window size for the RNN?

Window size selection depends on the autocorrelation structure of your trading series. For monthly FRED data, the repository uses 24-month windows to capture business cycle effects. Test windows between 12 and 36 periods, validating with the BIC or cross-validated MSE to prevent look-ahead bias in your backtests.

Do I need to normalize data before using create_multivariate_rnn_data?

Absolutely. Multivariate financial series operate on different scales (e.g., unemployment percentages vs. industrial production indices). Apply MinMaxScaler or StandardScaler before windowing, as shown in the data preparation section, to ensure stable gradients during backpropagation through time.

How does this compare to transformer models for time series?

While transformers capture long-range dependencies via attention mechanisms, LSTM/GRU models are more data-efficient for the limited history available in macroeconomic datasets like FRED. The repository demonstrates that RNNs outperform VAR baselines with significantly fewer parameters and training data than transformer architectures require.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →