# How to Use RNNs for Multivariate Time Series Prediction in Trading

> Master multivariate time series prediction in trading using RNNs and LSTM or GRU layers. Learn to capture non-linear dependencies and improve your trading strategies with this practical guide.

- Repository: [Stefan Jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading)
- Tags: how-to-guide
- Published: 2026-06-02

---

**Recurrent Neural Networks (RNNs) with LSTM or GRU layers capture non-linear dependencies across multiple financial time series by reshaping data into `(samples, window_size, n_features)` tensors and training sequence-to-vector models that outperform traditional VAR baselines.**

Multivariate time series prediction is essential for portfolio management and algorithmic trading, where asset prices move in complex interdependent patterns. The `machine-learning-for-trading` repository by Stefan Jansen provides a complete implementation using TensorFlow/Keras to forecast multiple economic indicators simultaneously. This guide explains how to prepare FRED macroeconomic data and architect RNN models that exploit temporal dependencies superior to vector autoregression.

## Why RNNs Outperform VAR for Trading Data

Vector Autoregression (VAR) models assume linear relationships and require stationarity, often failing to capture regime shifts in financial markets. In contrast, RNNs—specifically LSTM and GRU variants—maintain internal state vectors that remember long-term dependencies across variables.

As implemented in [`19_recurrent_neural_nets/04_multivariate_timeseries.ipynb`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/19_recurrent_neural_nets/04_multivariate_timeseries.ipynb), the architecture accepts input tensors of shape `(n_samples, window_size, n_series)`, where `window_size` represents the lookback period and `n_series` the number of macroeconomic indicators (e.g., industrial production, unemployment, consumer sentiment). The hidden layers use **tanh** activations with recurrent dropout for regularization, while the output **Dense** layer projects to the forecast horizon. The model compiles with the **RMSprop** optimizer and **Mean Squared Error** loss, converging faster than VAR maximum likelihood estimation on non-stationary FRED data loaded from [`00_build_dataset.ipynb`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/19_recurrent_neural_nets/00_build_dataset.ipynb).

## Preparing Multivariate Sequences

Financial time series require careful windowing to preserve temporal causality. The custom helper function [`create_multivariate_rnn_data`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/19_recurrent_neural_nets/04_multivariate_timeseries.ipynb#L45) transforms a 2D DataFrame into supervised learning samples by generating rolling windows.

```python
import numpy as np
import pandas as pd
from sklearn.preprocessing import MinMaxScaler

def create_multivariate_rnn_data(data, window_size):
    """
    Reshape time series into (samples, timesteps, features) format.
    Source: 19_recurrent_neural_nets/04_multivariate_timeseries.ipynb#L45
    """
    y = data[window_size:]
    n_features = data.shape[1]
    X = np.zeros((len(y), window_size, n_features))
    
    for i in range(len(y)):
        X[i] = data[i:i+window_size]
    
    return X, y

# Load FRED monthly data from helper notebook

df = pd.read_csv('fred_macro_data.csv', index_col=0, parse_dates=True)
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(df)

# Create sequences with 24-month lookback for business cycle capture

X, y = create_multivariate_rnn_data(data_scaled, window_size=24)
print(f"Input shape: {X.shape}")  # (n_samples, 24, n_series)

```

## Architecting the LSTM/GRU Model

The repository demonstrates a **stacked LSTM** configuration suitable for capturing hierarchical temporal patterns. The first LSTM layer returns sequences to feed subsequent recurrent layers, while the final LSTM feeds a Dense output predicting the next time step for all series simultaneously.

```python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import LSTM, Dense, Dropout

# Model definition as per 19_recurrent_neural_nets/04_multivariate_timeseries.ipynb#L88

model = Sequential([
    LSTM(50, activation='tanh', return_sequences=True, 
         input_shape=(X.shape[1], X.shape[2])),
    Dropout(0.2),
    LSTM(50, activation='tanh'),
    Dense(y.shape[1])  # Predicts all series simultaneously

])

model.compile(optimizer='rmsprop', loss='mse', metrics=['mae'])
model.summary()

```

## Training and Rolling Forecast Generation

Training uses early stopping to prevent overfitting on the training window. For trading applications, the repository implements a rolling forecast where the model recursively predicts the next period, then appends the prediction to the input window for multi-step ahead forecasting.

```python
from tensorflow.keras.callbacks import EarlyStopping

# Training configuration from 19_recurrent_neural_nets/04_multivariate_timeseries.ipynb#L120

early_stop = EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True)

history = model.fit(X_train, y_train, 
                    epochs=100, 
                    batch_size=32, 
                    validation_split=0.2,
                    callbacks=[early_stop],
                    verbose=1)

# Rolling forecast for trading signals (multi-step ahead)

predictions = []
current_batch = X_test[0:1].copy()  # Start with first test window

for i in range(len(y_test)):
    pred = model.predict(current_batch, verbose=0)
    predictions.append(pred[0])
    # Update window: drop oldest timestep, append new prediction

    current_batch = np.roll(current_batch, -1, axis=1)
    current_batch[0, -1, :] = pred

predictions = scaler.inverse_transform(predictions)

```

## Summary

- **RNNs** handle non-linear multivariate dependencies better than VAR models for financial forecasting, particularly during market regime shifts.
- The [`create_multivariate_rnn_data`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/19_recurrent_neural_nets/04_multivariate_timeseries.ipynb) function reshapes FRED data into `(samples, window_size, features)` tensors required by Keras LSTM layers.
- **LSTM** or **GRU** architectures with recurrent dropout prevent overfitting on noisy trading data while maintaining temporal memory.
- **RMSprop** optimizer is preferred over Adam for RNN stability on stationary-transformed economic series from the FRED database.
- Rolling forecasts enable realistic multi-step prediction for portfolio rebalancing strategies without future data leakage.

## Frequently Asked Questions

### Can I use GRU instead of LSTM for trading prediction?

Yes. GRU layers are computationally efficient and perform comparably on smaller macroeconomic datasets with fewer parameters. In [`04_multivariate_timeseries.ipynb`](https://github.com/stefan-jansen/machine-learning-for-trading/blob/main/19_recurrent_neural_nets/04_multivariate_timeseries.ipynb), replacing `LSTM` with `GRU` reduces training time by approximately 25% while maintaining similar MSE on the FRED test set.

### How do I select the optimal window size for the RNN?

Window size selection depends on the autocorrelation structure of your trading series. For monthly FRED data, the repository uses 24-month windows to capture business cycle effects. Test windows between 12 and 36 periods, validating with the **BIC** or **cross-validated MSE** to prevent look-ahead bias in your backtests.

### Do I need to normalize data before using `create_multivariate_rnn_data`?

Absolutely. Multivariate financial series operate on different scales (e.g., unemployment percentages vs. industrial production indices). Apply `MinMaxScaler` or `StandardScaler` before windowing, as shown in the data preparation section, to ensure stable gradients during backpropagation through time.

### How does this compare to transformer models for time series?

While transformers capture long-range dependencies via attention mechanisms, LSTM/GRU models are more data-efficient for the limited history available in macroeconomic datasets like FRED. The repository demonstrates that RNNs outperform VAR baselines with significantly fewer parameters and training data than transformer architectures require.