How to Implement Pairs Trading with Kalman Filter: A Python Guide

You can build a robust statistical arbitrage strategy by applying a Kalman filter to smooth price observations and calculate a dynamic, time-varying hedge ratio that updates with every new bar, then trade the z-score of the resulting mean-reverting spread.

Pairs trading with Kalman filter techniques provides an adaptive alternative to traditional static regression methods. This implementation leverages the stefan-jansen/machine-learning-for-trading repository to demonstrate how probabilistic state-space models can improve hedge ratio stability and signal quality in cointegrated relationships.

Why Use a Kalman Filter for Pairs Trading?

Traditional pairs trading relies on ordinary least squares (OLS) to compute a fixed hedge ratio over a historical window. This approach suffers when the price relationship between two assets drifts due to regime changes or market stress. The Kalman filter addresses this by treating the hedge ratio as a latent state variable that evolves gradually over time.

According to the source code in 09_time_series_models/06_statistical_arbitrage_with_cointegrated_pairs.ipynb, the implementation provides two key advantages:

  • Noise reduction: The KFSmoother function applies a Kalman smoother to each price series, separating observation noise from the underlying latent price level before any regression occurs.
  • Adaptive hedging: The KFHedgeRatio function performs Kalman regression where the state vector [beta, intercept] updates at each observation, allowing the hedge ratio to track slow structural changes in the pair's relationship.

Step-by-Step Implementation Architecture

The complete pipeline follows a clear sequence from raw data to trading signals, implemented across two primary notebooks in the repository.

Data Ingestion from HDF5 Storage

The workflow begins by reading adjusted close prices from the HDF5 store located at ../data/assets.h5. The repository uses pd.IndexSlice to efficiently query time-indexed price matrices for candidate pairs.

import pandas as pd
from pathlib import Path

DATA_STORE = Path('..', 'data', 'assets.h5')
prices = pd.read_hdf(DATA_STORE, 'stooq/us/nyse/stocks/prices')['close']

Smoothing Price Series with KFSmoother

Before estimating the hedge ratio, each price series undergoes Kalman smoothing to reduce observation noise. As implemented in lines 1573-1585 of the Statistical Arbitrage notebook, the KFSmoother function treats the price as a latent state with unit transition dynamics.

from pykalman import KalmanFilter
import numpy as np

def KFSmoother(prices):
    """Apply Kalman smoothing to reduce noise in price observations."""
    kf = KalmanFilter(
        transition_matrices=np.eye(1),
        observation_matrices=np.eye(1),
        initial_state_mean=0,
        initial_state_covariance=1,
        observation_covariance=1,
        transition_covariance=0.05,
    )
    state_means, _ = kf.filter(prices.values)
    return pd.Series(state_means.ravel(), index=prices.index)

The transition covariance (0.05) controls how much the filter trusts new observations versus the prior state, balancing responsiveness against smoothness.

Dynamic Hedge Ratio Estimation with KFHedgeRatio

The core innovation occurs in the KFHedgeRatio function (lines 1635-1660), which runs a Kalman regression on the two smoothed series. The state vector contains the slope (hedge ratio) and intercept, both updating dynamically.

def KFHedgeRatio(x, y):
    """Estimate time-varying hedge ratio using Kalman regression."""
    delta = 1e-3
    trans_cov = delta / (1 - delta) * np.eye(2)
    
    # Observation matrix [x_t, 1] for each time step

    obs_mat = np.expand_dims(np.vstack([x, np.ones_like(x)]).T, axis=1)
    
    kf = KalmanFilter(
        n_dim_obs=1,
        n_dim_state=2,
        initial_state_mean=[0, 0],
        initial_state_covariance=np.ones((2, 2)),
        transition_matrices=np.eye(2),
        observation_matrices=obs_mat,
        observation_covariance=2,
        transition_covariance=trans_cov,
    )
    state_means, _ = kf.filter(y)
    return -state_means  # Returns [slope, intercept] at each time t

The delta parameter (1e-3) determines the adaptivity of the hedge ratio. A smaller value makes the hedge ratio more responsive to recent changes, while larger values produce smoother estimates.

Spread Construction and Z-Score Calculation

With the dynamic hedge ratio beta_t and intercept alpha_t, the spread is constructed as spread = y - beta * x - alpha. The repository then calculates the half-life of mean reversion via an Ornstein-Uhlenbeck fit to determine the rolling window for z-score normalization.

def estimate_half_life(spread):
    """Calculate half-life of mean reversion using OLS."""
    delta_spread = spread.diff().dropna()
    lagged_spread = spread.shift(1).dropna()
    delta_spread = delta_spread[1:]  # Align indices

    
    # AR(1) coefficient

    lambda_ = np.polyfit(lagged_spread.values, delta_spread.values, 1)[0]
    half_life = -np.log(2) / lambda_
    return max(1, int(np.round(half_life)))

# Build spread and z-score

beta = hedge_ratios[:, 0]
intercept = hedge_ratios[:, 1]
spread = prices[y] - beta * prices[x] - intercept

half_life = estimate_half_life(spread)
window = min(2 * half_life, len(spread))

rolling_mean = spread.rolling(window).mean()
rolling_std = spread.rolling(window).std()
z_score = (spread - rolling_mean) / rolling_std

The half-life determines the optimal look-back period for the rolling statistics. A shorter half-life implies faster mean reversion and allows for tighter exit thresholds.

Signal Generation and Backtesting

The get_spread routine (lines 1697-1727) orchestrates the entire pipeline, returning a DataFrame ready for back-testing. Entry signals trigger when the z-score exceeds threshold bounds (typically ±2.0), indicating deviation from the statistical equilibrium.

entry_threshold = 2.0
exit_threshold = 0.5

# Generate trading signals

long_entries = z_score[z_score < -entry_threshold].index
short_entries = z_score[z_score > entry_threshold].index
exits = z_score[abs(z_score) < exit_threshold].index

Positions are closed when the z-score reverts to the exit threshold (e.g., ±0.5). The dynamic hedge ratio at entry determines the position sizing between the two legs, maintaining dollar neutrality according to the evolving relationship.

Complete Working Example

Below is a condensed, runnable implementation combining all components from 09_time_series_models/06_statistical_arbitrage_with_cointegrated_pairs.ipynb:

import warnings
warnings.filterwarnings('ignore')
import numpy as np
import pandas as pd
from pykalman import KalmanFilter
from pathlib import Path

def kf_smoother(series):
    kf = KalmanFilter(
        transition_matrices=np.eye(1),
        observation_matrices=np.eye(1),
        initial_state_mean=0,
        initial_state_covariance=1,
        observation_covariance=1,
        transition_covariance=0.05,
    )
    state_means, _ = kf.filter(series.values)
    return pd.Series(state_means.ravel(), index=series.index)

def kalman_hedge(x, y):
    delta = 1e-3
    trans_cov = delta / (1 - delta) * np.eye(2)
    obs_mat = np.expand_dims(np.vstack([x, np.ones_like(x)]).T, axis=1)
    
    kf = KalmanFilter(
        n_dim_obs=1,
        n_dim_state=2,
        initial_state_mean=[0, 0],
        initial_state_covariance=np.ones((2, 2)),
        transition_matrices=np.eye(2),
        observation_matrices=obs_mat,
        observation_covariance=2,
        transition_covariance=trans_cov,
    )
    state_means, _ = kf.filter(y)
    return -state_means  # [beta, intercept]

def build_pair(prices, ticker_x, ticker_y):
    # Smooth both series

    x_smooth = kf_smoother(prices[ticker_x])
    y_smooth = kf_smoother(prices[ticker_y])
    
    # Dynamic hedge ratio

    hedge = kalman_hedge(x_smooth.values, y_smooth.values)
    beta = hedge[:, 0]
    intercept = hedge[:, 1]
    
    # Calculate spread

    spread = prices[ticker_y] - beta * prices[ticker_x] - intercept
    
    # Half-life estimation (OU process)

    delta = spread.diff().dropna()
    lag = spread.shift(1).dropna()
    lambda_ = np.polyfit(lag[1:], delta[1:], 1)[0]
    hl = max(1, int(np.round(-np.log(2) / lambda_)))
    
    # Z-score with 2*half_life window

    window = min(2 * hl, len(spread))
    z = (spread - spread.rolling(window).mean()) / spread.rolling(window).std()
    
    return pd.DataFrame({
        'beta': beta,
        'spread': spread,
        'z_score': z,
        'half_life': hl
    }), hl

# Load data (assuming assets.h5 exists from repo setup)

STORE = Path('..', 'data', 'assets.h5')
prices = pd.read_hdf(STORE, 'stooq/us/nyse/stocks/prices')['close']
prices = prices.loc['2018':'2019']

pair_data, hl = build_pair(prices, 'AAPL.US', 'MSFT.US')
print(f'Half-life: {hl} days')
print(pair_data.dropna().tail())

Summary

Implementing pairs trading with Kalman filter techniques in the stefan-jansen/machine-learning-for-trading repository provides a production-ready framework for statistical arbitrage:

  • KFSmoother removes observation noise from price series before regression analysis
  • KFHedgeRatio calculates a time-varying hedge ratio that adapts to structural breaks and regime changes
  • The half-life of the spread determines optimal rolling windows for z-score calculation and position sizing
  • Signal generation relies on z-score thresholds that trigger when the filtered spread deviates significantly from its dynamic equilibrium

The complete source code resides in 04_alpha_factor_research/03_kalman_filter_and_wavelets.ipynb (introductory concepts) and 09_time_series_models/06_statistical_arbitrage_with_cointegrated_pairs.ipynb (full pipeline implementation).

Frequently Asked Questions

Why choose a Kalman filter over static OLS regression for pairs trading?

A static OLS hedge ratio assumes the relationship between two assets remains constant over the estimation window, making it vulnerable to structural breaks. The Kalman filter treats the hedge ratio as a latent state that updates with each observation, allowing it to track slow drift in the cointegrating relationship. As shown in the repository's KFHedgeRatio implementation, this adaptivity reduces false signals during regime changes while maintaining statistical efficiency.

How does the half-life parameter influence the trading strategy?

The half-life represents the expected time for the spread to revert halfway to its mean, derived from an Ornstein-Uhlenbeck process fit. In the repository's implementation, twice the half-life sets the rolling window for z-score calculation. A short half-life (1-5 days) permits tighter stop-losses and frequent trading, while a long half-life (20+ days) requires wider thresholds and longer holding periods to account for slower mean reversion.

Which Python library does the repository use for Kalman filter calculations?

The code relies on pykalman (version 0.9.5), a pure Python implementation of the Kalman filter and smoother. This library provides the KalmanFilter class used in both KFSmoother and KFHedgeRatio, with explicit parameters for transition_covariance and observation_covariance that control the filter's trust in the model versus data.

Can this implementation handle intraday data or multiple pairs simultaneously?

Yes. The get_spread function in the repository accepts a list of candidate pairs and applies the Kalman pipeline to each independently. For intraday data, you would adjust the delta parameter in KFHedgeRatio to allow faster adaptation to high-frequency noise, and ensure your HDF5 store contains timestamped data at the desired frequency. The resulting DataFrames can be concatenated for multi-pair portfolio back-testing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →