# How to Implement Pairs Trading with Kalman Filter: A Python Guide

> Implement pairs trading with Kalman filter using Python. Build a robust statistical arbitrage strategy using a dynamic hedge ratio and trade the z-score of the mean-reverting spread for profitable results.

- Repository: [Stefan Jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading)
- Tags: how-to-guide
- Published: 2026-06-02

---

**You can build a robust statistical arbitrage strategy by applying a Kalman filter to smooth price observations and calculate a dynamic, time-varying hedge ratio that updates with every new bar, then trade the z-score of the resulting mean-reverting spread.**

Pairs trading with Kalman filter techniques provides an adaptive alternative to traditional static regression methods. This implementation leverages the **stefan-jansen/machine-learning-for-trading** repository to demonstrate how probabilistic state-space models can improve hedge ratio stability and signal quality in cointegrated relationships.

## Why Use a Kalman Filter for Pairs Trading?

Traditional pairs trading relies on ordinary least squares (OLS) to compute a fixed hedge ratio over a historical window. This approach suffers when the price relationship between two assets drifts due to regime changes or market stress. The **Kalman filter** addresses this by treating the hedge ratio as a latent state variable that evolves gradually over time.

According to the source code in `09_time_series_models/06_statistical_arbitrage_with_cointegrated_pairs.ipynb`, the implementation provides two key advantages:

- **Noise reduction**: The `KFSmoother` function applies a Kalman smoother to each price series, separating observation noise from the underlying latent price level before any regression occurs.
- **Adaptive hedging**: The `KFHedgeRatio` function performs Kalman regression where the state vector `[beta, intercept]` updates at each observation, allowing the hedge ratio to track slow structural changes in the pair's relationship.

## Step-by-Step Implementation Architecture

The complete pipeline follows a clear sequence from raw data to trading signals, implemented across two primary notebooks in the repository.

### Data Ingestion from HDF5 Storage

The workflow begins by reading adjusted close prices from the HDF5 store located at `../data/assets.h5`. The repository uses `pd.IndexSlice` to efficiently query time-indexed price matrices for candidate pairs.

```python
import pandas as pd
from pathlib import Path

DATA_STORE = Path('..', 'data', 'assets.h5')
prices = pd.read_hdf(DATA_STORE, 'stooq/us/nyse/stocks/prices')['close']

```

### Smoothing Price Series with KFSmoother

Before estimating the hedge ratio, each price series undergoes Kalman smoothing to reduce observation noise. As implemented in lines 1573-1585 of the *Statistical Arbitrage* notebook, the `KFSmoother` function treats the price as a latent state with unit transition dynamics.

```python
from pykalman import KalmanFilter
import numpy as np

def KFSmoother(prices):
    """Apply Kalman smoothing to reduce noise in price observations."""
    kf = KalmanFilter(
        transition_matrices=np.eye(1),
        observation_matrices=np.eye(1),
        initial_state_mean=0,
        initial_state_covariance=1,
        observation_covariance=1,
        transition_covariance=0.05,
    )
    state_means, _ = kf.filter(prices.values)
    return pd.Series(state_means.ravel(), index=prices.index)

```

The **transition covariance** (0.05) controls how much the filter trusts new observations versus the prior state, balancing responsiveness against smoothness.

### Dynamic Hedge Ratio Estimation with KFHedgeRatio

The core innovation occurs in the `KFHedgeRatio` function (lines 1635-1660), which runs a Kalman regression on the two smoothed series. The state vector contains the **slope** (hedge ratio) and **intercept**, both updating dynamically.

```python
def KFHedgeRatio(x, y):
    """Estimate time-varying hedge ratio using Kalman regression."""
    delta = 1e-3
    trans_cov = delta / (1 - delta) * np.eye(2)
    
    # Observation matrix [x_t, 1] for each time step

    obs_mat = np.expand_dims(np.vstack([x, np.ones_like(x)]).T, axis=1)
    
    kf = KalmanFilter(
        n_dim_obs=1,
        n_dim_state=2,
        initial_state_mean=[0, 0],
        initial_state_covariance=np.ones((2, 2)),
        transition_matrices=np.eye(2),
        observation_matrices=obs_mat,
        observation_covariance=2,
        transition_covariance=trans_cov,
    )
    state_means, _ = kf.filter(y)
    return -state_means  # Returns [slope, intercept] at each time t

```

The **delta** parameter (1e-3) determines the adaptivity of the hedge ratio. A smaller value makes the hedge ratio more responsive to recent changes, while larger values produce smoother estimates.

### Spread Construction and Z-Score Calculation

With the dynamic hedge ratio `beta_t` and intercept `alpha_t`, the spread is constructed as `spread = y - beta * x - alpha`. The repository then calculates the **half-life** of mean reversion via an Ornstein-Uhlenbeck fit to determine the rolling window for z-score normalization.

```python
def estimate_half_life(spread):
    """Calculate half-life of mean reversion using OLS."""
    delta_spread = spread.diff().dropna()
    lagged_spread = spread.shift(1).dropna()
    delta_spread = delta_spread[1:]  # Align indices

    
    # AR(1) coefficient

    lambda_ = np.polyfit(lagged_spread.values, delta_spread.values, 1)[0]
    half_life = -np.log(2) / lambda_
    return max(1, int(np.round(half_life)))

# Build spread and z-score

beta = hedge_ratios[:, 0]
intercept = hedge_ratios[:, 1]
spread = prices[y] - beta * prices[x] - intercept

half_life = estimate_half_life(spread)
window = min(2 * half_life, len(spread))

rolling_mean = spread.rolling(window).mean()
rolling_std = spread.rolling(window).std()
z_score = (spread - rolling_mean) / rolling_std

```

The **half-life** determines the optimal look-back period for the rolling statistics. A shorter half-life implies faster mean reversion and allows for tighter exit thresholds.

### Signal Generation and Backtesting

The `get_spread` routine (lines 1697-1727) orchestrates the entire pipeline, returning a DataFrame ready for back-testing. Entry signals trigger when the **z-score** exceeds threshold bounds (typically ±2.0), indicating deviation from the statistical equilibrium.

```python
entry_threshold = 2.0
exit_threshold = 0.5

# Generate trading signals

long_entries = z_score[z_score < -entry_threshold].index
short_entries = z_score[z_score > entry_threshold].index
exits = z_score[abs(z_score) < exit_threshold].index

```

Positions are closed when the z-score reverts to the **exit threshold** (e.g., ±0.5). The dynamic hedge ratio at entry determines the position sizing between the two legs, maintaining dollar neutrality according to the evolving relationship.

## Complete Working Example

Below is a condensed, runnable implementation combining all components from `09_time_series_models/06_statistical_arbitrage_with_cointegrated_pairs.ipynb`:

```python
import warnings
warnings.filterwarnings('ignore')
import numpy as np
import pandas as pd
from pykalman import KalmanFilter
from pathlib import Path

def kf_smoother(series):
    kf = KalmanFilter(
        transition_matrices=np.eye(1),
        observation_matrices=np.eye(1),
        initial_state_mean=0,
        initial_state_covariance=1,
        observation_covariance=1,
        transition_covariance=0.05,
    )
    state_means, _ = kf.filter(series.values)
    return pd.Series(state_means.ravel(), index=series.index)

def kalman_hedge(x, y):
    delta = 1e-3
    trans_cov = delta / (1 - delta) * np.eye(2)
    obs_mat = np.expand_dims(np.vstack([x, np.ones_like(x)]).T, axis=1)
    
    kf = KalmanFilter(
        n_dim_obs=1,
        n_dim_state=2,
        initial_state_mean=[0, 0],
        initial_state_covariance=np.ones((2, 2)),
        transition_matrices=np.eye(2),
        observation_matrices=obs_mat,
        observation_covariance=2,
        transition_covariance=trans_cov,
    )
    state_means, _ = kf.filter(y)
    return -state_means  # [beta, intercept]

def build_pair(prices, ticker_x, ticker_y):
    # Smooth both series

    x_smooth = kf_smoother(prices[ticker_x])
    y_smooth = kf_smoother(prices[ticker_y])
    
    # Dynamic hedge ratio

    hedge = kalman_hedge(x_smooth.values, y_smooth.values)
    beta = hedge[:, 0]
    intercept = hedge[:, 1]
    
    # Calculate spread

    spread = prices[ticker_y] - beta * prices[ticker_x] - intercept
    
    # Half-life estimation (OU process)

    delta = spread.diff().dropna()
    lag = spread.shift(1).dropna()
    lambda_ = np.polyfit(lag[1:], delta[1:], 1)[0]
    hl = max(1, int(np.round(-np.log(2) / lambda_)))
    
    # Z-score with 2*half_life window

    window = min(2 * hl, len(spread))
    z = (spread - spread.rolling(window).mean()) / spread.rolling(window).std()
    
    return pd.DataFrame({
        'beta': beta,
        'spread': spread,
        'z_score': z,
        'half_life': hl
    }), hl

# Load data (assuming assets.h5 exists from repo setup)

STORE = Path('..', 'data', 'assets.h5')
prices = pd.read_hdf(STORE, 'stooq/us/nyse/stocks/prices')['close']
prices = prices.loc['2018':'2019']

pair_data, hl = build_pair(prices, 'AAPL.US', 'MSFT.US')
print(f'Half-life: {hl} days')
print(pair_data.dropna().tail())

```

## Summary

Implementing pairs trading with Kalman filter techniques in the **stefan-jansen/machine-learning-for-trading** repository provides a production-ready framework for statistical arbitrage:

- **`KFSmoother`** removes observation noise from price series before regression analysis
- **`KFHedgeRatio`** calculates a **time-varying hedge ratio** that adapts to structural breaks and regime changes
- The **half-life** of the spread determines optimal rolling windows for z-score calculation and position sizing
- Signal generation relies on z-score thresholds that trigger when the filtered spread deviates significantly from its dynamic equilibrium

The complete source code resides in `04_alpha_factor_research/03_kalman_filter_and_wavelets.ipynb` (introductory concepts) and `09_time_series_models/06_statistical_arbitrage_with_cointegrated_pairs.ipynb` (full pipeline implementation).

## Frequently Asked Questions

### Why choose a Kalman filter over static OLS regression for pairs trading?

A **static OLS hedge ratio** assumes the relationship between two assets remains constant over the estimation window, making it vulnerable to structural breaks. The **Kalman filter** treats the hedge ratio as a latent state that updates with each observation, allowing it to track slow drift in the cointegrating relationship. As shown in the repository's `KFHedgeRatio` implementation, this adaptivity reduces false signals during regime changes while maintaining statistical efficiency.

### How does the half-life parameter influence the trading strategy?

The **half-life** represents the expected time for the spread to revert halfway to its mean, derived from an Ornstein-Uhlenbeck process fit. In the repository's implementation, twice the half-life sets the rolling window for z-score calculation. A short half-life (1-5 days) permits tighter stop-losses and frequent trading, while a long half-life (20+ days) requires wider thresholds and longer holding periods to account for slower mean reversion.

### Which Python library does the repository use for Kalman filter calculations?

The code relies on **pykalman** (version 0.9.5), a pure Python implementation of the Kalman filter and smoother. This library provides the `KalmanFilter` class used in both `KFSmoother` and `KFHedgeRatio`, with explicit parameters for `transition_covariance` and `observation_covariance` that control the filter's trust in the model versus data.

### Can this implementation handle intraday data or multiple pairs simultaneously?

Yes. The `get_spread` function in the repository accepts a list of candidate pairs and applies the Kalman pipeline to each independently. For intraday data, you would adjust the **delta** parameter in `KFHedgeRatio` to allow faster adaptation to high-frequency noise, and ensure your HDF5 store contains timestamped data at the desired frequency. The resulting DataFrames can be concatenated for multi-pair portfolio back-testing.