How to Implement Pairs Trading with Cointegration: A Complete Python Workflow
You can implement pairs trading with cointegration by testing asset pairs for statistical stationarity using Engle-Granger and Johansen tests, selecting those that pass both thresholds, and executing mean-reversion trades via a backtrader strategy that monitors the spread.
The machine-learning-for-trading repository by Stefan Jansen provides a production-ready framework for statistical arbitrage. It demonstrates how to implement pairs trading with cointegration through a modular pipeline that moves from data acquisition to backtested execution using statsmodels for econometric testing and backtrader for event-driven simulation.
Step 1: Data Preparation and Alignment
Before testing for cointegration, you must align price series on a common calendar. The repository uses the notebook data/create_stooq_data.ipynb to download ETF and equity data, clean tickers, and resample to a consistent frequency.
Proper alignment ensures that subsequent cointegration tests compare contemporaneous prices. The cleaned data feeds directly into the screening process without leakage between training and test periods.
Step 2: Testing for Cointegration
The core logic resides in 09_time_series_models/06_statistical_arbitrage_with_cointegrated_pairs.ipynb. The test_cointegration function (lines 86-114) loops through all possible pairs, applies two distinct statistical tests, and persists results to an HDF5 store for efficient retrieval.
Engle-Granger Two-Step Method
The Engle-Granger test runs a two-directional Augmented Dickey-Fuller (ADF) regression to check if a linear combination of the two price series is stationary.
from statsmodels.tsa.stattools import coint
# Bidirectional test ensures robustness regardless of which series is regressed on which
p1 = coint(df[s1], df[s2], trend='c')[1]
p2 = coint(df[s2], df[s1], trend='c')[1]
A p-value below 0.05 in both directions indicates that the spread is mean-reverting.
Johansen Multivariate Test
While Engle-Granger tests single cointegrating relationships, the Johansen test handles multiple vectors and provides trace statistics and eigenvectors for portfolio construction.
from statsmodels.tsa.vector_ar.vecm import coint_johansen
# det_order=0 for no deterministic trend, k_ar_diff from VAR lag selection
cj = coint_johansen(df[[s1, s2]], det_order=0, k_ar_diff=order)
trace_stat = cj.lr1[0]
crit_value = cj.cvt[0, 1] # 5% critical value
eigenvector = cj.evec[:, cj.ind[0]] # Hedge ratio
Vector Autoregression Lag Selection
Before running the Johansen test, the code fits a VAR model to select the optimal lag order via AIC. This prevents over-differencing and aligns the multivariate test with the underlying dynamics of the pair.
from statsmodels.tsa.vector_ar.var_model import VAR
var = VAR(df[[s1, s2]].dropna())
order = var.select_order().selected_orders['aic']
Step 3: Selecting Robust Trading Pairs
After computing statistics for the entire universe, the notebook filters candidates using a strict dual-threshold mask around line 935.
# Significance at 5% level for both directions of Engle-Granger
test_results['eg_sig'] = (
(test_results.coint_eg1 < 0.05) &
(test_results.coint_eg2 < 0.05)
)
# Johansen trace statistic exceeds 5% critical value
test_results['joh_sig'] = test_results.j_trace0 > test_results.j_crit0
# Union of both conditions to minimize false positives
test_results['coint'] = test_results.eg_sig & test_results.joh_sig
selected_pairs = test_results[test_results.coint]
Only pairs passing both tests proceed to the trading stage. This dual requirement reduces spurious regressions that might arise from random walk behavior.
Step 4: Building the Trading Strategy with Backtrader
The execution engine lives in 09_time_series_models/07_pairs_trading_backtest.ipynb. It defines a StatisticalArbitrageCointegration class inheriting from bt.Strategy, which manages position sizing, spread monitoring, and risk limits.
The Pair Dataclass
The strategy uses a type-safe Pair dataclass to encapsulate hedge ratios, position sizes, and active status.
from dataclasses import dataclass
@dataclass
class Pair:
s1: str
s2: str
hedge: float # From Johansen eigenvector
size1: int
size2: int
long: bool = True
active: bool = False
def compute_spread_return(self, prices_s1, prices_s2):
return prices_s1 - self.hedge * prices_s2
Entry and Exit Logic
The strategy opens positions when the normalized spread deviates beyond a Bollinger-Band-type threshold (typically 2 standard deviations) and closes them when the spread reverts to the mean or hits a risk limit.
import backtrader as bt
from pathlib import Path
import csv
class StatisticalArbitrageCointegration(bt.Strategy):
params = (('risk_limit', -0.2), ('log_file', 'backtest.csv'))
def __init__(self):
self.active_pairs = {}
self.clear_log()
def clear_log(self):
Path(self.p.log_file).unlink(missing_ok=True)
with open(self.p.log_file, 'a', newline='') as f:
csv.writer(f).writerow(['Date', 'Pair', 'Symbol', 'Order #',
'Reason', 'Status', 'Long', 'Price',
'Size', 'Position'])
def next(self):
# Iterate active pairs, check spread deviation
for pair_id, pair in self.active_pairs.items():
spread = pair.compute_spread_return(
self.data.close[pair.s1],
self.data.close[pair.s2]
)
# Entry/exit logic and risk checks omitted for brevity
if not pair.active and self.is_entry_signal(spread):
self.buy_pair(pair)
elif pair.active and self.check_risk_limit(spread):
self.sell_pair(pair)
The check_risk_limit method enforces a hard stop at -20% (risk_limit=-0.2) to prevent catastrophic drift during structural breaks.
Step 5: Backtesting and Performance Evaluation
The notebook wires the strategy to backtrader.Cerebro, feeds the pre-processed price series, and outputs a CSV log containing timestamps, fills, and P&L.
cerebro = bt.Cerebro()
cerebro.addstrategy(StatisticalArbitrageCointegration,
risk_limit=-0.2,
log_file='backtest.csv')
# Load aligned price data
for ticker, df in price_data.items():
data = bt.feeds.PandasData(
dataname=df,
name=ticker,
datetime='date',
open='open',
high='high',
low='low',
close='close',
volume='volume'
)
cerebro.adddata(data)
cerebro.broker.setcash(100_000)
cerebro.run()
cerebro.plot()
Post-backtest analysis reads backtest.csv to calculate Sharpe ratios, cumulative returns, and maximum drawdown using pandas.
Summary
- Dual Testing Requirement: The repository requires both Engle-Granger and Johansen tests to pass at the 5% significance level, filtering out spurious pairs.
- VAR Lag Selection: Optimal lag order selection via AIC ensures the Johansen test respects the time-series dynamics of each pair.
- Backtrader Integration: The
StatisticalArbitrageCointegrationclass leverages event-driven architecture for realistic order handling and slippage simulation. - Persistent Storage: Cointegration results are stored in HDF5 format, separating expensive computation from trading logic for rapid iteration.
- Risk Management: A hardcoded -20% risk limit on individual pair positions prevents runaway losses during regime changes.
Frequently Asked Questions
What is the difference between Engle-Granger and Johansen cointegration tests?
The Engle-Granger test is a two-step single-equation method that regresses one series on another and tests the residuals for stationarity. It is implemented via statsmodels.tsa.stattools.coint. The Johansen test is a multivariate maximum likelihood approach that can detect multiple cointegrating vectors and provides eigenvectors for hedge ratio construction via statsmodels.tsa.vector_ar.vecm.coint_johansen. The repository uses both to ensure robust pair selection.
How does the strategy determine position sizes for each leg of the pair?
Position sizes are derived from the hedge ratio extracted from the Johansen eigenvector (cj.evec). This ratio determines how many shares of the second asset to trade for each share of the first asset, creating a dollar-neutral spread. The Pair dataclass stores these as size1 and size2, ensuring the portfolio remains market-neutral at entry.
Why does the code use a Vector Autoregression (VAR) before the Johansen test?
The VAR model selects the optimal lag order using the AIC criterion (var.select_order().selected_orders['aic']). The Johansen test requires the number of lagged differences (k_ar_diff) as an input. Using AIC ensures the test specification matches the data-generating process, preventing size distortions in the cointegration test statistics.
What file contains the complete backtesting implementation?
The complete backtesting logic resides in 09_time_series_models/07_pairs_trading_backtest.ipynb. This notebook contains the StatisticalArbitrageCointegration strategy class, the Pair dataclass definition, and the cerebro.run() execution block that generates performance reports.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →