Handling Survivorship Bias in Backtesting Datasets: Methods and Implementation
Survivorship bias inflates backtesting performance by excluding delisted securities, and can be eliminated by filtering raw price panels to include only tickers that existed throughout the entire test horizon before running any strategy.
Handling survivorship bias in backtesting datasets is critical for realistic performance metrics in quantitative trading research. The awesome-systematic-trading repository provides a strategy-agnostic framework that separates data cleaning from strategy implementation. By using pure functions that consume pre-processed DataFrames, the codebase ensures that delisted and merged securities are properly excluded before any factor calculations begin.
Understanding Survivorship Bias in Financial Data
Survivorship bias occurs when a backtesting sample includes only assets that survived until the end of the observation window, systematically omitting firms that were delisted, merged, or went bankrupt. This selection error artificially inflates strategy returns because failed companies—which often exhibit poor performance—are excluded from the analysis. According to the awesome-systematic-trading source code, mitigating this bias requires explicit data handling before any strategy logic executes.
Architectural Approach to Data Integrity
The repository addresses survivorship bias through three core architectural principles:
Comprehensive Data Ingestion
The data loaders referenced in README.md emphasize sourcing raw price and fundamental data from providers that retain delisted securities. CRSP and Compustat databases are specifically recommended because they maintain historical records for securities that no longer trade, ensuring the complete universe is available for filtering.
Explicit Data-Cleaning Utilities
Helper scripts within static/utils/ demonstrate how to flag and remove entries that only appear after a certain date. These utilities—including conceptual filters like filter_delisted and align_dates—align the backtest horizon with the actual existence period of each ticker, preventing lookahead bias and survivorship bias simultaneously.
Strategy-Agnostic Design
Every strategy under static/strategies/ is implemented as a pure function receiving a cleaned DataFrame. This decoupling ensures that bias removal occurs once during preprocessing, making the cleaning logic reusable across all factor-based studies in the repository.
Step-by-Step: Removing Survivorship Bias Before Backtesting
The following implementation demonstrates how to clean raw CRSP data to eliminate survivorship bias before passing it to any strategy from the repository.
import pandas as pd
import numpy as np
# ----------------------------------------------------------------------
# 1️⃣ Load raw price data (example CSV exported from CRSP)
# ----------------------------------------------------------------------
raw_prices = pd.read_csv(
"data/crsp_daily_prices.csv",
parse_dates=["date"],
dtype={"ticker": str}
)
# ----------------------------------------------------------------------
# 2️⃣ Identify delisted securities
# ----------------------------------------------------------------------
# A ticker is considered delisted if its last appearance is before the
# end-of-sample date.
END_DATE = pd.Timestamp("2023-12-31")
last_trade = raw_prices.groupby("ticker")["date"].max()
delisted = last_trade[last_trade < END_DATE].index
# ----------------------------------------------------------------------
# 3️⃣ Remove survivorship-biased rows
# ----------------------------------------------------------------------
clean_prices = raw_prices[~raw_prices["ticker"].isin(delisted)]
# ----------------------------------------------------------------------
# 4️⃣ Align the panel to a common start date
# ----------------------------------------------------------------------
START_DATE = pd.Timestamp("2000-01-01")
clean_prices = clean_prices[
(clean_prices["date"] >= START_DATE) &
(clean_prices["date"] <= END_DATE)
]
# ----------------------------------------------------------------------
# 5️⃣ Feed the cleaned DataFrame into a strategy from the repo
# ----------------------------------------------------------------------
# Example: momentum_factor_effect_in_stocks.py expects a DataFrame with
# columns ['date', 'ticker', 'close'].
from static.strategies.momentum_factor_effect_in_stocks import run_momentum_strategy
results = run_momentum_strategy(clean_prices)
print(results.head())
Key Implementation Details
- Step 2 isolates tickers whose last trading day precedes the desired backtest horizon—the classic survivorship bias filter identified in the source analysis.
- Step 3 drops all rows for those tickers, preserving only the universe that truly existed throughout the test period.
- The cleaned
clean_pricesDataFrame can be passed directly to any strategy (e.g.,run_momentum_strategy) because each implementation instatic/strategies/expects pre-processed data with standardized columns.
Repository Structure for Bias-Free Backtesting
The following files demonstrate how the awesome-systematic-trading repository organizes data integrity and strategy implementation:
| File | Role |
|---|---|
README.md |
Central documentation including data-source recommendations and bias-handling guidelines |
static/strategies/momentum_factor_effect_in_stocks.py |
Example strategy consuming cleaned price panels via run_momentum_strategy |
static/strategies/value-factor-effect-within-countries.py |
Demonstrates factor-based research requiring delisted-security filtering |
static/strategies/turn-of-the-month-in-equity-indexes.py |
Calendar-based effect strategy relying on clean input data |
static/strategies/volatility-risk-premium-effect.py |
Volatility strategy particularly sensitive to survivorship bias omission |
Why This Matters for Strategy Performance
Strategies such as momentum and value factors are particularly susceptible to survivorship bias because distressed companies often exhibit extreme returns before delisting. By implementing the filtering logic demonstrated above—specifically checking last_trade < END_DATE before executing any portfolio formation—the awesome-systematic-trading framework ensures that backtests reflect realistic execution costs and availability constraints.
Summary
- Survivorship bias artificially inflates returns by excluding delisted securities from historical datasets.
- The awesome-systematic-trading repository mitigates this by ingesting data from sources like CRSP and Compustat that retain delisted records.
- Explicit filtering (grouping by ticker and checking last trade dates against the test horizon) must occur before strategy execution.
- All strategies in
static/strategies/are designed as pure functions that accept pre-cleaned DataFrames, separating data integrity logic from factor calculation. - The
run_momentum_strategyfunction inmomentum_factor_effect_in_stocks.pyexemplifies how cleaned data flows directly into backtesting logic.
Frequently Asked Questions
What is survivorship bias in backtesting datasets?
Survivorship bias occurs when historical datasets only include securities that currently exist or survived through the entire backtest period, excluding companies that were delisted, merged, or bankrupted. This omission creates an upward bias in performance metrics because poorly performing assets that failed are systematically removed from the sample.
How do I identify delisted securities in CRSP data?
Group your price DataFrame by ticker using groupby("ticker")["date"].max() to find the last trading date for each security. Compare these dates against your backtest end date; any ticker whose last appearance precedes the end date was delisted during the sample period and must be excluded to avoid survivorship bias.
Why does the awesome-systematic-trading repository use pure functions for strategies?
The repository implements strategies as pure functions receiving pre-processed DataFrames to enforce a strict separation between data cleaning and strategy logic. This design ensures that survivorship bias removal and other preprocessing steps occur consistently across all strategies in static/strategies/, preventing implementation errors that could reintroduce bias into specific factor calculations.
Which strategies in the repository are most affected by survivorship bias?
Momentum and volatility strategies are particularly vulnerable because delisted securities often exhibit extreme price movements before failure. The volatility-risk-premium-effect.py and momentum_factor_effect_in_stocks.py strategies require complete historical panels including distressed assets to accurately measure factor premiums, making the delisted-filtering step essential for valid results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →