# Handling Survivorship Bias in Backtesting Datasets: Methods and Implementation

> Eliminate survivorship bias in backtesting datasets to prevent inflated performance. Learn methods to filter raw price panels and ensure accurate trading strategy results.

- Repository: [Papers With Backtest/awesome-systematic-trading](https://github.com/paperswithbacktest/awesome-systematic-trading)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Survivorship bias inflates backtesting performance by excluding delisted securities, and can be eliminated by filtering raw price panels to include only tickers that existed throughout the entire test horizon before running any strategy.**

Handling survivorship bias in backtesting datasets is critical for realistic performance metrics in quantitative trading research. The **awesome-systematic-trading** repository provides a strategy-agnostic framework that separates data cleaning from strategy implementation. By using pure functions that consume pre-processed DataFrames, the codebase ensures that delisted and merged securities are properly excluded before any factor calculations begin.

## Understanding Survivorship Bias in Financial Data

Survivorship bias occurs when a backtesting sample includes only assets that survived until the end of the observation window, systematically omitting firms that were delisted, merged, or went bankrupt. This selection error artificially inflates strategy returns because failed companies—which often exhibit poor performance—are excluded from the analysis. According to the awesome-systematic-trading source code, mitigating this bias requires explicit data handling before any strategy logic executes.

## Architectural Approach to Data Integrity

The repository addresses survivorship bias through three core architectural principles:

### Comprehensive Data Ingestion

The data loaders referenced in [`README.md`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/README.md) emphasize sourcing raw price and fundamental data from providers that retain delisted securities. **CRSP** and **Compustat** databases are specifically recommended because they maintain historical records for securities that no longer trade, ensuring the complete universe is available for filtering.

### Explicit Data-Cleaning Utilities

Helper scripts within `static/utils/` demonstrate how to flag and remove entries that only appear after a certain date. These utilities—including conceptual filters like `filter_delisted` and `align_dates`—align the backtest horizon with the actual existence period of each ticker, preventing lookahead bias and survivorship bias simultaneously.

### Strategy-Agnostic Design

Every strategy under `static/strategies/` is implemented as a pure function receiving a cleaned `DataFrame`. This decoupling ensures that bias removal occurs once during preprocessing, making the cleaning logic reusable across all factor-based studies in the repository.

## Step-by-Step: Removing Survivorship Bias Before Backtesting

The following implementation demonstrates how to clean raw CRSP data to eliminate survivorship bias before passing it to any strategy from the repository.

```python
import pandas as pd
import numpy as np

# ----------------------------------------------------------------------

# 1️⃣ Load raw price data (example CSV exported from CRSP)

# ----------------------------------------------------------------------

raw_prices = pd.read_csv(
    "data/crsp_daily_prices.csv",
    parse_dates=["date"],
    dtype={"ticker": str}
)

# ----------------------------------------------------------------------

# 2️⃣ Identify delisted securities

# ----------------------------------------------------------------------

# A ticker is considered delisted if its last appearance is before the

# end-of-sample date.

END_DATE = pd.Timestamp("2023-12-31")
last_trade = raw_prices.groupby("ticker")["date"].max()
delisted = last_trade[last_trade < END_DATE].index

# ----------------------------------------------------------------------

# 3️⃣ Remove survivorship-biased rows

# ----------------------------------------------------------------------

clean_prices = raw_prices[~raw_prices["ticker"].isin(delisted)]

# ----------------------------------------------------------------------

# 4️⃣ Align the panel to a common start date

# ----------------------------------------------------------------------

START_DATE = pd.Timestamp("2000-01-01")
clean_prices = clean_prices[
    (clean_prices["date"] >= START_DATE) &
    (clean_prices["date"] <= END_DATE)
]

# ----------------------------------------------------------------------

# 5️⃣ Feed the cleaned DataFrame into a strategy from the repo

# ----------------------------------------------------------------------

# Example: momentum_factor_effect_in_stocks.py expects a DataFrame with

# columns ['date', 'ticker', 'close'].

from static.strategies.momentum_factor_effect_in_stocks import run_momentum_strategy

results = run_momentum_strategy(clean_prices)
print(results.head())

```

### Key Implementation Details

- **Step 2** isolates tickers whose last trading day precedes the desired backtest horizon—the classic survivorship bias filter identified in the source analysis.
- **Step 3** drops all rows for those tickers, preserving only the universe that truly existed throughout the test period.
- The cleaned `clean_prices` DataFrame can be passed directly to any strategy (e.g., `run_momentum_strategy`) because each implementation in `static/strategies/` expects pre-processed data with standardized columns.

## Repository Structure for Bias-Free Backtesting

The following files demonstrate how the awesome-systematic-trading repository organizes data integrity and strategy implementation:

| File | Role |
|------|------|
| [`README.md`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/README.md) | Central documentation including data-source recommendations and bias-handling guidelines |
| [`static/strategies/momentum_factor_effect_in_stocks.py`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/static/strategies/momentum_factor_effect_in_stocks.py) | Example strategy consuming cleaned price panels via `run_momentum_strategy` |
| [`static/strategies/value-factor-effect-within-countries.py`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/static/strategies/value-factor-effect-within-countries.py) | Demonstrates factor-based research requiring delisted-security filtering |
| [`static/strategies/turn-of-the-month-in-equity-indexes.py`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/static/strategies/turn-of-the-month-in-equity-indexes.py) | Calendar-based effect strategy relying on clean input data |
| [`static/strategies/volatility-risk-premium-effect.py`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/static/strategies/volatility-risk-premium-effect.py) | Volatility strategy particularly sensitive to survivorship bias omission |

## Why This Matters for Strategy Performance

Strategies such as momentum and value factors are particularly susceptible to survivorship bias because distressed companies often exhibit extreme returns before delisting. By implementing the filtering logic demonstrated above—specifically checking `last_trade < END_DATE` before executing any portfolio formation—the awesome-systematic-trading framework ensures that backtests reflect realistic execution costs and availability constraints.

## Summary

- **Survivorship bias** artificially inflates returns by excluding delisted securities from historical datasets.
- The **awesome-systematic-trading** repository mitigates this by ingesting data from sources like **CRSP** and **Compustat** that retain delisted records.
- **Explicit filtering** (grouping by ticker and checking last trade dates against the test horizon) must occur before strategy execution.
- All strategies in `static/strategies/` are designed as pure functions that accept pre-cleaned DataFrames, separating data integrity logic from factor calculation.
- The `run_momentum_strategy` function in [`momentum_factor_effect_in_stocks.py`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/momentum_factor_effect_in_stocks.py) exemplifies how cleaned data flows directly into backtesting logic.

## Frequently Asked Questions

### What is survivorship bias in backtesting datasets?

Survivorship bias occurs when historical datasets only include securities that currently exist or survived through the entire backtest period, excluding companies that were delisted, merged, or bankrupted. This omission creates an upward bias in performance metrics because poorly performing assets that failed are systematically removed from the sample.

### How do I identify delisted securities in CRSP data?

Group your price DataFrame by ticker using `groupby("ticker")["date"].max()` to find the last trading date for each security. Compare these dates against your backtest end date; any ticker whose last appearance precedes the end date was delisted during the sample period and must be excluded to avoid survivorship bias.

### Why does the awesome-systematic-trading repository use pure functions for strategies?

The repository implements strategies as pure functions receiving pre-processed DataFrames to enforce a strict separation between data cleaning and strategy logic. This design ensures that survivorship bias removal and other preprocessing steps occur consistently across all strategies in `static/strategies/`, preventing implementation errors that could reintroduce bias into specific factor calculations.

### Which strategies in the repository are most affected by survivorship bias?

Momentum and volatility strategies are particularly vulnerable because delisted securities often exhibit extreme price movements before failure. The [`volatility-risk-premium-effect.py`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/volatility-risk-premium-effect.py) and [`momentum_factor_effect_in_stocks.py`](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/momentum_factor_effect_in_stocks.py) strategies require complete historical panels including distressed assets to accurately measure factor premiums, making the delisted-filtering step essential for valid results.