How to Get Free Historical Stock Market Data for Backtesting: A Complete Guide

You can obtain free historical stock market data for backtesting by using open-source Python libraries such as yfinance, AkShare, and pandas-datareader to download price series, then cleaning and feeding that data into backtesting engines like Backtrader, Zipline, or vectorbt.

The Awesome Systematic Trading repository maintains a curated list of free data-access libraries that provide programmatic interfaces to historical market data without requiring paid subscriptions. These tools return tidy pandas DataFrames containing OHLCV (Open, High, Low, Close, Volume) data, enabling you to build reproducible quantitative strategies entirely within the open-source ecosystem.

Choosing a Free Data Provider

Selecting the right library depends on which markets you need to access. The repository's [README.md](https://github.com/paperswithbacktest/awesome-systematic-trading/blob/main/README.md#data-sources) lists several actively maintained options under the Data Sources section:

  • yfinance: Ideal for U.S. equities and ETFs, offering adjusted daily prices and dividends.
  • AkShare: Comprehensive coverage of Chinese A-shares plus global macro data.
  • TuShare: Specialized for Chinese market data with fundamental datasets.
  • pandas-datareader: Unified interface for Yahoo Finance, Alpha Vantage, and other sources.

Different providers expose different exchanges and data fields, so verify that your target assets are supported before committing to a pipeline.

Installing and Fetching Raw Data

Once you select a provider, install the package via pip to ensure you have the latest API surface and compiled binaries.

pip install yfinance akshare pandas-datareader

Downloading U.S. Equity Data with yfinance

The yfinance library provides a simple interface to Yahoo Finance data. The download() function returns a DataFrame with OHLCV columns ready for analysis:

import yfinance as yf

# Pull Apple data from 2000 to 2024

data = yf.download("AAPL", start="2000-01-01", end="2024-12-31")
print(data.head())

Accessing Chinese A-Share Data with AkShare

For Chinese markets, AkShare offers granular access to individual securities. The stock_zh_a_daily() function retrieves historical prices for specific symbols:

import akshare as ak

# Ping An Insurance (Shanghai)

cn_data = ak.stock_zh_a_daily(
    symbol="sh600036", 
    start_date="20000101", 
    end_date="20241231"
)
cn_data = cn_data.set_index("date")["close"]

Cleaning and Aligning Time Series

Raw data often contains missing trading days or requires adjustment for splits and dividends. Clean data avoids look-ahead bias and ensures a consistent time index for vectorized backtests.

Fill missing calendar days using forward-fill to maintain continuous series:

import pandas as pd

# Reindex to complete date range and forward fill

date_range = pd.date_range(start="2000-01-01", end="2024-12-31", freq="D")
data = data.reindex(date_range, method="ffill")

Most libraries like yfinance return split and dividend-adjusted prices by default, but verify the Adj Close column exists to ensure your returns reflect corporate actions accurately.

Caching Data Locally with Marketstore

Re-downloading large datasets for every experiment wastes bandwidth and time. The repository lists Marketstore under its Databases section as a high-performance time-series database optimized for financial data.

Start Marketstore using Docker:

docker run -d -p 5993:5993 -v $(pwd)/data:/data \
  marketstore/marketstore

Then cache your DataFrames to avoid redundant network calls:

import marketstore as ms

# Write to Marketstore (requires datetime index)

ms.write("AAPL", data)

# Read back for backtesting

df_cached = ms.read("AAPL")

Storing data in Feather or Parquet format (data.to_parquet("AAPL.parquet")) provides a lightweight alternative for smaller datasets.

Feeding Data to Backtesting Frameworks

Each backtesting engine expects a specific data wrapper, but conversion is trivial once you have a clean DataFrame.

Backtrader Integration

Backtrader accepts pandas DataFrames through the PandasData feed:

import backtrader as bt

datafeed = bt.feeds.PandasData(dataname=data)
cerebro = bt.Cerebro()
cerebro.adddata(datafeed)

vectorbt Integration

For vectorized backtesting, pass the price series directly to Portfolio.from_signals():

import vectorbt as vbt

close = data["Adj Close"]
sma_fast = close.rolling(20).mean()
sma_slow = close.rolling(60).mean()

entries = sma_fast > sma_slow
exits = sma_fast < sma_slow

portfolio = vbt.Portfolio.from_signals(close, entries, exits, freq="1D")
print(portfolio.stats())

Zipline Integration

For Zipline, register your DataFrame via the data portal and ensure compatibility with the trading calendar system. The static/strategies/ directory in the repository contains example QuantConnect strategies that demonstrate data ingestion patterns adaptable to Zipline.

Enriching with Fundamental Data

Price data alone limits factor-based strategies. Use AkShare to pull balance-sheet items, earnings dates, or macro variables and join them on your existing datetime index:


# Fetch fundamental data

fundamentals = ak.stock_financial_report_sina(stock="600036")

# Join with price data on date index

enriched_data = data.join(fundamentals, how="left")

This enables construction of value and quality factors without purchasing commercial fundamental datasets.

Summary

  • Choose market-specific libraries: yfinance for U.S. equities, AkShare or TuShare for Chinese markets, and pandas-datareader for multi-source aggregation.
  • Clean data meticulously: Handle missing days via reindexing and verify adjusted close prices to avoid split/dividend distortions.
  • Cache aggressively: Use Marketstore or Parquet files to persist downloaded data and accelerate iterative backtesting.
  • Framework flexibility: Convert DataFrames to Backtrader feeds, vectorbt signals, or Zipline data portals with minimal boilerplate.
  • Enrich strategically: Combine price series with fundamental data from AkShare to build multi-factor models.

Frequently Asked Questions

What is the best free library for downloading U.S. stock market data?

yfinance is the most popular choice for U.S. equities, providing adjusted daily OHLCV data through a simple Python interface. According to the Awesome Systematic Trading repository's README.md, it is actively maintained and supports dividends and stock splits automatically.

Can I get free historical data for international markets like China?

Yes, AkShare and TuShare provide comprehensive coverage of Chinese A-shares and other Asian markets. These libraries offer both price history and fundamental data, making them suitable for region-specific systematic strategies without data vendor costs.

How do I prevent look-ahead bias when preparing free data for backtesting?

Ensure you use only data available at each historical timestamp by avoiding future-peeking during cleaning. Fill missing values using forward-fill (method="ffill") rather than backward-fill, and confirm that adjusted prices are calculated based on corporate actions announced on or before each date.

What is the most efficient way to store historical data for repeated backtests?

Use Marketstore for high-performance storage or Apache Parquet for file-based caching. Marketstore, listed under the repository's Databases section, allows millisecond-level queries across large universes, while Parquet provides compression and fast I/O for smaller datasets stored locally.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →