How to Backtest ML-Driven Trading Strategies Using Backtrader
Backtrader enables event-driven backtesting of machine learning trading signals by extending PandasData to include prediction columns and implementing a custom Strategy class that dynamically allocates capital based on ML model outputs.
The notebook 08_ml4t_workflow/03_backtesting_with_backtrader.ipynb in the stefan-jansen/machine-learning-for-trading repository demonstrates a production-ready workflow for integrating scikit-learn predictions with the Backtrader engine. This approach supports realistic simulation features including custom commission schemes, multi-asset portfolios, and PyFolio-compatible performance analytics.
Setting Up the Backtrader Environment
Begin by importing the required libraries and defining a custom commission structure. The repository silences warnings and configures pandas display options to handle large equity curves cleanly.
import warnings
warnings.filterwarnings('ignore')
import backtrader as bt
from backtrader.feeds import PandasData
import pandas as pd
Custom Commission Schemes
For realistic cost modeling, the repository implements a fixed-fee commission scheme rather than percentage-based fees. This is critical when testing high-frequency ML signals where flat fees dominate transaction costs.
class FixedCommisionScheme(bt.CommInfoBase):
params = (('commission', .02), ('stocklike', True),
('commtype', bt.CommInfoBase.COMM_FIXED),)
def _getcommission(self, size, price, pseudoexec):
return abs(size) * self.p.commission
Attach this scheme to the broker using cerebro.broker.addcommissioninfo(FixedCommisionScheme()) after initializing the Cerebro engine.
Creating a Custom Data Feed for ML Signals
Standard Backtrader data feeds only recognize OHLCV columns. To inject ML predictions into the backtest, you must extend PandasData to declare a new predicted line that stores model outputs (positive values for long signals, negative for short).
class SignalData(PandasData):
cols = ['open', 'high', 'low', 'close', 'volume', 'predicted']
lines = tuple(cols)
params = {c: -1 for c in cols}
params.update({'datetime': None})
params = tuple(params.items())
Key implementation details:
- The
linestuple creates accessible attributes (.open,.high, ...,.predicted) inside the strategy - Parameter values of
-1instruct Backtrader to auto-locate columns by name in the input DataFrame - The input DataFrame must contain a
predictedcolumn with numerical signals derived from your ML model
Implementing the ML Strategy Logic
The MLStrategy class implements a long/short equity market-neutral approach. It ranks all universes by predicted returns each day and rebalances to hold equal-weight positions in the top-N longs and bottom-N shorts.
class MLStrategy(bt.Strategy):
params = (('n_positions', 10), ('min_positions', 5), ('verbose', False))
def next(self):
today = self.datas[0].datetime.date()
up, down = {}, {}
# Collect predictions for current date
for data in self.datas:
if data.datetime.date() == today:
if data.predicted[0] > 0:
up[data._name] = data.predicted[0]
elif data.predicted[0] < 0:
down[data._name] = data.predicted[0]
# Rank and select top/bottom signals
shorts = sorted(down, key=down.get)[:self.p.n_positions]
longs = sorted(up, key=up.get, reverse=True)[:self.p.n_positions]
# Risk management: require minimum positions to avoid concentration
if len(shorts) < self.p.min_positions or len(longs) < self.p.min_positions:
longs, shorts = [], []
# Close positions not in current signals
for ticker in self.getdatanames():
if ticker not in longs + shorts:
self.order_target_percent(data=ticker, target=0)
# Equal-weight allocation
short_target = -1 / max(self.p.n_positions, len(shorts))
long_target = 1 / max(self.p.n_positions, len(longs))
for ticker in shorts:
self.order_target_percent(data=ticker, target=short_target)
for ticker in longs:
self.order_target_percent(data=ticker, target=long_target)
Critical logic components:
- Signal aggregation: The strategy iterates through
self.datas(all registered tickers) and checksdata.predicted[0]to access the current bar's ML prediction - Ranking mechanism: Python's
sorted()function creates priority queues based on signal strength - Position sizing:
order_target_percentenforces equal-weight exposure, automatically handling existing positions and rounding - Safety constraints: The
min_positionsparameter prevents the strategy from taking concentrated bets when few signals are available
Configuring Cerebro and Running the Backtest
Load pre-computed predictions (typically stored in HDF5 format from an earlier ML training step) and attach each ticker as a separate data feed. The example uses pd.read_hdf to retrieve OHLCV plus predictions from 00_data/backtest.h5.
cerebro = bt.Cerebro()
cerebro.broker.setcash(10000.0)
# Load backtest data with predictions
data = pd.read_hdf('00_data/backtest.h5', 'data').sort_index()
# Add each ticker individually
for ticker in data.index.get_level_values(0).unique():
df = data.loc[pd.IndexSlice[ticker, :], :].droplevel('ticker', axis=0)
df.index.name = 'datetime'
cerebro.adddata(SignalData(dataname=df), name=ticker)
# Attach strategy and analyzer
cerebro.addstrategy(MLStrategy, n_positions=25, min_positions=20)
cerebro.addanalyzer(bt.analyzers.PyFolio, _name='pyfolio')
results = cerebro.run()
ending_value = cerebro.broker.getvalue()
print(f'Final Portfolio Value: {ending_value:,.2f}')
Analyzing Results with PyFolio
Backtrader's built-in PyFolio analyzer extracts four objects required for institutional-grade performance reporting: returns, positions, transactions, and gross leverage. Export these to HDF5 for persistent storage or pass directly to PyFolio's tear sheet generator.
pyfolio_analyzer = results[0].analyzers.getbyname('pyfolio')
returns, positions, transactions, gross_lev = pyfolio_analyzer.get_pf_items()
# Export for downstream analysis
returns.to_hdf('backtrader.h5', 'returns')
positions.to_hdf('backtrader.h5', 'positions')
transactions.to_hdf('backtrader.h5', 'transactions')
# Generate full tear sheet with benchmark
import pandas_datareader.data as web
benchmark = web.DataReader('SP500', 'fred', '2014', '2018').pct_change().tz_localize('UTC')
import pyfolio as pf
pf.create_full_tear_sheet(returns, transactions=transactions,
positions=positions, benchmark_rets=benchmark.dropna())
This workflow produces Sharpe ratios, drawdown analysis, and beta-adjusted returns essential for validating ML-driven alpha generation.
Complete Minimal Example
For rapid prototyping, use this condensed version that requires only a single CSV containing OHLCV columns plus a predicted column:
import backtrader as bt
import pandas as pd
from backtrader.feeds import PandasData
class SignalData(PandasData):
lines = ('predicted',)
params = (('datetime', None), ('open', -1), ('high', -1),
('low', -1), ('close', -1), ('volume', -1),
('predicted', -1))
class SimpleML(bt.Strategy):
def next(self):
for data in self.datas:
if data.predicted[0] > 0:
self.order_target_percent(data=data, target=0.1)
elif data.predicted[0] < 0:
self.order_target_percent(data=data, target=-0.1)
df = pd.read_csv('predictions.csv', parse_dates=['datetime'])
df.set_index('datetime', inplace=True)
cerebro = bt.Cerebro()
cerebro.broker.setcash(100000.0)
cerebro.adddata(SignalData(dataname=df), name='asset')
cerebro.addstrategy(SimpleML)
cerebro.run()
print(f'Final Value: {cerebro.broker.getvalue():,.2f}')
Summary
- Custom data feeds: Extend
PandasDatain08_ml4t_workflow/03_backtesting_with_backtrader.ipynbto map ML prediction columns to Backtrader lines using thelinesandparamsclass attributes. - Dynamic rebalancing: The
MLStrategy.next()method ranks universe constituents bydata.predicted[0]values and rebalances daily usingorder_target_percentfor equal-weight exposures. - Risk controls: Implement minimum position thresholds and concentration limits via strategy parameters (
n_positions,min_positions). - Performance analytics: Attach
bt.analyzers.PyFolioto export returns, positions, and transactions compatible with PyFolio tear sheets for Sharpe ratio and drawdown analysis. - Commission modeling: Override
bt.CommInfoBaseto implement fixed-cost or percentage-based fee structures that accurately reflect trading costs on ML signal turnover.
Frequently Asked Questions
What input format does Backtrader require for ML predictions?
Backtrader requires a pandas DataFrame indexed by datetime containing standard OHLCV columns plus your model's predictions. In SignalData, the params dictionary maps these column names to line indices, with -1 enabling automatic column detection. The DataFrame is passed via SignalData(dataname=df) when calling cerebro.adddata().
How does the MLStrategy handle position sizing and rebalancing?
The strategy calculates equal-weight targets for long and short sides separately using 1 / max(self.p.n_positions, len(longs)). It calls self.order_target_percent() for each selected ticker, which automatically sizes orders to reach the target percentage of current portfolio value while closing stale positions not present in the current signal set.
Can I backtest multiple assets simultaneously with different ML models?
Yes. The repository demonstrates multi-asset backtesting by iterating through tickers in an HDF5 file and calling cerebro.adddata() for each. Each ticker maintains its own predicted line. The strategy accesses these via self.datas and filters by data._name to apply ticker-specific logic if using different models per asset.
How do I export backtest results for risk analysis outside Backtrader?
Add bt.analyzers.PyFolio to the Cerebro instance before running. After cerebro.run(), access the analyzer via results[0].analyzers.getbyname('pyfolio') and call .get_pf_items() to extract returns, positions, transactions, and leverage. These pandas objects can be saved to HDF5 or fed directly into PyFolio's create_full_tear_sheet() for benchmark comparison.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →