How to Backtest ML-Driven Trading Strategies Using Backtrader

Backtrader enables event-driven backtesting of machine learning trading signals by extending PandasData to include prediction columns and implementing a custom Strategy class that dynamically allocates capital based on ML model outputs.

The notebook 08_ml4t_workflow/03_backtesting_with_backtrader.ipynb in the stefan-jansen/machine-learning-for-trading repository demonstrates a production-ready workflow for integrating scikit-learn predictions with the Backtrader engine. This approach supports realistic simulation features including custom commission schemes, multi-asset portfolios, and PyFolio-compatible performance analytics.

Setting Up the Backtrader Environment

Begin by importing the required libraries and defining a custom commission structure. The repository silences warnings and configures pandas display options to handle large equity curves cleanly.

import warnings
warnings.filterwarnings('ignore')
import backtrader as bt
from backtrader.feeds import PandasData
import pandas as pd

Custom Commission Schemes

For realistic cost modeling, the repository implements a fixed-fee commission scheme rather than percentage-based fees. This is critical when testing high-frequency ML signals where flat fees dominate transaction costs.

class FixedCommisionScheme(bt.CommInfoBase):
    params = (('commission', .02), ('stocklike', True),
              ('commtype', bt.CommInfoBase.COMM_FIXED),)

    def _getcommission(self, size, price, pseudoexec):
        return abs(size) * self.p.commission

Attach this scheme to the broker using cerebro.broker.addcommissioninfo(FixedCommisionScheme()) after initializing the Cerebro engine.

Creating a Custom Data Feed for ML Signals

Standard Backtrader data feeds only recognize OHLCV columns. To inject ML predictions into the backtest, you must extend PandasData to declare a new predicted line that stores model outputs (positive values for long signals, negative for short).

class SignalData(PandasData):
    cols = ['open', 'high', 'low', 'close', 'volume', 'predicted']
    lines = tuple(cols)
    params = {c: -1 for c in cols}
    params.update({'datetime': None})
    params = tuple(params.items())

Key implementation details:

  • The lines tuple creates accessible attributes (.open, .high, ..., .predicted) inside the strategy
  • Parameter values of -1 instruct Backtrader to auto-locate columns by name in the input DataFrame
  • The input DataFrame must contain a predicted column with numerical signals derived from your ML model

Implementing the ML Strategy Logic

The MLStrategy class implements a long/short equity market-neutral approach. It ranks all universes by predicted returns each day and rebalances to hold equal-weight positions in the top-N longs and bottom-N shorts.

class MLStrategy(bt.Strategy):
    params = (('n_positions', 10), ('min_positions', 5), ('verbose', False))

    def next(self):
        today = self.datas[0].datetime.date()
        up, down = {}, {}
        
        # Collect predictions for current date

        for data in self.datas:
            if data.datetime.date() == today:
                if data.predicted[0] > 0:
                    up[data._name] = data.predicted[0]
                elif data.predicted[0] < 0:
                    down[data._name] = data.predicted[0]
        
        # Rank and select top/bottom signals

        shorts = sorted(down, key=down.get)[:self.p.n_positions]
        longs = sorted(up, key=up.get, reverse=True)[:self.p.n_positions]
        
        # Risk management: require minimum positions to avoid concentration

        if len(shorts) < self.p.min_positions or len(longs) < self.p.min_positions:
            longs, shorts = [], []
        
        # Close positions not in current signals

        for ticker in self.getdatanames():
            if ticker not in longs + shorts:
                self.order_target_percent(data=ticker, target=0)
        
        # Equal-weight allocation

        short_target = -1 / max(self.p.n_positions, len(shorts))
        long_target = 1 / max(self.p.n_positions, len(longs))
        
        for ticker in shorts:
            self.order_target_percent(data=ticker, target=short_target)
        for ticker in longs:
            self.order_target_percent(data=ticker, target=long_target)

Critical logic components:

  • Signal aggregation: The strategy iterates through self.datas (all registered tickers) and checks data.predicted[0] to access the current bar's ML prediction
  • Ranking mechanism: Python's sorted() function creates priority queues based on signal strength
  • Position sizing: order_target_percent enforces equal-weight exposure, automatically handling existing positions and rounding
  • Safety constraints: The min_positions parameter prevents the strategy from taking concentrated bets when few signals are available

Configuring Cerebro and Running the Backtest

Load pre-computed predictions (typically stored in HDF5 format from an earlier ML training step) and attach each ticker as a separate data feed. The example uses pd.read_hdf to retrieve OHLCV plus predictions from 00_data/backtest.h5.

cerebro = bt.Cerebro()
cerebro.broker.setcash(10000.0)

# Load backtest data with predictions

data = pd.read_hdf('00_data/backtest.h5', 'data').sort_index()

# Add each ticker individually

for ticker in data.index.get_level_values(0).unique():
    df = data.loc[pd.IndexSlice[ticker, :], :].droplevel('ticker', axis=0)
    df.index.name = 'datetime'
    cerebro.adddata(SignalData(dataname=df), name=ticker)

# Attach strategy and analyzer

cerebro.addstrategy(MLStrategy, n_positions=25, min_positions=20)
cerebro.addanalyzer(bt.analyzers.PyFolio, _name='pyfolio')

results = cerebro.run()
ending_value = cerebro.broker.getvalue()
print(f'Final Portfolio Value: {ending_value:,.2f}')

Analyzing Results with PyFolio

Backtrader's built-in PyFolio analyzer extracts four objects required for institutional-grade performance reporting: returns, positions, transactions, and gross leverage. Export these to HDF5 for persistent storage or pass directly to PyFolio's tear sheet generator.

pyfolio_analyzer = results[0].analyzers.getbyname('pyfolio')
returns, positions, transactions, gross_lev = pyfolio_analyzer.get_pf_items()

# Export for downstream analysis

returns.to_hdf('backtrader.h5', 'returns')
positions.to_hdf('backtrader.h5', 'positions')
transactions.to_hdf('backtrader.h5', 'transactions')

# Generate full tear sheet with benchmark

import pandas_datareader.data as web
benchmark = web.DataReader('SP500', 'fred', '2014', '2018').pct_change().tz_localize('UTC')

import pyfolio as pf
pf.create_full_tear_sheet(returns, transactions=transactions,
                          positions=positions, benchmark_rets=benchmark.dropna())

This workflow produces Sharpe ratios, drawdown analysis, and beta-adjusted returns essential for validating ML-driven alpha generation.

Complete Minimal Example

For rapid prototyping, use this condensed version that requires only a single CSV containing OHLCV columns plus a predicted column:

import backtrader as bt
import pandas as pd
from backtrader.feeds import PandasData

class SignalData(PandasData):
    lines = ('predicted',)
    params = (('datetime', None), ('open', -1), ('high', -1),
              ('low', -1), ('close', -1), ('volume', -1),
              ('predicted', -1))

class SimpleML(bt.Strategy):
    def next(self):
        for data in self.datas:
            if data.predicted[0] > 0:
                self.order_target_percent(data=data, target=0.1)
            elif data.predicted[0] < 0:
                self.order_target_percent(data=data, target=-0.1)

df = pd.read_csv('predictions.csv', parse_dates=['datetime'])
df.set_index('datetime', inplace=True)

cerebro = bt.Cerebro()
cerebro.broker.setcash(100000.0)
cerebro.adddata(SignalData(dataname=df), name='asset')
cerebro.addstrategy(SimpleML)

cerebro.run()
print(f'Final Value: {cerebro.broker.getvalue():,.2f}')

Summary

  • Custom data feeds: Extend PandasData in 08_ml4t_workflow/03_backtesting_with_backtrader.ipynb to map ML prediction columns to Backtrader lines using the lines and params class attributes.
  • Dynamic rebalancing: The MLStrategy.next() method ranks universe constituents by data.predicted[0] values and rebalances daily using order_target_percent for equal-weight exposures.
  • Risk controls: Implement minimum position thresholds and concentration limits via strategy parameters (n_positions, min_positions).
  • Performance analytics: Attach bt.analyzers.PyFolio to export returns, positions, and transactions compatible with PyFolio tear sheets for Sharpe ratio and drawdown analysis.
  • Commission modeling: Override bt.CommInfoBase to implement fixed-cost or percentage-based fee structures that accurately reflect trading costs on ML signal turnover.

Frequently Asked Questions

What input format does Backtrader require for ML predictions?

Backtrader requires a pandas DataFrame indexed by datetime containing standard OHLCV columns plus your model's predictions. In SignalData, the params dictionary maps these column names to line indices, with -1 enabling automatic column detection. The DataFrame is passed via SignalData(dataname=df) when calling cerebro.adddata().

How does the MLStrategy handle position sizing and rebalancing?

The strategy calculates equal-weight targets for long and short sides separately using 1 / max(self.p.n_positions, len(longs)). It calls self.order_target_percent() for each selected ticker, which automatically sizes orders to reach the target percentage of current portfolio value while closing stale positions not present in the current signal set.

Can I backtest multiple assets simultaneously with different ML models?

Yes. The repository demonstrates multi-asset backtesting by iterating through tickers in an HDF5 file and calling cerebro.adddata() for each. Each ticker maintains its own predicted line. The strategy accesses these via self.datas and filters by data._name to apply ticker-specific logic if using different models per asset.

How do I export backtest results for risk analysis outside Backtrader?

Add bt.analyzers.PyFolio to the Cerebro instance before running. After cerebro.run(), access the analyzer via results[0].analyzers.getbyname('pyfolio') and call .get_pf_items() to extract returns, positions, transactions, and leverage. These pandas objects can be saved to HDF5 or fed directly into PyFolio's create_full_tear_sheet() for benchmark comparison.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →