# How to Backtest ML-Driven Trading Strategies Using Backtrader

> Learn to backtest ML-driven trading strategies with Backtrader. Extend PandasData and create custom strategies for dynamic capital allocation based on ML model outputs.

- Repository: [Stefan Jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading)
- Tags: how-to-guide
- Published: 2026-06-02

---

**Backtrader enables event-driven backtesting of machine learning trading signals by extending `PandasData` to include prediction columns and implementing a custom `Strategy` class that dynamically allocates capital based on ML model outputs.**

The notebook `08_ml4t_workflow/03_backtesting_with_backtrader.ipynb` in the [stefan-jansen/machine-learning-for-trading](https://github.com/stefan-jansen/machine-learning-for-trading) repository demonstrates a production-ready workflow for integrating scikit-learn predictions with the Backtrader engine. This approach supports realistic simulation features including custom commission schemes, multi-asset portfolios, and PyFolio-compatible performance analytics.

## Setting Up the Backtrader Environment

Begin by importing the required libraries and defining a custom commission structure. The repository silences warnings and configures pandas display options to handle large equity curves cleanly.

```python
import warnings
warnings.filterwarnings('ignore')
import backtrader as bt
from backtrader.feeds import PandasData
import pandas as pd

```

### Custom Commission Schemes

For realistic cost modeling, the repository implements a fixed-fee commission scheme rather than percentage-based fees. This is critical when testing high-frequency ML signals where flat fees dominate transaction costs.

```python
class FixedCommisionScheme(bt.CommInfoBase):
    params = (('commission', .02), ('stocklike', True),
              ('commtype', bt.CommInfoBase.COMM_FIXED),)

    def _getcommission(self, size, price, pseudoexec):
        return abs(size) * self.p.commission

```

Attach this scheme to the broker using `cerebro.broker.addcommissioninfo(FixedCommisionScheme())` after initializing the Cerebro engine.

## Creating a Custom Data Feed for ML Signals

Standard Backtrader data feeds only recognize OHLCV columns. To inject ML predictions into the backtest, you must extend `PandasData` to declare a new `predicted` line that stores model outputs (positive values for long signals, negative for short).

```python
class SignalData(PandasData):
    cols = ['open', 'high', 'low', 'close', 'volume', 'predicted']
    lines = tuple(cols)
    params = {c: -1 for c in cols}
    params.update({'datetime': None})
    params = tuple(params.items())

```

**Key implementation details:**
- The `lines` tuple creates accessible attributes (`.open`, `.high`, ..., `.predicted`) inside the strategy
- Parameter values of `-1` instruct Backtrader to auto-locate columns by name in the input DataFrame
- The input DataFrame must contain a `predicted` column with numerical signals derived from your ML model

## Implementing the ML Strategy Logic

The `MLStrategy` class implements a long/short equity market-neutral approach. It ranks all universes by predicted returns each day and rebalances to hold equal-weight positions in the top-N longs and bottom-N shorts.

```python
class MLStrategy(bt.Strategy):
    params = (('n_positions', 10), ('min_positions', 5), ('verbose', False))

    def next(self):
        today = self.datas[0].datetime.date()
        up, down = {}, {}
        
        # Collect predictions for current date

        for data in self.datas:
            if data.datetime.date() == today:
                if data.predicted[0] > 0:
                    up[data._name] = data.predicted[0]
                elif data.predicted[0] < 0:
                    down[data._name] = data.predicted[0]
        
        # Rank and select top/bottom signals

        shorts = sorted(down, key=down.get)[:self.p.n_positions]
        longs = sorted(up, key=up.get, reverse=True)[:self.p.n_positions]
        
        # Risk management: require minimum positions to avoid concentration

        if len(shorts) < self.p.min_positions or len(longs) < self.p.min_positions:
            longs, shorts = [], []
        
        # Close positions not in current signals

        for ticker in self.getdatanames():
            if ticker not in longs + shorts:
                self.order_target_percent(data=ticker, target=0)
        
        # Equal-weight allocation

        short_target = -1 / max(self.p.n_positions, len(shorts))
        long_target = 1 / max(self.p.n_positions, len(longs))
        
        for ticker in shorts:
            self.order_target_percent(data=ticker, target=short_target)
        for ticker in longs:
            self.order_target_percent(data=ticker, target=long_target)

```

**Critical logic components:**
- **Signal aggregation**: The strategy iterates through `self.datas` (all registered tickers) and checks `data.predicted[0]` to access the current bar's ML prediction
- **Ranking mechanism**: Python's `sorted()` function creates priority queues based on signal strength
- **Position sizing**: `order_target_percent` enforces equal-weight exposure, automatically handling existing positions and rounding
- **Safety constraints**: The `min_positions` parameter prevents the strategy from taking concentrated bets when few signals are available

## Configuring Cerebro and Running the Backtest

Load pre-computed predictions (typically stored in HDF5 format from an earlier ML training step) and attach each ticker as a separate data feed. The example uses `pd.read_hdf` to retrieve OHLCV plus predictions from `00_data/backtest.h5`.

```python
cerebro = bt.Cerebro()
cerebro.broker.setcash(10000.0)

# Load backtest data with predictions

data = pd.read_hdf('00_data/backtest.h5', 'data').sort_index()

# Add each ticker individually

for ticker in data.index.get_level_values(0).unique():
    df = data.loc[pd.IndexSlice[ticker, :], :].droplevel('ticker', axis=0)
    df.index.name = 'datetime'
    cerebro.adddata(SignalData(dataname=df), name=ticker)

# Attach strategy and analyzer

cerebro.addstrategy(MLStrategy, n_positions=25, min_positions=20)
cerebro.addanalyzer(bt.analyzers.PyFolio, _name='pyfolio')

results = cerebro.run()
ending_value = cerebro.broker.getvalue()
print(f'Final Portfolio Value: {ending_value:,.2f}')

```

## Analyzing Results with PyFolio

Backtrader's built-in `PyFolio` analyzer extracts four objects required for institutional-grade performance reporting: returns, positions, transactions, and gross leverage. Export these to HDF5 for persistent storage or pass directly to PyFolio's tear sheet generator.

```python
pyfolio_analyzer = results[0].analyzers.getbyname('pyfolio')
returns, positions, transactions, gross_lev = pyfolio_analyzer.get_pf_items()

# Export for downstream analysis

returns.to_hdf('backtrader.h5', 'returns')
positions.to_hdf('backtrader.h5', 'positions')
transactions.to_hdf('backtrader.h5', 'transactions')

# Generate full tear sheet with benchmark

import pandas_datareader.data as web
benchmark = web.DataReader('SP500', 'fred', '2014', '2018').pct_change().tz_localize('UTC')

import pyfolio as pf
pf.create_full_tear_sheet(returns, transactions=transactions,
                          positions=positions, benchmark_rets=benchmark.dropna())

```

This workflow produces Sharpe ratios, drawdown analysis, and beta-adjusted returns essential for validating ML-driven alpha generation.

## Complete Minimal Example

For rapid prototyping, use this condensed version that requires only a single CSV containing OHLCV columns plus a `predicted` column:

```python
import backtrader as bt
import pandas as pd
from backtrader.feeds import PandasData

class SignalData(PandasData):
    lines = ('predicted',)
    params = (('datetime', None), ('open', -1), ('high', -1),
              ('low', -1), ('close', -1), ('volume', -1),
              ('predicted', -1))

class SimpleML(bt.Strategy):
    def next(self):
        for data in self.datas:
            if data.predicted[0] > 0:
                self.order_target_percent(data=data, target=0.1)
            elif data.predicted[0] < 0:
                self.order_target_percent(data=data, target=-0.1)

df = pd.read_csv('predictions.csv', parse_dates=['datetime'])
df.set_index('datetime', inplace=True)

cerebro = bt.Cerebro()
cerebro.broker.setcash(100000.0)
cerebro.adddata(SignalData(dataname=df), name='asset')
cerebro.addstrategy(SimpleML)

cerebro.run()
print(f'Final Value: {cerebro.broker.getvalue():,.2f}')

```

## Summary

- **Custom data feeds**: Extend `PandasData` in `08_ml4t_workflow/03_backtesting_with_backtrader.ipynb` to map ML prediction columns to Backtrader lines using the `lines` and `params` class attributes.
- **Dynamic rebalancing**: The `MLStrategy.next()` method ranks universe constituents by `data.predicted[0]` values and rebalances daily using `order_target_percent` for equal-weight exposures.
- **Risk controls**: Implement minimum position thresholds and concentration limits via strategy parameters (`n_positions`, `min_positions`).
- **Performance analytics**: Attach `bt.analyzers.PyFolio` to export returns, positions, and transactions compatible with PyFolio tear sheets for Sharpe ratio and drawdown analysis.
- **Commission modeling**: Override `bt.CommInfoBase` to implement fixed-cost or percentage-based fee structures that accurately reflect trading costs on ML signal turnover.

## Frequently Asked Questions

### What input format does Backtrader require for ML predictions?

Backtrader requires a pandas DataFrame indexed by datetime containing standard OHLCV columns plus your model's predictions. In `SignalData`, the `params` dictionary maps these column names to line indices, with `-1` enabling automatic column detection. The DataFrame is passed via `SignalData(dataname=df)` when calling `cerebro.adddata()`.

### How does the MLStrategy handle position sizing and rebalancing?

The strategy calculates equal-weight targets for long and short sides separately using `1 / max(self.p.n_positions, len(longs))`. It calls `self.order_target_percent()` for each selected ticker, which automatically sizes orders to reach the target percentage of current portfolio value while closing stale positions not present in the current signal set.

### Can I backtest multiple assets simultaneously with different ML models?

Yes. The repository demonstrates multi-asset backtesting by iterating through tickers in an HDF5 file and calling `cerebro.adddata()` for each. Each ticker maintains its own `predicted` line. The strategy accesses these via `self.datas` and filters by `data._name` to apply ticker-specific logic if using different models per asset.

### How do I export backtest results for risk analysis outside Backtrader?

Add `bt.analyzers.PyFolio` to the Cerebro instance before running. After `cerebro.run()`, access the analyzer via `results[0].analyzers.getbyname('pyfolio')` and call `.get_pf_items()` to extract returns, positions, transactions, and leverage. These pandas objects can be saved to HDF5 or fed directly into PyFolio's `create_full_tear_sheet()` for benchmark comparison.