How to Run a Backtest with Custom CSV Data in Nautilus Trader

To run a backtest with custom CSV data in Nautilus Trader, load your file using TestDataProvider.read_csv_bars(), convert the DataFrame to Bar objects via BarDataWrangler, and inject them into a BacktestEngine using add_data() before calling run().

Nautilus Trader provides a flexible data-loading layer that can ingest user-supplied CSV files for both bars (OHLCV) and ticks (quotes/trades). By leveraging the built-in TestDataProvider and BarDataWrangler, you can run a backtest with custom CSV data without modifying the core engine. This workflow mirrors the official example fx_ema_cross_bracket_gbpusd_bars_external.py—requiring only a file path change to use your own datasets.

Preparing Your CSV File Format

Before loading data, ensure your CSV follows the expected schema defined in nautilus_trader/persistence/loaders.py. The CSVBarDataLoader expects specific column names and timestamp formats.

For bar data (OHLCV), your CSV must contain:

  • timestamp: ISO-8601 strings, Unix epoch seconds, or nanoseconds
  • open, high, low, close: Price values as floats
  • volume: Trading volume as integer or float

For tick data (quotes or trades), include:

  • timestamp: As above
  • price: Execution or quote price
  • size: Quantity
  • Optional: side, trade_id for trade tick identification

Loading CSV Data with TestDataProvider

The TestDataProvider class in nautilus_trader/test_kit/providers.py provides convenience methods to read CSV files without manual pandas configuration. It uses CSVBarDataLoader from nautilus_trader/persistence/loaders.py under the hood.

from nautilus_trader.test_kit.providers import TestDataProvider

provider = TestDataProvider()
df = provider.read_csv_bars("path/to/your_data.csv")

This returns a pandas DataFrame with a properly parsed datetime index. The loader automatically handles mixed timestamp formats using format="mixed".

Converting DataFrames to Nautilus Bar Objects

Raw DataFrames must be converted to Nautilus data objects before ingestion. Use BarDataWrangler from nautilus_trader/persistence/wranglers.py to handle type safety and timestamp alignment.

from nautilus_trader.persistence.wranglers import BarDataWrangler
from nautilus_trader.model.data import BarType
from nautilus_trader.test_kit.providers import TestInstrumentProvider
from nautilus_trader.model.identifiers import Venue

venue = Venue("SIM")
instrument = TestInstrumentProvider.default_fx_ccy("GBP/USD", venue)

bar_type = BarType.from_str("GBP/USD.SIM-1-MINUTE-BID-EXTERNAL")
wrangler = BarDataWrangler(bar_type=bar_type, instrument=instrument)
bars = wrangler.process(data=df)

The process method returns a list of Bar objects with correct bar_type metadata and instrument references, ensuring the backtest engine interprets the data correctly.

Configuring the Backtest Engine

With data prepared, initialize the BacktestEngine from nautilus_trader/backtest/engine.py. Configure venues, instruments, and starting capital before injecting your custom data.

from nautilus_trader.backtest.engine import BacktestEngine
from nautilus_trader.backtest.config import BacktestEngineConfig
from nautilus_trader.model.identifiers import TraderId, Venue
from nautilus_trader.model.enums import AccountType, OmsType
from nautilus_trader.model.objects import Money
from nautilus_trader.model.currencies import USD

engine = BacktestEngine(
    config=BacktestEngineConfig(trader_id=TraderId("BACKTESTER-001"))
)

venue = Venue("SIM")
engine.add_venue(
    venue=venue,
    oms_type=OmsType.HEDGING,
    account_type=AccountType.MARGIN,
    base_currency=USD,
    starting_balances=[Money(100_000, USD)],
)

engine.add_instrument(instrument)
engine.add_data(bars)

The add_data method accepts the list of Bar objects created by the wrangler, registering them for the simulation timeline.

Attaching Strategies and Executing the Backtest

After configuring the engine, attach your strategy implementation and execute the simulation. The following example uses the built-in OrderBookImbalance strategy, but you can substitute any custom Strategy subclass.

from nautilus_trader.examples.strategies.orderbook_imbalance import (
    OrderBookImbalance,
    OrderBookImbalanceConfig,
)

strategy_config = OrderBookImbalanceConfig(
    instrument_id=instrument.id,
    max_trade_size=10,
    min_seconds_between_triggers=1.0,
)
engine.add_strategy(OrderBookImbalance(config=strategy_config))

engine.run()

The run method executes the backtest from the first to the last timestamp in your custom CSV data.

Generating Performance Reports

Once the simulation completes, extract performance metrics using the trader's report generators available in the engine.

import pandas as pd

with pd.option_context("display.max_rows", 100, "display.width", 300):
    print(engine.trader.generate_account_report(venue))
    print(engine.trader.generate_order_fills_report())
    print(engine.trader.generate_positions_report())

These reports provide account balances, fill details, and position snapshots throughout the backtest period.

Optimizing for Large CSV Files

For CSV files containing millions of rows that exceed available memory, process the data in chunks rather than loading the entire file at once. This approach prevents memory exhaustion while allowing the backtest to span extensive historical periods.

import pandas as pd

chunk_size = 100_000
for chunk in pd.read_csv("large_file.csv", chunksize=chunk_size):
    bars_chunk = wrangler.process(chunk)
    engine.add_data(bars_chunk)

engine.run()

By incrementally adding each chunk to the engine via engine.add_data() before calling engine.run(), you can backtest multi-gigabyte datasets on standard hardware.

Summary

  • Prepare CSV files with timestamp, open, high, low, close, and volume columns for bar data, ensuring timestamps are ISO-8601 or Unix epoch format.
  • Load data using TestDataProvider.read_csv_bars() from nautilus_trader/test_kit/providers.py, which handles timestamp parsing automatically via CSVBarDataLoader.
  • Transform DataFrames into Bar objects with BarDataWrangler from nautilus_trader/persistence/wranglers.py to ensure type safety and instrument alignment.
  • Configure the BacktestEngine from nautilus_trader/backtest/engine.py with venues, instruments, and starting balances using add_venue() and add_instrument().
  • Execute by calling engine.run() after attaching your strategy with add_strategy(), and generate reports using engine.trader methods.
  • Optimize memory usage by processing large CSVs in chunks via pandas.read_csv(chunksize=...) before adding to the engine.

Frequently Asked Questions

What timestamp formats does Nautilus Trader accept in CSV files?

Nautilus Trader accepts ISO-8601 formatted strings, Unix epoch seconds, and Unix epoch nanoseconds. The CSVBarDataLoader in nautilus_trader/persistence/loaders.py uses pandas with format="mixed" by default, automatically inferring the correct format from your timestamp column without manual specification.

Can I backtest with tick-level CSV data instead of bars?

Yes. Use TestDataProvider.read_csv_ticks() to load tick data, then process the DataFrame with TickDataWrangler from nautilus_trader/persistence/wranglers.py to generate QuoteTick or TradeTick objects. The workflow remains identical to bar data: load, wrangle, add to engine via add_data(), and run.

How do I handle CSV files with non-standard column names?

If your CSV uses different column names (e.g., Open instead of open), load the CSV into a pandas DataFrame manually using pd.read_csv(), rename the columns to match the expected schema (timestamp, open, high, low, close, volume), then pass the DataFrame to BarDataWrangler.process() instead of using TestDataProvider.read_csv_bars().

Is there a limit to how much CSV data I can load into a backtest?

There is no hardcoded limit, but available RAM constrains the process. For datasets exceeding memory capacity, process the CSV in chunks using pandas.read_csv(chunksize=...) and incrementally add each chunk to the engine via engine.add_data() before calling engine.run(). This allows backtesting multi-gigabyte datasets on standard hardware without memory exhaustion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →