How to Run a Backtest with Custom CSV Data in Nautilus Trader
To run a backtest with custom CSV data in Nautilus Trader, load your file using TestDataProvider.read_csv_bars(), convert the DataFrame to Bar objects via BarDataWrangler, and inject them into a BacktestEngine using add_data() before calling run().
Nautilus Trader provides a flexible data-loading layer that can ingest user-supplied CSV files for both bars (OHLCV) and ticks (quotes/trades). By leveraging the built-in TestDataProvider and BarDataWrangler, you can run a backtest with custom CSV data without modifying the core engine. This workflow mirrors the official example fx_ema_cross_bracket_gbpusd_bars_external.py—requiring only a file path change to use your own datasets.
Preparing Your CSV File Format
Before loading data, ensure your CSV follows the expected schema defined in nautilus_trader/persistence/loaders.py. The CSVBarDataLoader expects specific column names and timestamp formats.
For bar data (OHLCV), your CSV must contain:
timestamp: ISO-8601 strings, Unix epoch seconds, or nanosecondsopen,high,low,close: Price values as floatsvolume: Trading volume as integer or float
For tick data (quotes or trades), include:
timestamp: As aboveprice: Execution or quote pricesize: Quantity- Optional:
side,trade_idfor trade tick identification
Loading CSV Data with TestDataProvider
The TestDataProvider class in nautilus_trader/test_kit/providers.py provides convenience methods to read CSV files without manual pandas configuration. It uses CSVBarDataLoader from nautilus_trader/persistence/loaders.py under the hood.
from nautilus_trader.test_kit.providers import TestDataProvider
provider = TestDataProvider()
df = provider.read_csv_bars("path/to/your_data.csv")
This returns a pandas DataFrame with a properly parsed datetime index. The loader automatically handles mixed timestamp formats using format="mixed".
Converting DataFrames to Nautilus Bar Objects
Raw DataFrames must be converted to Nautilus data objects before ingestion. Use BarDataWrangler from nautilus_trader/persistence/wranglers.py to handle type safety and timestamp alignment.
from nautilus_trader.persistence.wranglers import BarDataWrangler
from nautilus_trader.model.data import BarType
from nautilus_trader.test_kit.providers import TestInstrumentProvider
from nautilus_trader.model.identifiers import Venue
venue = Venue("SIM")
instrument = TestInstrumentProvider.default_fx_ccy("GBP/USD", venue)
bar_type = BarType.from_str("GBP/USD.SIM-1-MINUTE-BID-EXTERNAL")
wrangler = BarDataWrangler(bar_type=bar_type, instrument=instrument)
bars = wrangler.process(data=df)
The process method returns a list of Bar objects with correct bar_type metadata and instrument references, ensuring the backtest engine interprets the data correctly.
Configuring the Backtest Engine
With data prepared, initialize the BacktestEngine from nautilus_trader/backtest/engine.py. Configure venues, instruments, and starting capital before injecting your custom data.
from nautilus_trader.backtest.engine import BacktestEngine
from nautilus_trader.backtest.config import BacktestEngineConfig
from nautilus_trader.model.identifiers import TraderId, Venue
from nautilus_trader.model.enums import AccountType, OmsType
from nautilus_trader.model.objects import Money
from nautilus_trader.model.currencies import USD
engine = BacktestEngine(
config=BacktestEngineConfig(trader_id=TraderId("BACKTESTER-001"))
)
venue = Venue("SIM")
engine.add_venue(
venue=venue,
oms_type=OmsType.HEDGING,
account_type=AccountType.MARGIN,
base_currency=USD,
starting_balances=[Money(100_000, USD)],
)
engine.add_instrument(instrument)
engine.add_data(bars)
The add_data method accepts the list of Bar objects created by the wrangler, registering them for the simulation timeline.
Attaching Strategies and Executing the Backtest
After configuring the engine, attach your strategy implementation and execute the simulation. The following example uses the built-in OrderBookImbalance strategy, but you can substitute any custom Strategy subclass.
from nautilus_trader.examples.strategies.orderbook_imbalance import (
OrderBookImbalance,
OrderBookImbalanceConfig,
)
strategy_config = OrderBookImbalanceConfig(
instrument_id=instrument.id,
max_trade_size=10,
min_seconds_between_triggers=1.0,
)
engine.add_strategy(OrderBookImbalance(config=strategy_config))
engine.run()
The run method executes the backtest from the first to the last timestamp in your custom CSV data.
Generating Performance Reports
Once the simulation completes, extract performance metrics using the trader's report generators available in the engine.
import pandas as pd
with pd.option_context("display.max_rows", 100, "display.width", 300):
print(engine.trader.generate_account_report(venue))
print(engine.trader.generate_order_fills_report())
print(engine.trader.generate_positions_report())
These reports provide account balances, fill details, and position snapshots throughout the backtest period.
Optimizing for Large CSV Files
For CSV files containing millions of rows that exceed available memory, process the data in chunks rather than loading the entire file at once. This approach prevents memory exhaustion while allowing the backtest to span extensive historical periods.
import pandas as pd
chunk_size = 100_000
for chunk in pd.read_csv("large_file.csv", chunksize=chunk_size):
bars_chunk = wrangler.process(chunk)
engine.add_data(bars_chunk)
engine.run()
By incrementally adding each chunk to the engine via engine.add_data() before calling engine.run(), you can backtest multi-gigabyte datasets on standard hardware.
Summary
- Prepare CSV files with
timestamp,open,high,low,close, andvolumecolumns for bar data, ensuring timestamps are ISO-8601 or Unix epoch format. - Load data using
TestDataProvider.read_csv_bars()fromnautilus_trader/test_kit/providers.py, which handles timestamp parsing automatically viaCSVBarDataLoader. - Transform DataFrames into
Barobjects withBarDataWranglerfromnautilus_trader/persistence/wranglers.pyto ensure type safety and instrument alignment. - Configure the
BacktestEnginefromnautilus_trader/backtest/engine.pywith venues, instruments, and starting balances usingadd_venue()andadd_instrument(). - Execute by calling
engine.run()after attaching your strategy withadd_strategy(), and generate reports usingengine.tradermethods. - Optimize memory usage by processing large CSVs in chunks via
pandas.read_csv(chunksize=...)before adding to the engine.
Frequently Asked Questions
What timestamp formats does Nautilus Trader accept in CSV files?
Nautilus Trader accepts ISO-8601 formatted strings, Unix epoch seconds, and Unix epoch nanoseconds. The CSVBarDataLoader in nautilus_trader/persistence/loaders.py uses pandas with format="mixed" by default, automatically inferring the correct format from your timestamp column without manual specification.
Can I backtest with tick-level CSV data instead of bars?
Yes. Use TestDataProvider.read_csv_ticks() to load tick data, then process the DataFrame with TickDataWrangler from nautilus_trader/persistence/wranglers.py to generate QuoteTick or TradeTick objects. The workflow remains identical to bar data: load, wrangle, add to engine via add_data(), and run.
How do I handle CSV files with non-standard column names?
If your CSV uses different column names (e.g., Open instead of open), load the CSV into a pandas DataFrame manually using pd.read_csv(), rename the columns to match the expected schema (timestamp, open, high, low, close, volume), then pass the DataFrame to BarDataWrangler.process() instead of using TestDataProvider.read_csv_bars().
Is there a limit to how much CSV data I can load into a backtest?
There is no hardcoded limit, but available RAM constrains the process. For datasets exceeding memory capacity, process the CSV in chunks using pandas.read_csv(chunksize=...) and incrementally add each chunk to the engine via engine.add_data() before calling engine.run(). This allows backtesting multi-gigabyte datasets on standard hardware without memory exhaustion.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →