# How to Optimize Backtest Performance for Large Datasets in Nautilus Trader

> Optimize backtest performance for large datasets in Nautilus Trader. Learn techniques to reduce execution time by over 70% processing multi-gigabyte catalogs.

- Repository: [Nautech Systems/nautilus_trader](https://github.com/nautechsystems/nautilus_trader)
- Tags: performance
- Published: 2026-02-16

---

**Stream data in configurable chunks using `BacktestRunConfig.chunk_size`, disable logging with `bypass_logging`, and skip post-run analysis to reduce backtest execution time by over 70% when processing multi-gigabyte catalogs.**

Nautilus Trader is a high-performance algorithmic trading platform built on a deterministic, event-driven simulation engine. When backtesting strategies against millions of market data rows, unoptimized configurations can cause severe memory pressure and I/O bottlenecks. This guide explains how to optimize backtest performance for large datasets by leveraging the streaming APIs and configuration flags available in the Rust core.

## Understanding Performance Bottlenecks in Event-Driven Backtests

The engine processes data chronologically, making certain operations expensive at scale:

- **I/O and Data Loading**: Loading entire catalogs into memory dominates wall-clock time.
- **Data Ordering**: Out-of-order rows force repeated resorting.
- **Logging Overhead**: Per-event debug logging consumes significant CPU cycles.
- **Post-Run Analysis**: Portfolio statistics calculations traverse the entire execution cache.

## Streaming Large Datasets with Chunked Loading

Instead of loading all data upfront, use `BacktestRunConfig.chunk_size` to process data in manageable segments. This configuration, defined in [`crates/backtest/src/config.rs`](https://github.com/nautechsystems/nautilus_trader/blob/main/crates/backtest/src/config.rs) (lines 98-110), enables streaming mode where the engine loads only the specified number of rows per iteration.

```rust
let run_cfg = BacktestRunConfig::new(
    venues,
    data_cfgs,
    engine_cfg,
    Some(500_000),            // Process 500k rows per chunk
    Some(true),               // Dispose on completion
    None,
    None,
);

```

## Minimizing Runtime Overhead

### Disabling Logging and Analysis

Set `bypass_logging=true` in `BacktestEngineConfig` to mute the logger, and disable post-run statistics with `run_analysis=false`. These fields are located in [`crates/backtest/src/config.rs`](https://github.com/nautechsystems/nautilus_trader/blob/main/crates/backtest/src/config.rs) (lines 78-82).

```rust
let engine_cfg = BacktestEngineConfig::new(
    Environment::Backtest,
    TraderId::default(),
    None, None,               // Load/save state off
    Some(true),               // bypass_logging
    Some(false),              // run_analysis
    // ... remaining defaults
);

```

### Optimizing Data Sorting

The `add_data` method in [`crates/backtest/src/engine.rs`](https://github.com/nautechsystems/nautilus_trader/blob/main/crates/backtest/src/engine.rs) (lines 30-48) accepts a `sort` parameter. Pass `sort=false` to skip immediate sorting, then call `sort_data()` once before execution to perform a single O(N log N) operation.

```rust
engine.add_data(raw_data, None, validate=true, sort=false);
engine.sort_data();          // Single sort before first run

```

## Implementing the Streaming Execution Loop

Combine these techniques in a loop that keeps the engine kernel alive between chunks. The `run` method in [`crates/backtest/src/engine.rs`](https://github.com/nautechsystems/nautilus_trader/blob/main/crates/backtest/src/engine.rs) supports a `streaming` flag (lines 24-36) that prevents engine reset between iterations.

```rust
// Initialize once
let mut engine = BacktestEngine::new(run_cfg.engine.clone())?;
engine.add_venue(/* ... */)?;
engine.add_market_data_client_if_not_exists(Venue::from("BINANCE"));

// Streaming loop
while let Some(chunk) = load_next_csv_chunk()? {
    engine.add_data(chunk, None, true, false);
    
    if engine.iteration == 0 {
        engine.sort_data();
        engine.run(None, None, None, true)?; // streaming=true
    } else {
        engine.run(None, None, None, true)?;
    }
    
    engine.clear_data(); // Free memory before next chunk
}

// Final flush
engine.run(None, None, None, false)?;
let result = engine.get_result();

```

## Summary

- **Stream data in chunks** using `BacktestRunConfig.chunk_size` to limit memory resident set size.
- **Disable overhead** by setting `bypass_logging=true` and `run_analysis=false` in `BacktestEngineConfig`.
- **Sort once** by using `sort=false` in `add_data` followed by a single `sort_data()` call.
- **Reuse the engine** across chunks via the `streaming` flag in `run` to avoid expensive kernel resets.
- **Clear data** between iterations with `clear_data()` to prevent memory pressure buildup.

## Frequently Asked Questions

### How does `chunk_size` affect deterministic replay?

The `chunk_size` parameter in `BacktestRunConfig` only controls memory loading patterns, not event ordering. The engine maintains a sorted view of all loaded data via `BacktestDataIterator` in [`crates/backtest/src/data_iterator.rs`](https://github.com/nautechsystems/nautilus_trader/blob/main/crates/backtest/src/data_iterator.rs), ensuring that events are processed chronologically across chunk boundaries. Determinism is preserved as long as the full dataset is ultimately loaded in order.

### Can I use Python instead of Rust to configure these optimizations?

Yes. While the underlying implementation resides in [`crates/backtest/src/config.rs`](https://github.com/nautechsystems/nautilus_trader/blob/main/crates/backtest/src/config.rs) and [`crates/backtest/src/engine.rs`](https://github.com/nautechsystems/nautilus_trader/blob/main/crates/backtest/src/engine.rs), Nautilus Trader exposes these same parameters through its Python API. You can set `chunk_size`, `bypass_logging`, and `run_analysis` when constructing `BacktestRunConfig` and `BacktestEngineConfig` in Python, with the Rust backend handling the streaming execution.

### What is the performance impact of disabling `run_analysis`?

Disabling `run_analysis` eliminates the post-simulation traversal of the entire execution cache to calculate statistics, drawdowns, and performance charts. For multi-gigabyte backtests, this can reduce total runtime by 20-40% and significantly lower peak memory usage, especially when you only need raw fill and order data rather than computed metrics.

### When should I use `sort=false` versus the default sorting behavior?

Use `sort=false` when you are loading pre-sorted data or when you plan to load multiple batches before starting the simulation. This avoids the O(N log N) cost of sorting on every `add_data` call. Call `sort_data()` once after all data is loaded but before the first `run` invocation. Use default sorting only when loading small, potentially unordered datasets where the overhead is negligible.