# When Should You Use `memory_saving_mode` in TabPFN to Prevent OOM Errors

> Prevent OOM errors with TabPFN's memory_saving_mode. Learn when to enable it for limited hardware, large test sets, or low memory training to optimize performance and avoid crashes.

- Repository: [Prior Labs/TabPFN](https://github.com/PriorLabs/TabPFN)
- Tags: performance
- Published: 2026-05-06

---

**Enable `memory_saving_mode` when working with limited GPU or CPU memory, large test datasets where the number of test instances exceeds training instances, or when `fit_mode="low_memory"` is active to automatically split tensor operations into smaller chunks and avoid out-of-memory failures.**

The `memory_saving_mode` parameter in PriorLabs/TabPFN is a runtime optimization that reduces peak memory consumption by chunking expensive matrix operations. According to the source code in [`src/tabpfn/architectures/base/memory.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/architectures/base/memory.py), this feature is defined as `MemorySavingMode = Union[bool, Literal["auto"], float, int]` and propagates through the inference engine to prevent allocation of massive intermediate tensors.

## When to Activate Memory Saving Mode

TabPFN builds inference-time matrices with dimensions `[n_test × n_train]`, which can explode memory usage when datasets are large. You should enable this feature in specific hardware and data scenarios.

### Limited GPU or CPU Memory Resources

When running TabPFN on a single-GPU workstation or CPU-only environment with strict memory caps, set `memory_saving_mode=True` or `memory_saving_mode="auto"`. The mode wraps core computations in `torch.chunk`-style loops, keeping intermediate tensors below available limits by processing sub-batches sequentially.

### Large Test Sets Where n_test Exceeds n_train

If your test set is significantly larger than your training set, the `[n_test × n_train]` distance matrix can consume gigabytes of VRAM. Activate `memory_saving_mode` to chunk this matrix construction, trading modest additional runtime for a bounded memory peak that scales with available hardware rather than dataset size.

### Low-Memory Training Configurations

When using `fit_mode="low_memory"` during model fitting, you should pair it with `memory_saving_mode` during inference. The low-memory fit mode minimizes preprocessing overhead, but the inference path can still allocate large prediction matrices. Enabling memory saving adds a second layer of protection against OOM errors in [`src/tabpfn/inference.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference.py).

### Repeated Prediction Workflows

If you plan to call `.predict()` multiple times on the same dataset in a single session, `memory_saving_mode` helps by managing tensor allocation patterns efficiently, preventing cumulative memory fragmentation that often triggers OOM errors in iterative pipelines.

## How `memory_saving_mode` Is Implemented

The implementation spans multiple critical files in the repository:

- **[`src/tabpfn/regressor.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/regressor.py)** (line 236): The `TabPFNRegressor` constructor accepts `memory_saving_mode` and stores it as `self.memory_saving_mode`. Lines 350-383 contain the docstring explaining the `"auto"`, `True`/`False`, and numeric behaviors.
- **[`src/tabpfn/classifier.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/classifier.py)**: Mirrors the regressor implementation for `TabPFNClassifier`.
- **[`src/tabpfn/inference.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference.py)**: The flag is passed as `save_peak_mem` to low-level architecture calls at lines 108, 245, 283, 342, 503, 624, 811, and 1021, where it controls chunking logic.
- **[`src/tabpfn/architectures/base/memory.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/architectures/base/memory.py)**: Defines the type hints and default behavior for the memory saving system.

The test suite verifies equivalence between standard and memory-saving paths in [`tests/test_architectures/test_tabpfn_v3.py`](https://github.com/PriorLabs/TabPFN/blob/main/tests/test_architectures/test_tabpfn_v3.py) (lines 33-48), ensuring predictions remain identical regardless of chunking strategy.

## Configuration Options and Code Examples

The `memory_saving_mode` parameter accepts booleans, the string `"auto"`, or numeric values representing memory fractions.

### Automatic OOM Prevention

Use `"auto"` to let TabPFN detect when a batch would exceed memory caps and automatically switch to chunked evaluation:

```python
from tabpfn import TabPFNRegressor

reg = TabPFNRegressor.create_default_for_version(
    fit_mode="low_memory",
    memory_saving_mode="auto",
    random_state=0,
)

reg.fit(X_train, y_train)
y_pred = reg.predict(X_test)

```

### Explicit Memory Budget Control

Pass a float between 0 and 1 to target a specific fraction of available memory, or use `True` to force chunked mode:

```python

# Target 50% of available GPU/CPU memory

reg = TabPFNRegressor(
    memory_saving_mode=0.5,
    random_state=0,
)

# Force chunked evaluation regardless of memory detection

reg = TabPFNRegressor(
    memory_saving_mode=True,
    random_state=0,
)

```

### Classifier Implementation

The identical API works for classification tasks:

```python
from tabpfn import TabPFNClassifier

clf = TabPFNClassifier(
    fit_mode="low_memory",
    memory_saving_mode=True,
    random_state=42,
)

clf.fit(X_train, y_train)
pred = clf.predict(X_test)

```

### Advanced Performance Options

You can also control memory via `PerformanceOptions` with the `save_peak_memory_factor` argument, which specifies how aggressively to reduce peak memory (higher values = more chunking):

```python
from tabpfn import PerformanceOptions

perf_opts = PerformanceOptions(save_peak_memory_factor=4)
y_pred = reg.predict(X_test, performance_options=perf_opts)

```

## Performance Trade-offs and Verification

Enabling `memory_saving_mode` introduces a modest runtime overhead proportional to the number of chunks required. However, the implementation ensures numerical equivalence—tests in [`tests/test_regressor_interface.py`](https://github.com/PriorLabs/TabPFN/blob/main/tests/test_regressor_interface.py) and [`tests/test_classifier_interface.py`](https://github.com/PriorLabs/TabPFN/blob/main/tests/test_classifier_interface.py) confirm that `fit_mode="low_memory"` combined with memory saving produces identical predictions to standard inference paths.

## Summary

- **`memory_saving_mode`** is defined in [`src/tabpfn/architectures/base/memory.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/architectures/base/memory.py) as `Union[bool, Literal["auto"], float, int]`.
- Enable it when processing large test sets, working with limited GPU/CPU memory, or using `fit_mode="low_memory"`.
- The flag propagates through [`src/tabpfn/inference.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference.py) as `save_peak_mem`, triggering chunked matrix operations.
- Set to `"auto"` for automatic detection, `True` to force chunking, or a float (e.g., `0.5`) to target specific memory fractions.
- Predictions remain numerically identical regardless of chunking strategy, as verified by the test suite.

## Frequently Asked Questions

### What value should I set for `memory_saving_mode` if I have 8GB GPU memory?

Use `memory_saving_mode="auto"` to let TabPFN dynamically detect memory pressure, or set `memory_saving_mode=0.7` to explicitly target 70% of your available VRAM as a safety buffer. If you still encounter OOM errors, switch to `memory_saving_mode=True` to force maximum chunking.

### Does enabling `memory_saving_mode` change my prediction results?

No. The implementation in [`src/tabpfn/inference.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference.py) ensures chunked operations produce mathematically identical results to standard batched operations. The test suite in [`tests/test_architectures/test_tabpfn_v3.py`](https://github.com/PriorLabs/TabPFN/blob/main/tests/test_architectures/test_tabpfn_v3.py) specifically validates output equivalence between memory-saving and standard modes.

### Can I use `memory_saving_mode` with `TabPFNClassifier`?

Yes. Both `TabPFNRegressor` and `TabPFNClassifier` support the parameter in their constructors, as implemented in [`src/tabpfn/regressor.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/regressor.py) and [`src/tabpfn/classifier.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/classifier.py) respectively. The underlying inference engine handles chunking identically for both regression and classification tasks.

### How does `memory_saving_mode` interact with `fit_mode="low_memory"`?

These settings operate at different stages. `fit_mode="low_memory"` reduces memory during preprocessing and model fitting, while `memory_saving_mode` activates during the prediction phase in [`src/tabpfn/inference.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference.py). Using them together provides comprehensive memory protection throughout the entire TabPFN workflow, preventing OOM errors during both fitting and inference on constrained hardware.