When Should You Use `memory_saving_mode` in TabPFN to Prevent OOM Errors

Enable memory_saving_mode when working with limited GPU or CPU memory, large test datasets where the number of test instances exceeds training instances, or when fit_mode="low_memory" is active to automatically split tensor operations into smaller chunks and avoid out-of-memory failures.

The memory_saving_mode parameter in PriorLabs/TabPFN is a runtime optimization that reduces peak memory consumption by chunking expensive matrix operations. According to the source code in src/tabpfn/architectures/base/memory.py, this feature is defined as MemorySavingMode = Union[bool, Literal["auto"], float, int] and propagates through the inference engine to prevent allocation of massive intermediate tensors.

When to Activate Memory Saving Mode

TabPFN builds inference-time matrices with dimensions [n_test × n_train], which can explode memory usage when datasets are large. You should enable this feature in specific hardware and data scenarios.

Limited GPU or CPU Memory Resources

When running TabPFN on a single-GPU workstation or CPU-only environment with strict memory caps, set memory_saving_mode=True or memory_saving_mode="auto". The mode wraps core computations in torch.chunk-style loops, keeping intermediate tensors below available limits by processing sub-batches sequentially.

Large Test Sets Where n_test Exceeds n_train

If your test set is significantly larger than your training set, the [n_test × n_train] distance matrix can consume gigabytes of VRAM. Activate memory_saving_mode to chunk this matrix construction, trading modest additional runtime for a bounded memory peak that scales with available hardware rather than dataset size.

Low-Memory Training Configurations

When using fit_mode="low_memory" during model fitting, you should pair it with memory_saving_mode during inference. The low-memory fit mode minimizes preprocessing overhead, but the inference path can still allocate large prediction matrices. Enabling memory saving adds a second layer of protection against OOM errors in src/tabpfn/inference.py.

Repeated Prediction Workflows

If you plan to call .predict() multiple times on the same dataset in a single session, memory_saving_mode helps by managing tensor allocation patterns efficiently, preventing cumulative memory fragmentation that often triggers OOM errors in iterative pipelines.

How memory_saving_mode Is Implemented

The implementation spans multiple critical files in the repository:

  • src/tabpfn/regressor.py (line 236): The TabPFNRegressor constructor accepts memory_saving_mode and stores it as self.memory_saving_mode. Lines 350-383 contain the docstring explaining the "auto", True/False, and numeric behaviors.
  • src/tabpfn/classifier.py: Mirrors the regressor implementation for TabPFNClassifier.
  • src/tabpfn/inference.py: The flag is passed as save_peak_mem to low-level architecture calls at lines 108, 245, 283, 342, 503, 624, 811, and 1021, where it controls chunking logic.
  • src/tabpfn/architectures/base/memory.py: Defines the type hints and default behavior for the memory saving system.

The test suite verifies equivalence between standard and memory-saving paths in tests/test_architectures/test_tabpfn_v3.py (lines 33-48), ensuring predictions remain identical regardless of chunking strategy.

Configuration Options and Code Examples

The memory_saving_mode parameter accepts booleans, the string "auto", or numeric values representing memory fractions.

Automatic OOM Prevention

Use "auto" to let TabPFN detect when a batch would exceed memory caps and automatically switch to chunked evaluation:

from tabpfn import TabPFNRegressor

reg = TabPFNRegressor.create_default_for_version(
    fit_mode="low_memory",
    memory_saving_mode="auto",
    random_state=0,
)

reg.fit(X_train, y_train)
y_pred = reg.predict(X_test)

Explicit Memory Budget Control

Pass a float between 0 and 1 to target a specific fraction of available memory, or use True to force chunked mode:


# Target 50% of available GPU/CPU memory

reg = TabPFNRegressor(
    memory_saving_mode=0.5,
    random_state=0,
)

# Force chunked evaluation regardless of memory detection

reg = TabPFNRegressor(
    memory_saving_mode=True,
    random_state=0,
)

Classifier Implementation

The identical API works for classification tasks:

from tabpfn import TabPFNClassifier

clf = TabPFNClassifier(
    fit_mode="low_memory",
    memory_saving_mode=True,
    random_state=42,
)

clf.fit(X_train, y_train)
pred = clf.predict(X_test)

Advanced Performance Options

You can also control memory via PerformanceOptions with the save_peak_memory_factor argument, which specifies how aggressively to reduce peak memory (higher values = more chunking):

from tabpfn import PerformanceOptions

perf_opts = PerformanceOptions(save_peak_memory_factor=4)
y_pred = reg.predict(X_test, performance_options=perf_opts)

Performance Trade-offs and Verification

Enabling memory_saving_mode introduces a modest runtime overhead proportional to the number of chunks required. However, the implementation ensures numerical equivalence—tests in tests/test_regressor_interface.py and tests/test_classifier_interface.py confirm that fit_mode="low_memory" combined with memory saving produces identical predictions to standard inference paths.

Summary

  • memory_saving_mode is defined in src/tabpfn/architectures/base/memory.py as Union[bool, Literal["auto"], float, int].
  • Enable it when processing large test sets, working with limited GPU/CPU memory, or using fit_mode="low_memory".
  • The flag propagates through src/tabpfn/inference.py as save_peak_mem, triggering chunked matrix operations.
  • Set to "auto" for automatic detection, True to force chunking, or a float (e.g., 0.5) to target specific memory fractions.
  • Predictions remain numerically identical regardless of chunking strategy, as verified by the test suite.

Frequently Asked Questions

What value should I set for memory_saving_mode if I have 8GB GPU memory?

Use memory_saving_mode="auto" to let TabPFN dynamically detect memory pressure, or set memory_saving_mode=0.7 to explicitly target 70% of your available VRAM as a safety buffer. If you still encounter OOM errors, switch to memory_saving_mode=True to force maximum chunking.

Does enabling memory_saving_mode change my prediction results?

No. The implementation in src/tabpfn/inference.py ensures chunked operations produce mathematically identical results to standard batched operations. The test suite in tests/test_architectures/test_tabpfn_v3.py specifically validates output equivalence between memory-saving and standard modes.

Can I use memory_saving_mode with TabPFNClassifier?

Yes. Both TabPFNRegressor and TabPFNClassifier support the parameter in their constructors, as implemented in src/tabpfn/regressor.py and src/tabpfn/classifier.py respectively. The underlying inference engine handles chunking identically for both regression and classification tasks.

How does memory_saving_mode interact with fit_mode="low_memory"?

These settings operate at different stages. fit_mode="low_memory" reduces memory during preprocessing and model fitting, while memory_saving_mode activates during the prediction phase in src/tabpfn/inference.py. Using them together provides comprehensive memory protection throughout the entire TabPFN workflow, preventing OOM errors during both fitting and inference on constrained hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →