TabPFN inference_precision Settings: How Auto, Autocast, Float32, and Float64 Affect Performance

The inference_precision parameter controls whether TabPFN uses mixed-precision autocast (FP16), single-precision (FP32), or double-precision (FP64) tensors during inference, directly determining computation speed, memory usage, and numerical reproducibility.

TabPFN is a neural-network-based predictor for tabular data developed by PriorLabs. The inference_precision argument dictates the numeric datatype for model weights, inputs, and intermediate activations during the forward pass, with significant implications for hardware acceleration and result stability.

Understanding the Four inference_precision Modes

TabPFN accepts four distinct settings for inference_precision: "auto", "autocast", torch.float32, and torch.float64. Each option triggers specific behavior in the underlying PyTorch inference engine.

"auto" Mode (Default)

When inference_precision="auto" is selected, TabPFN invokes the determine_precision function in tabpfn/base.py to inspect the available devices. If the hardware supports FP16 tensor operations (such as NVIDIA GPUs with tensor cores), the library enables autocast automatically; otherwise, it falls back to standard float32 precision.

This mode offers a balance between performance and compatibility. On supported GPUs, you receive the speed benefits of mixed precision without manual configuration. On CPUs or older GPUs, the system defaults to safe single-precision arithmetic.

"autocast" Mode

Setting inference_precision="autocast" forces PyTorch’s mixed-precision autocast mode (torch.autocast) regardless of automatic heuristics. This mode executes many matrix operations in FP16 (half-precision) while maintaining select operations in FP32 for numerical stability.

Performance characteristics:

  • Speed: Fastest option on modern CUDA devices (up to 2× speed-up on RTX/A100 series)
  • Memory: Lowest footprint (2 bytes per element via AUTOCAST_DTYPE_BYTE_SIZE in tabpfn/utils.py)
  • Reproducibility: May introduce small nondeterministic rounding errors (≈1e-3 relative error)

torch.float32 (Single Precision)

Explicitly setting inference_precision=torch.float32 disables autocast and forces all tensors to use single-precision floating-point. This mode provides deterministic, reproducible results across different runs and hardware platforms.

While slower than autocast on tensor-core-equipped GPUs, float32 ensures consistent predictions and wider hardware compatibility. It requires 4 bytes per element, doubling the memory consumption compared to autocast mode.

torch.float64 (Double Precision)

The torch.float64 setting forces double-precision computation throughout the inference pipeline. This provides the highest numerical stability for datasets with extremely large value ranges or when debugging precision-sensitive calculations.

However, this mode incurs significant penalties:

  • Speed: Slowest performance, as most consumer GPUs lack FP64 hardware acceleration and fall back to CPU computation
  • Memory: Highest consumption (8 bytes per element)
  • Compatibility: Raises a TabPFNValidationError on Apple Silicon MPS devices, as implemented in determine_precision within tabpfn/base.py

Core Implementation Logic

The precision handling logic resides in tabpfn/base.py within the determine_precision function:

def determine_precision(
    inference_precision: torch.dtype | Literal["autocast", "auto"],
    devices_: Sequence[torch.device]
) -> tuple[bool, torch.dtype | None, int]:
    ...

This function returns three critical values:

  • use_autocast_: Boolean flag enabling PyTorch autocast
  • forced_inference_dtype_: Concrete torch.dtype when specified by the user
  • byte_size: Integer representing bytes per element (2 for autocast, 4 for float32, 8 for float64)

When processing "auto" or "autocast", the function calls infer_fp16_inference_mode from tabpfn/utils.py to validate FP16 support on the target devices. If validation fails (e.g., on CPU-only systems), the library gracefully degrades to float32.

Performance Comparison by Setting

Setting Speed Impact Memory Usage Reproducibility
"auto" Fast on tensor-core GPUs; standard otherwise Low on GPUs (autocast), normal otherwise May vary slightly if autocast activates
"autocast" Fastest (1.5–2× speed-up on modern CUDA) Lowest (half-precision tensors) Slight precision loss (≈1e-4 to 1e-3 error)
float32 Standard; slower than autocast on GPUs Moderate (4 bytes/element) Fully deterministic and reproducible
float64 Slowest (CPU fallback on most GPUs) Highest (8 bytes/element) Maximum numerical stability

The default precision configuration is defined in tabpfn/settings.py, which sets "auto" as the global default for all TabPFN estimators.

Practical Usage Examples

The following code demonstrates how to instantiate TabPFN regressors with each precision mode:

import time
import torch
from tabpfn import TabPFNRegressor
from sklearn.datasets import fetch_california_housing

X, y = fetch_california_housing(return_X_y=True, as_frame=False)

# 1. Auto mode - uses autocast on CUDA, float32 otherwise

reg_auto = TabPFNRegressor(inference_precision="auto")
t0 = time.time()
reg_auto.fit(X, y)
print(f"Auto fit time: {time.time() - t0:.2f}s")

# 2. Explicit autocast - forces mixed precision

reg_autocast = TabPFNRegressor(inference_precision="autocast")
t0 = time.time()
reg_autocast.fit(X, y)
print(f"Autocast fit time: {time.time() - t0:.2f}s")

# 3. Float32 - deterministic single precision

reg_f32 = TabPFNRegressor(inference_precision=torch.float32)
t0 = time.time()
reg_f32.fit(X, y)
print(f"Float32 fit time: {time.time() - t0:.2f}s")

# 4. Float64 - double precision (high memory, slow)

reg_f64 = TabPFNRegressor(inference_precision=torch.float64)
t0 = time.time()
reg_f64.fit(X, y)
print(f"Float64 fit time: {time.time() - t0:.2f}s")

Summary

  • Use "auto" for general-purpose inference that automatically leverages GPU acceleration when available.
  • Use "autocast" when maximizing throughput on modern NVIDIA GPUs and slight numerical variation is acceptable.
  • Use torch.float32 when you require deterministic, reproducible predictions across different hardware or runs.
  • Use torch.float64 only when debugging numerical instability or processing data with extreme dynamic ranges, keeping in mind the performance penalty and MPS incompatibility.
  • The precision logic is centralized in tabpfn/base.py and validated through unit tests in tests/test_regressor_interface.py and tests/test_classifier_interface.py.

Frequently Asked Questions

What is the default inference_precision setting in TabPFN?

The default value is "auto", defined in tabpfn/settings.py. This setting automatically enables autocast on CUDA devices with FP16 support while falling back to float32 on CPUs or incompatible hardware.

Does autocast mode reduce prediction accuracy?

Autocast introduces minor numerical differences due to FP16 rounding, typically resulting in relative errors around 1e-3 or smaller. For most tabular prediction tasks, this difference is negligible, but for applications requiring strict reproducibility, use torch.float32 instead.

Can I use float64 precision on Apple Silicon Macs?

No. Attempting to set inference_precision=torch.float64 on MPS devices raises a TabPFNValidationError according to the validation logic in tabpfn/base.py. Double precision is not supported on Apple Silicon hardware.

How do I verify which precision mode is actually being used?

Check the estimator's use_autocast_ and forced_inference_dtype_ attributes after fitting. These attributes are set by the determine_precision function and indicate whether autocast is active or a specific dtype is enforced.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →