TabPFN inference_precision Settings: How Auto, Autocast, Float32, and Float64 Affect Performance
The inference_precision parameter controls whether TabPFN uses mixed-precision autocast (FP16), single-precision (FP32), or double-precision (FP64) tensors during inference, directly determining computation speed, memory usage, and numerical reproducibility.
TabPFN is a neural-network-based predictor for tabular data developed by PriorLabs. The inference_precision argument dictates the numeric datatype for model weights, inputs, and intermediate activations during the forward pass, with significant implications for hardware acceleration and result stability.
Understanding the Four inference_precision Modes
TabPFN accepts four distinct settings for inference_precision: "auto", "autocast", torch.float32, and torch.float64. Each option triggers specific behavior in the underlying PyTorch inference engine.
"auto" Mode (Default)
When inference_precision="auto" is selected, TabPFN invokes the determine_precision function in tabpfn/base.py to inspect the available devices. If the hardware supports FP16 tensor operations (such as NVIDIA GPUs with tensor cores), the library enables autocast automatically; otherwise, it falls back to standard float32 precision.
This mode offers a balance between performance and compatibility. On supported GPUs, you receive the speed benefits of mixed precision without manual configuration. On CPUs or older GPUs, the system defaults to safe single-precision arithmetic.
"autocast" Mode
Setting inference_precision="autocast" forces PyTorch’s mixed-precision autocast mode (torch.autocast) regardless of automatic heuristics. This mode executes many matrix operations in FP16 (half-precision) while maintaining select operations in FP32 for numerical stability.
Performance characteristics:
- Speed: Fastest option on modern CUDA devices (up to 2× speed-up on RTX/A100 series)
- Memory: Lowest footprint (2 bytes per element via
AUTOCAST_DTYPE_BYTE_SIZEintabpfn/utils.py) - Reproducibility: May introduce small nondeterministic rounding errors (≈1e-3 relative error)
torch.float32 (Single Precision)
Explicitly setting inference_precision=torch.float32 disables autocast and forces all tensors to use single-precision floating-point. This mode provides deterministic, reproducible results across different runs and hardware platforms.
While slower than autocast on tensor-core-equipped GPUs, float32 ensures consistent predictions and wider hardware compatibility. It requires 4 bytes per element, doubling the memory consumption compared to autocast mode.
torch.float64 (Double Precision)
The torch.float64 setting forces double-precision computation throughout the inference pipeline. This provides the highest numerical stability for datasets with extremely large value ranges or when debugging precision-sensitive calculations.
However, this mode incurs significant penalties:
- Speed: Slowest performance, as most consumer GPUs lack FP64 hardware acceleration and fall back to CPU computation
- Memory: Highest consumption (8 bytes per element)
- Compatibility: Raises a
TabPFNValidationErroron Apple Silicon MPS devices, as implemented indetermine_precisionwithintabpfn/base.py
Core Implementation Logic
The precision handling logic resides in tabpfn/base.py within the determine_precision function:
def determine_precision(
inference_precision: torch.dtype | Literal["autocast", "auto"],
devices_: Sequence[torch.device]
) -> tuple[bool, torch.dtype | None, int]:
...
This function returns three critical values:
use_autocast_: Boolean flag enabling PyTorch autocastforced_inference_dtype_: Concretetorch.dtypewhen specified by the userbyte_size: Integer representing bytes per element (2 for autocast, 4 for float32, 8 for float64)
When processing "auto" or "autocast", the function calls infer_fp16_inference_mode from tabpfn/utils.py to validate FP16 support on the target devices. If validation fails (e.g., on CPU-only systems), the library gracefully degrades to float32.
Performance Comparison by Setting
| Setting | Speed Impact | Memory Usage | Reproducibility |
|---|---|---|---|
| "auto" | Fast on tensor-core GPUs; standard otherwise | Low on GPUs (autocast), normal otherwise | May vary slightly if autocast activates |
| "autocast" | Fastest (1.5–2× speed-up on modern CUDA) | Lowest (half-precision tensors) | Slight precision loss (≈1e-4 to 1e-3 error) |
| float32 | Standard; slower than autocast on GPUs | Moderate (4 bytes/element) | Fully deterministic and reproducible |
| float64 | Slowest (CPU fallback on most GPUs) | Highest (8 bytes/element) | Maximum numerical stability |
The default precision configuration is defined in tabpfn/settings.py, which sets "auto" as the global default for all TabPFN estimators.
Practical Usage Examples
The following code demonstrates how to instantiate TabPFN regressors with each precision mode:
import time
import torch
from tabpfn import TabPFNRegressor
from sklearn.datasets import fetch_california_housing
X, y = fetch_california_housing(return_X_y=True, as_frame=False)
# 1. Auto mode - uses autocast on CUDA, float32 otherwise
reg_auto = TabPFNRegressor(inference_precision="auto")
t0 = time.time()
reg_auto.fit(X, y)
print(f"Auto fit time: {time.time() - t0:.2f}s")
# 2. Explicit autocast - forces mixed precision
reg_autocast = TabPFNRegressor(inference_precision="autocast")
t0 = time.time()
reg_autocast.fit(X, y)
print(f"Autocast fit time: {time.time() - t0:.2f}s")
# 3. Float32 - deterministic single precision
reg_f32 = TabPFNRegressor(inference_precision=torch.float32)
t0 = time.time()
reg_f32.fit(X, y)
print(f"Float32 fit time: {time.time() - t0:.2f}s")
# 4. Float64 - double precision (high memory, slow)
reg_f64 = TabPFNRegressor(inference_precision=torch.float64)
t0 = time.time()
reg_f64.fit(X, y)
print(f"Float64 fit time: {time.time() - t0:.2f}s")
Summary
- Use
"auto"for general-purpose inference that automatically leverages GPU acceleration when available. - Use
"autocast"when maximizing throughput on modern NVIDIA GPUs and slight numerical variation is acceptable. - Use
torch.float32when you require deterministic, reproducible predictions across different hardware or runs. - Use
torch.float64only when debugging numerical instability or processing data with extreme dynamic ranges, keeping in mind the performance penalty and MPS incompatibility. - The precision logic is centralized in
tabpfn/base.pyand validated through unit tests intests/test_regressor_interface.pyandtests/test_classifier_interface.py.
Frequently Asked Questions
What is the default inference_precision setting in TabPFN?
The default value is "auto", defined in tabpfn/settings.py. This setting automatically enables autocast on CUDA devices with FP16 support while falling back to float32 on CPUs or incompatible hardware.
Does autocast mode reduce prediction accuracy?
Autocast introduces minor numerical differences due to FP16 rounding, typically resulting in relative errors around 1e-3 or smaller. For most tabular prediction tasks, this difference is negligible, but for applications requiring strict reproducibility, use torch.float32 instead.
Can I use float64 precision on Apple Silicon Macs?
No. Attempting to set inference_precision=torch.float64 on MPS devices raises a TabPFNValidationError according to the validation logic in tabpfn/base.py. Double precision is not supported on Apple Silicon hardware.
How do I verify which precision mode is actually being used?
Check the estimator's use_autocast_ and forced_inference_dtype_ attributes after fitting. These attributes are set by the determine_precision function and indicate whether autocast is active or a specific dtype is enforced.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →