# TabPFN inference_precision Settings: How Auto, Autocast, Float32, and Float64 Affect Performance

> Explore TabPFN inference_precision settings. Learn how Auto, Autocast, Float32, and Float64 impact speed, memory, and reproducibility for optimal performance.

- Repository: [Prior Labs/TabPFN](https://github.com/PriorLabs/TabPFN)
- Tags: performance
- Published: 2026-05-06

---

**The `inference_precision` parameter controls whether TabPFN uses mixed-precision autocast (FP16), single-precision (FP32), or double-precision (FP64) tensors during inference, directly determining computation speed, memory usage, and numerical reproducibility.**

TabPFN is a neural-network-based predictor for tabular data developed by PriorLabs. The `inference_precision` argument dictates the numeric datatype for model weights, inputs, and intermediate activations during the forward pass, with significant implications for hardware acceleration and result stability.

## Understanding the Four inference_precision Modes

TabPFN accepts four distinct settings for `inference_precision`: `"auto"`, `"autocast"`, `torch.float32`, and `torch.float64`. Each option triggers specific behavior in the underlying PyTorch inference engine.

### "auto" Mode (Default)

When `inference_precision="auto"` is selected, TabPFN invokes the `determine_precision` function in [`tabpfn/base.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/base.py) to inspect the available devices. If the hardware supports FP16 tensor operations (such as NVIDIA GPUs with tensor cores), the library enables **autocast** automatically; otherwise, it falls back to standard `float32` precision.

This mode offers a balance between performance and compatibility. On supported GPUs, you receive the speed benefits of mixed precision without manual configuration. On CPUs or older GPUs, the system defaults to safe single-precision arithmetic.

### "autocast" Mode

Setting `inference_precision="autocast"` forces PyTorch’s mixed-precision autocast mode (`torch.autocast`) regardless of automatic heuristics. This mode executes many matrix operations in **FP16** (half-precision) while maintaining select operations in FP32 for numerical stability.

**Performance characteristics:**
- **Speed**: Fastest option on modern CUDA devices (up to 2× speed-up on RTX/A100 series)
- **Memory**: Lowest footprint (2 bytes per element via `AUTOCAST_DTYPE_BYTE_SIZE` in [`tabpfn/utils.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/utils.py))
- **Reproducibility**: May introduce small nondeterministic rounding errors (≈1e-3 relative error)

### torch.float32 (Single Precision)

Explicitly setting `inference_precision=torch.float32` disables autocast and forces all tensors to use **single-precision floating-point**. This mode provides deterministic, reproducible results across different runs and hardware platforms.

While slower than autocast on tensor-core-equipped GPUs, float32 ensures consistent predictions and wider hardware compatibility. It requires 4 bytes per element, doubling the memory consumption compared to autocast mode.

### torch.float64 (Double Precision)

The `torch.float64` setting forces **double-precision** computation throughout the inference pipeline. This provides the highest numerical stability for datasets with extremely large value ranges or when debugging precision-sensitive calculations.

However, this mode incurs significant penalties:
- **Speed**: Slowest performance, as most consumer GPUs lack FP64 hardware acceleration and fall back to CPU computation
- **Memory**: Highest consumption (8 bytes per element)
- **Compatibility**: Raises a `TabPFNValidationError` on Apple Silicon MPS devices, as implemented in `determine_precision` within [`tabpfn/base.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/base.py)

## Core Implementation Logic

The precision handling logic resides in [`tabpfn/base.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/base.py) within the `determine_precision` function:

```python
def determine_precision(
    inference_precision: torch.dtype | Literal["autocast", "auto"],
    devices_: Sequence[torch.device]
) -> tuple[bool, torch.dtype | None, int]:
    ...

```

This function returns three critical values:
- `use_autocast_`: Boolean flag enabling PyTorch autocast
- `forced_inference_dtype_`: Concrete `torch.dtype` when specified by the user
- `byte_size`: Integer representing bytes per element (2 for autocast, 4 for float32, 8 for float64)

When processing `"auto"` or `"autocast"`, the function calls `infer_fp16_inference_mode` from [`tabpfn/utils.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/utils.py) to validate FP16 support on the target devices. If validation fails (e.g., on CPU-only systems), the library gracefully degrades to float32.

## Performance Comparison by Setting

| Setting | Speed Impact | Memory Usage | Reproducibility |
|---------|-------------|--------------|-----------------|
| **"auto"** | Fast on tensor-core GPUs; standard otherwise | Low on GPUs (autocast), normal otherwise | May vary slightly if autocast activates |
| **"autocast"** | Fastest (1.5–2× speed-up on modern CUDA) | Lowest (half-precision tensors) | Slight precision loss (≈1e-4 to 1e-3 error) |
| **float32** | Standard; slower than autocast on GPUs | Moderate (4 bytes/element) | Fully deterministic and reproducible |
| **float64** | Slowest (CPU fallback on most GPUs) | Highest (8 bytes/element) | Maximum numerical stability |

The default precision configuration is defined in [`tabpfn/settings.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/settings.py), which sets `"auto"` as the global default for all TabPFN estimators.

## Practical Usage Examples

The following code demonstrates how to instantiate TabPFN regressors with each precision mode:

```python
import time
import torch
from tabpfn import TabPFNRegressor
from sklearn.datasets import fetch_california_housing

X, y = fetch_california_housing(return_X_y=True, as_frame=False)

# 1. Auto mode - uses autocast on CUDA, float32 otherwise

reg_auto = TabPFNRegressor(inference_precision="auto")
t0 = time.time()
reg_auto.fit(X, y)
print(f"Auto fit time: {time.time() - t0:.2f}s")

# 2. Explicit autocast - forces mixed precision

reg_autocast = TabPFNRegressor(inference_precision="autocast")
t0 = time.time()
reg_autocast.fit(X, y)
print(f"Autocast fit time: {time.time() - t0:.2f}s")

# 3. Float32 - deterministic single precision

reg_f32 = TabPFNRegressor(inference_precision=torch.float32)
t0 = time.time()
reg_f32.fit(X, y)
print(f"Float32 fit time: {time.time() - t0:.2f}s")

# 4. Float64 - double precision (high memory, slow)

reg_f64 = TabPFNRegressor(inference_precision=torch.float64)
t0 = time.time()
reg_f64.fit(X, y)
print(f"Float64 fit time: {time.time() - t0:.2f}s")

```

## Summary

- **Use `"auto"`** for general-purpose inference that automatically leverages GPU acceleration when available.
- **Use `"autocast"`** when maximizing throughput on modern NVIDIA GPUs and slight numerical variation is acceptable.
- **Use `torch.float32`** when you require deterministic, reproducible predictions across different hardware or runs.
- **Use `torch.float64`** only when debugging numerical instability or processing data with extreme dynamic ranges, keeping in mind the performance penalty and MPS incompatibility.
- The precision logic is centralized in [`tabpfn/base.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/base.py) and validated through unit tests in [`tests/test_regressor_interface.py`](https://github.com/PriorLabs/TabPFN/blob/main/tests/test_regressor_interface.py) and [`tests/test_classifier_interface.py`](https://github.com/PriorLabs/TabPFN/blob/main/tests/test_classifier_interface.py).

## Frequently Asked Questions

### What is the default inference_precision setting in TabPFN?

The default value is `"auto"`, defined in [`tabpfn/settings.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/settings.py). This setting automatically enables autocast on CUDA devices with FP16 support while falling back to float32 on CPUs or incompatible hardware.

### Does autocast mode reduce prediction accuracy?

Autocast introduces minor numerical differences due to FP16 rounding, typically resulting in relative errors around 1e-3 or smaller. For most tabular prediction tasks, this difference is negligible, but for applications requiring strict reproducibility, use `torch.float32` instead.

### Can I use float64 precision on Apple Silicon Macs?

No. Attempting to set `inference_precision=torch.float64` on MPS devices raises a `TabPFNValidationError` according to the validation logic in [`tabpfn/base.py`](https://github.com/PriorLabs/TabPFN/blob/main/tabpfn/base.py). Double precision is not supported on Apple Silicon hardware.

### How do I verify which precision mode is actually being used?

Check the estimator's `use_autocast_` and `forced_inference_dtype_` attributes after fitting. These attributes are set by the `determine_precision` function and indicate whether autocast is active or a specific dtype is enforced.