# TabPFN Fit Pipeline Preprocessing Transformations: Complete Technical Guide

> Explore TabPFN's fit pipeline preprocessing. Learn about quantile scaling, SVD, categorical encoding, fingerprint addition, and target transformation for optimal model performance.

- Repository: [Prior Labs/TabPFN](https://github.com/PriorLabs/TabPFN)
- Tags: technical-guide
- Published: 2026-05-06

---

**TabPFN applies a multi-stage preprocessing pipeline during fit that includes quantile scaling, SVD feature generation, categorical encoding, fingerprint addition, and target transformation, orchestrated through PreprocessorConfig objects defined in the PriorLabs/TabPFN repository.**

When you call `fit()` on a `TabPFNClassifier` or `TabPFNRegressor`, your data flows through a sophisticated preprocessing architecture before reaching the neural network. This pipeline handles data type coercion, missing value treatment, feature engineering, and distribution transformations, with distinct configurations for classification versus regression tasks.

## How the Preprocessing Pipeline is Structured

The preprocessing logic is defined by two configuration sources in [`src/tabpfn/inference_config.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference_config.py):

1. **`PREPROCESS_TRANSFORMS`** – A list of `PreprocessorConfig` objects that determine how input features (X) are transformed
2. **`REGRESSION_Y_PREPROCESS_TRANSFORMS`** – A tuple defining how target variables (y) are processed, specifically for regression tasks

The actual execution is orchestrated by the `TabPFNEnsemblePreprocessor` class, which parallelizes these transformations across ensemble members via `fit_preprocessing` in [`src/tabpfn/preprocessing/transform.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/transform.py) (lines 113-134). The pipeline assembly occurs in `TabPFNRegressor._initialize_dataset_preprocessing` ([`src/tabpfn/regressor.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/regressor.py), lines 630-642) for regression models, with an equivalent entry point in [`src/tabpfn/classifier.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/classifier.py) for classification.

## Core Preprocessing Transformations

### Data Cleaning and Modality Detection

Before any feature engineering occurs, the pipeline executes fundamental data cleaning:

- **`clean_data`** – A utility function that casts dtypes, repairs text-NA columns, and returns an ordinal encoder for categorical columns
- **`detect_feature_modalities`** – Heuristically determines which columns are categorical versus numerical based on cardinality and data types

These steps ensure downstream transformers receive properly typed data and correct modality classifications.

### Quantile Scaling and Distribution Normalization

Numerical features undergo uniform quantile transformation to standardize their distributions:

- **`quantile_uni`** or **`quantile_uni_coarse`** – Applied per-feature as specified in `PreprocessorConfig`
- Configured in [`src/tabpfn/preprocessing/presets.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/presets.py) where `v2_regressor_preprocessor_configs()` defines the default scaling strategy
- The `append_original=True` parameter retains original features alongside transformed versions

### SVD Feature Generation

Global structure is captured through low-rank projections:

- **`svd`** or **`svd_quarter_components`** – Adds singular value decomposition components as supplementary features
- Implemented in [`src/tabpfn/preprocessing/steps/add_svd_features_step.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/steps/add_svd_features_step.py)
- Acts as a global transformer across all features to capture inter-feature correlations

### Categorical Encoding Strategies

Categorical variables are encoded using different strategies across ensemble members:

- **`ordinal_very_common_categories_shuffled`** – Ordinal encoding with randomized category ordering (used in the first preprocessor configuration)
- **`onehot`** – Standard one-hot encoding (typically used in the second preprocessor configuration)
- **`max_onehot_cardinality`** – Controls cardinality limits for one-hot encoding to prevent dimensionality explosion
- Implementation resides in [`src/tabpfn/preprocessing/steps/encode_categorical_features_step.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/steps/encode_categorical_features_step.py)

### Fingerprint and Auxiliary Features

Additional engineered features improve model robustness:

- **`add_fingerprint_features_step`** – Adds a hash-based fingerprint column derived from row values to help the model distinguish duplicate or near-duplicate rows
- **Polynomial features** – Optionally added when `POLYNOMIAL_FEATURES` is enabled in `InferenceConfig`
- **Feature shuffling** – Applied via `shuffle_features_step` when `FEATURE_SHIFT_METHOD` is set to `"shuffle"`

### Target Transformation for Regression

Regression tasks specifically transform the target variable to handle skewed distributions:

- **`safepower`** transformation – Applied via `reshape_feature_distribution_step` in [`src/tabpfn/preprocessing/steps/reshape_feature_distribution_step.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/steps/reshape_feature_distribution_step.py)
- Default configuration uses `REGRESSION_Y_PREPROCESS_TRANSFORMS: tuple[str | None, ...] = (None, "safepower")` in `InferenceConfig`, where the first ensemble member sees raw targets and the second receives power-transformed targets

### Specialized Transformations

Additional steps controlled by `InferenceConfig` flags include:

- **Outlier removal** – Activated when `OUTLIER_REMOVAL_STD` is not `None`, performed within cleaning and feature addition steps
- **Differentiable Z-normalization** – Implemented in `differentiable_z_norm_step`, though currently only supported when `differentiable_input=True` (not enabled for regressors)

## Default V2 Regression Configuration

The default preprocessing for TabPFN regression models (version 2) is defined in [`src/tabpfn/preprocessing/presets.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/presets.py):

```python
def v2_regressor_preprocessor_configs() -> list[PreprocessorConfig]:
    return [
        PreprocessorConfig(
            "quantile_uni",                # quantile scaling

            append_original=True,
            categorical_name="ordinal_very_common_categories_shuffled",
            global_transformer_name="svd",
        ),
        PreprocessorConfig("safepower", categorical_name="onehot"),
    ]

```

This configuration creates two distinct preprocessing branches:

1. **First preprocessor** – Applies quantile scaling with SVD features and shuffled ordinal encoding for categoricals, appending results to original features
2. **Second preprocessor** – Applies safe-power transformation with one-hot encoding for categoricals

## Inspecting the Preprocessing Pipeline

### View Default Configuration

```python
from tabpfn import TabPFNRegressor
from tabpfn.preprocessing import v2_regressor_preprocessor_configs

# Display the two default preprocessor configs for v2 regression

for cfg in v2_regressor_preprocessor_configs():
    print(cfg)

```

This outputs the `PreprocessorConfig` objects showing `quantile_uni` with `ordinal_very_common_categories_shuffled` and `svd` transformer, followed by the `safepower` config with `onehot` encoding.

### Access Trained Preprocessing

```python
import sklearn.datasets
from tabpfn import TabPFNRegressor

X, y = sklearn.datasets.make_regression(n_samples=200, n_features=20, random_state=0)
model = TabPFNRegressor()
model.fit(X, y)

# Examine ensemble preprocessing configurations

for i, cfg in enumerate(model.ensemble_configs_[:2]):
    print(f"Ensemble member {i}:")
    print("  preprocessor:", cfg.preprocess_config.name)
    print("  categorical encoding:", cfg.preprocess_config.categorical_name)
    print("  global transformer:", cfg.preprocess_config.global_transformer_name)

```

### Manual Preprocessing Application

```python
from tabpfn.preprocessing import TabPFNEnsemblePreprocessor
from tabpfn.preprocessing import v2_regressor_preprocessor_configs
from tabpfn.inference_config import InferenceConfig
from tabpfn.constants import ModelVersion, TaskType

# Build inference configuration matching defaults

ic = InferenceConfig.get_default(task_type="regression", model_version=ModelVersion.V2)

# Initialize preprocessor without GPU acceleration

preproc = TabPFNEnsemblePreprocessor(
    n_preprocessing_jobs=1,
    enable_gpu_preprocessing=False,
    preprocessor_configs=ic.PREPROCESS_TRANSFORMS,
    target_transforms=[None, "safepower"]
)

# Fit and transform

preproc.fit(X, y)
X_preprocessed, y_preprocessed = preproc.transform(X, y)

```

## Summary

- **TabPFN** uses `PreprocessorConfig` objects defined in [`src/tabpfn/inference_config.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference_config.py) and [`src/tabpfn/preprocessing/presets.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/presets.py) to specify transformation sequences
- **Feature preprocessing** includes cleaning, modality detection, quantile scaling, SVD feature generation, categorical encoding (ordinal with shuffling or one-hot), and fingerprint addition
- **Regression targets** undergo safe-power transformation by default via `REGRESSION_Y_PREPROCESS_TRANSFORMS` set to `(None, "safepower")`
- **Parallel execution** occurs across ensemble members through `TabPFNEnsemblePreprocessor` in [`src/tabpfn/preprocessing/transform.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/transform.py)
- **Configuration assembly** happens in `TabPFNRegressor._initialize_dataset_preprocessing` ([`src/tabpfn/regressor.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/regressor.py), lines 630-642)

## Frequently Asked Questions

### What preprocessing does TabPFN apply to categorical features?

TabPFN uses two encoding strategies across its ensemble: **`ordinal_very_common_categories_shuffled`** for the first preprocessor (which randomizes category order to improve robustness) and **`onehot`** for the second preprocessor. These are implemented in [`src/tabpfn/preprocessing/steps/encode_categorical_features_step.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/steps/encode_categorical_features_step.py). The choice between these is controlled by the `categorical_name` parameter in each `PreprocessorConfig`.

### How does TabPFN handle target variable transformation in regression?

For regression tasks, TabPFN applies a **`safepower`** transformation to the target variable in the second ensemble member, while keeping the first member untransformed (`None`). This configuration, defined as `REGRESSION_Y_PREPROCESS_TRANSFORMS: tuple[str | None, ...] = (None, "safepower")` in [`src/tabpfn/inference_config.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/inference_config.py), helps the model handle skewed target distributions through the `reshape_feature_distribution_step` implementation.

### Where is the preprocessing pipeline assembled in the source code?

The preprocessing pipeline is assembled in **`TabPFNRegressor._initialize_dataset_preprocessing`** at lines 630-642 of [`src/tabpfn/regressor.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/regressor.py) (with equivalent logic in [`src/tabpfn/classifier.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/classifier.py) for classification). This method instantiates the `TabPFNEnsemblePreprocessor`, which then executes `fit_preprocessing` from [`src/tabpfn/preprocessing/transform.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/transform.py) to parallelize transformations across ensemble members.

### Can preprocessing steps be customized or disabled?

Yes, you can customize preprocessing by modifying `InferenceConfig.PREPROCESS_TRANSFORMS` before model initialization or by passing custom `PreprocessorConfig` objects. However, the default configurations in `v2_regressor_preprocessor_configs()` are highly optimized for heterogeneous tabular data. Disabling critical steps like `quantile_uni` or SVD generation is not recommended as these transformations align the data distribution with what the pre-trained neural network expects.