# Feature Subsampling in TabPFN: How FeatureSubsamplingMethod Controls Ensemble Diversity

> Explore feature subsampling in TabPFN and how FEATURE_SUBSAMPLING_METHOD controls ensemble diversity through random, balanced, or Gini importance feature selection for improved model performance.

- Repository: [Prior Labs/TabPFN](https://github.com/PriorLabs/TabPFN)
- Tags: internals
- Published: 2026-05-06

---

**Feature subsampling in TabPFN limits the number of input features per ensemble estimator to `max_features_per_estimator`, using the `FeatureSubsamplingMethod` enum to determine whether features are selected randomly, in balanced round-robin fashion, or via Gini importance rankings.**

TabPFN trains an ensemble of small neural-network estimators to deliver high performance on tabular data. When datasets contain more features than the per-estimator budget allows, the library automatically **subsamples** the feature set for each model. This ensures individual estimators stay within computational limits while the ensemble as a whole maintains a diverse view of the data. The `FeatureSubsamplingMethod` configuration—defined in [`src/tabpfn/preprocessing/configs.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/configs.py)—governs exactly how these feature subsets are selected across the ensemble.

## Why TabPFN Uses Feature Subsampling

TabPFN builds hundreds of small transformer-based estimators that vote together to produce final predictions. Each estimator has a fixed **feature budget** controlled by the `max_features_per_estimator` parameter. When your dataset exceeds this budget, the preprocessing pipeline in [`src/tabpfn/preprocessing/ensemble.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/ensemble.py) must select a subset of features for each estimator.

Subsampling serves two purposes. It prevents any single model from being overwhelmed by high-dimensional input, and it injects diversity into the ensemble by exposing different estimators to different feature views. The `FeatureSubsamplingMethod` enum defines the strategy for this selection, ranging from simple random sampling to importance-based selection using LightGBM Gini gain.

## The Five FeatureSubsamplingMethod Strategies

The enumeration offers five concrete strategies for feature selection. Each strategy implements a different trade-off between randomness, computational cost, and preservation of important predictive signals.

### BALANCED (Round-Robin Pool Sampling)

**BALANCED** implements a globally shuffled feature pool with round-robin sampling. The algorithm ensures each feature appears approximately the same number of times across the entire ensemble, guaranteeing equidistant coverage of the feature space. This method is the default fallback when `AUTO` resolution determines that importance-based sampling is unnecessary.

### RANDOM (Independent Random Draws)

**RANDOM** draws an independent random subset of size `max_features_per_estimator` for each estimator without reconciliation across the ensemble. While simple to compute, this approach does not guarantee uniform feature coverage and may over- or under-represent specific features across the ensemble.

### CONSTANT_AND_BALANCED (Fixed Leading Features)

**CONSTANT_AND_BALANCED** always retains a fixed number of leading features—controlled by `constant_feature_count` (default 50)—and fills the remaining budget with balanced sampling from the remaining features. This method is useful when you know that the first N features contain essential metadata or identifiers that every estimator should see.

### GINI_FEATURE_IMPORTANCE (Importance-Based Selection)

**GINI_FEATURE_IMPORTANCE** uses LightGBM to compute Gini gain importance scores on the training data. It keeps the top-K most important features—controlled by `importance_top_k_count`—and fills the rest of the budget with balanced sampling of the remaining features. This method requires sufficient data for reliable importance estimation and is only viable when the sample size exceeds `AUTO_FEATURE_SUBSAMPLING_IMPORTANCE_MIN_SAMPLES`.

### AUTO (Automatic Resolution)

**AUTO** is the default configuration that automatically selects between **GINI_FEATURE_IMPORTANCE** and **BALANCED** based on dataset characteristics. When the dataset is large enough for reliable importance scoring and subsampling is actually required, it selects the importance-based method. Otherwise, it falls back to balanced sampling to avoid unstable importance estimates on small datasets.

## How AUTO Resolution Works

The resolution logic lives in the helper `_resolve_feature_subsampling_method` within [`src/tabpfn/preprocessing/ensemble.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/ensemble.py) (lines 1043–1060). The implementation evaluates the following conditions:

```python
if method is not FeatureSubsamplingMethod.AUTO:
    return method
if needs_subsampling and n_samples > auto_min_samples:
    return FeatureSubsamplingMethod.GINI_FEATURE_IMPORTANCE
return FeatureSubsamplingMethod.BALANCED

```

If the user specifies any concrete method other than `AUTO`, the function returns that method immediately. When `AUTO` is active, the logic checks whether subsampling is necessary (`needs_subsampling`) and whether the dataset contains enough samples (`n_samples > auto_min_samples`) to compute reliable Gini importance. If both conditions are met, it returns `GINI_FEATURE_IMPORTANCE`; otherwise, it selects `BALANCED`.

## Index Generation Implementation

The actual feature index generation occurs in `_get_subsample_feature_indices`, which delegates to specialized helper functions based on the resolved method:

- **`_subsample_features_balanced`** – Implements the round-robin pool sampling for **BALANCED**.
- **`_subsample_features_random`** – Implements independent random draws for **RANDOM**.
- **`_subsample_features_constant_and_balanced`** – Keeps constant leading features then applies balanced sampling for **CONSTANT_AND_BALANCED**.
- **`_subsample_features_importance_based`** – Computes LightGBM importance and combines top-K features with balanced sampling for **GINI_FEATURE_IMPORTANCE**.

All helper functions respect the `max_features_per_estimator` budget and return `None` when the dataset already fits within the budget (i.e., no subsampling is required).

## Configuring Feature Subsampling in Practice

You control feature subsampling through the `TabPFNClassifier` or `TabPFNRegressor` interface. Import `FeatureSubsamplingMethod` from [`src/tabpfn/preprocessing/configs.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/configs.py) to specify your strategy:

```python
from tabpfn.preprocessing.configs import FeatureSubsamplingMethod
from tabpfn.classifier import TabPFNClassifier

# Use default AUTO behavior (balanced for small data, importance-based for large)

clf = TabPFNClassifier(max_features_per_estimator=300)

# Force random feature subsampling

clf = TabPFNClassifier(
    max_features_per_estimator=300,
    feature_subsampling_method=FeatureSubsamplingMethod.RANDOM,
)

# Keep the first 50 features always, balance the remainder

clf = TabPFNClassifier(
    max_features_per_estimator=300,
    feature_subsampling_method=FeatureSubsamplingMethod.CONSTANT_AND_BALANCED,
    constant_feature_count=50,
)

# Use Gini-based importance (requires sufficient training data)

clf = TabPFNClassifier(
    max_features_per_estimator=300,
    feature_subsampling_method=FeatureSubsamplingMethod.GINI_FEATURE_IMPORTANCE,
    importance_top_k_count=150,  # keep top-150 important features per estimator

)

```

For regression tasks, identical parameters apply to `TabPFNRegressor` in [`src/tabpfn/regressor.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/regressor.py).

## Summary

- **Feature subsampling** ensures each TabPFN estimator respects the `max_features_per_estimator` budget while maintaining ensemble diversity.
- **`FeatureSubsamplingMethod`** defines five strategies: **BALANCED**, **RANDOM**, **CONSTANT_AND_BALANCED**, **GINI_FEATURE_IMPORTANCE**, and **AUTO**.
- **AUTO resolution** automatically selects between Gini importance and balanced sampling based on dataset size and subsampling necessity.
- Implementation resides in [`src/tabpfn/preprocessing/ensemble.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/ensemble.py), with configuration defined in [`src/tabpfn/preprocessing/configs.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/configs.py).
- Configure subsampling via classifier/regressor parameters, using `constant_feature_count` or `importance_top_k_count` for method-specific tuning.

## Frequently Asked Questions

### What happens if my dataset has fewer features than max_features_per_estimator?

When the dataset dimensionality is below the budget, the subsampling logic returns `None` and each estimator trains on the full feature set. No subsampling occurs, and the `FeatureSubsamplingMethod` has no practical effect on that specific fit.

### How does CONSTANT_AND_BALANCED differ from GINI_FEATURE_IMPORTANCE?

**CONSTANT_AND_BALANCED** mechanically keeps the first N features (indices 0 to `constant_feature_count-1`) regardless of their predictive value, making it suitable when feature order carries semantic meaning. **GINI_FEATURE_IMPORTANCE** dynamically selects features based on LightGBM Gini gain scores calculated from the training data, prioritizing predictive power over position.

### When does AUTO select GINI_FEATURE_IMPORTANCE over BALANCED?

The `AUTO` resolver selects **GINI_FEATURE_IMPORTANCE** only when two conditions are met: the dataset actually requires subsampling (`needs_subsampling` is True), and the number of samples exceeds `AUTO_FEATURE_SUBSAMPLING_IMPORTANCE_MIN_SAMPLES`. If either condition fails, it defaults to **BALANCED** to avoid unstable importance estimates.

### Can I use feature subsampling with TabPFNRegressor?

Yes. Both `TabPFNClassifier` and `TabPFNRegressor` in [`src/tabpfn/classifier.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/classifier.py) and [`src/tabpfn/regressor.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/regressor.py) respectively accept the `feature_subsampling_method` parameter and pass it to the shared preprocessing pipeline in [`src/tabpfn/preprocessing/ensemble.py`](https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/preprocessing/ensemble.py). The behavior is identical for regression and classification tasks.