Feature Subsampling in TabPFN: How FeatureSubsamplingMethod Controls Ensemble Diversity
Feature subsampling in TabPFN limits the number of input features per ensemble estimator to max_features_per_estimator, using the FeatureSubsamplingMethod enum to determine whether features are selected randomly, in balanced round-robin fashion, or via Gini importance rankings.
TabPFN trains an ensemble of small neural-network estimators to deliver high performance on tabular data. When datasets contain more features than the per-estimator budget allows, the library automatically subsamples the feature set for each model. This ensures individual estimators stay within computational limits while the ensemble as a whole maintains a diverse view of the data. The FeatureSubsamplingMethod configuration—defined in src/tabpfn/preprocessing/configs.py—governs exactly how these feature subsets are selected across the ensemble.
Why TabPFN Uses Feature Subsampling
TabPFN builds hundreds of small transformer-based estimators that vote together to produce final predictions. Each estimator has a fixed feature budget controlled by the max_features_per_estimator parameter. When your dataset exceeds this budget, the preprocessing pipeline in src/tabpfn/preprocessing/ensemble.py must select a subset of features for each estimator.
Subsampling serves two purposes. It prevents any single model from being overwhelmed by high-dimensional input, and it injects diversity into the ensemble by exposing different estimators to different feature views. The FeatureSubsamplingMethod enum defines the strategy for this selection, ranging from simple random sampling to importance-based selection using LightGBM Gini gain.
The Five FeatureSubsamplingMethod Strategies
The enumeration offers five concrete strategies for feature selection. Each strategy implements a different trade-off between randomness, computational cost, and preservation of important predictive signals.
BALANCED (Round-Robin Pool Sampling)
BALANCED implements a globally shuffled feature pool with round-robin sampling. The algorithm ensures each feature appears approximately the same number of times across the entire ensemble, guaranteeing equidistant coverage of the feature space. This method is the default fallback when AUTO resolution determines that importance-based sampling is unnecessary.
RANDOM (Independent Random Draws)
RANDOM draws an independent random subset of size max_features_per_estimator for each estimator without reconciliation across the ensemble. While simple to compute, this approach does not guarantee uniform feature coverage and may over- or under-represent specific features across the ensemble.
CONSTANT_AND_BALANCED (Fixed Leading Features)
CONSTANT_AND_BALANCED always retains a fixed number of leading features—controlled by constant_feature_count (default 50)—and fills the remaining budget with balanced sampling from the remaining features. This method is useful when you know that the first N features contain essential metadata or identifiers that every estimator should see.
GINI_FEATURE_IMPORTANCE (Importance-Based Selection)
GINI_FEATURE_IMPORTANCE uses LightGBM to compute Gini gain importance scores on the training data. It keeps the top-K most important features—controlled by importance_top_k_count—and fills the rest of the budget with balanced sampling of the remaining features. This method requires sufficient data for reliable importance estimation and is only viable when the sample size exceeds AUTO_FEATURE_SUBSAMPLING_IMPORTANCE_MIN_SAMPLES.
AUTO (Automatic Resolution)
AUTO is the default configuration that automatically selects between GINI_FEATURE_IMPORTANCE and BALANCED based on dataset characteristics. When the dataset is large enough for reliable importance scoring and subsampling is actually required, it selects the importance-based method. Otherwise, it falls back to balanced sampling to avoid unstable importance estimates on small datasets.
How AUTO Resolution Works
The resolution logic lives in the helper _resolve_feature_subsampling_method within src/tabpfn/preprocessing/ensemble.py (lines 1043–1060). The implementation evaluates the following conditions:
if method is not FeatureSubsamplingMethod.AUTO:
return method
if needs_subsampling and n_samples > auto_min_samples:
return FeatureSubsamplingMethod.GINI_FEATURE_IMPORTANCE
return FeatureSubsamplingMethod.BALANCED
If the user specifies any concrete method other than AUTO, the function returns that method immediately. When AUTO is active, the logic checks whether subsampling is necessary (needs_subsampling) and whether the dataset contains enough samples (n_samples > auto_min_samples) to compute reliable Gini importance. If both conditions are met, it returns GINI_FEATURE_IMPORTANCE; otherwise, it selects BALANCED.
Index Generation Implementation
The actual feature index generation occurs in _get_subsample_feature_indices, which delegates to specialized helper functions based on the resolved method:
_subsample_features_balanced– Implements the round-robin pool sampling for BALANCED._subsample_features_random– Implements independent random draws for RANDOM._subsample_features_constant_and_balanced– Keeps constant leading features then applies balanced sampling for CONSTANT_AND_BALANCED._subsample_features_importance_based– Computes LightGBM importance and combines top-K features with balanced sampling for GINI_FEATURE_IMPORTANCE.
All helper functions respect the max_features_per_estimator budget and return None when the dataset already fits within the budget (i.e., no subsampling is required).
Configuring Feature Subsampling in Practice
You control feature subsampling through the TabPFNClassifier or TabPFNRegressor interface. Import FeatureSubsamplingMethod from src/tabpfn/preprocessing/configs.py to specify your strategy:
from tabpfn.preprocessing.configs import FeatureSubsamplingMethod
from tabpfn.classifier import TabPFNClassifier
# Use default AUTO behavior (balanced for small data, importance-based for large)
clf = TabPFNClassifier(max_features_per_estimator=300)
# Force random feature subsampling
clf = TabPFNClassifier(
max_features_per_estimator=300,
feature_subsampling_method=FeatureSubsamplingMethod.RANDOM,
)
# Keep the first 50 features always, balance the remainder
clf = TabPFNClassifier(
max_features_per_estimator=300,
feature_subsampling_method=FeatureSubsamplingMethod.CONSTANT_AND_BALANCED,
constant_feature_count=50,
)
# Use Gini-based importance (requires sufficient training data)
clf = TabPFNClassifier(
max_features_per_estimator=300,
feature_subsampling_method=FeatureSubsamplingMethod.GINI_FEATURE_IMPORTANCE,
importance_top_k_count=150, # keep top-150 important features per estimator
)
For regression tasks, identical parameters apply to TabPFNRegressor in src/tabpfn/regressor.py.
Summary
- Feature subsampling ensures each TabPFN estimator respects the
max_features_per_estimatorbudget while maintaining ensemble diversity. FeatureSubsamplingMethoddefines five strategies: BALANCED, RANDOM, CONSTANT_AND_BALANCED, GINI_FEATURE_IMPORTANCE, and AUTO.- AUTO resolution automatically selects between Gini importance and balanced sampling based on dataset size and subsampling necessity.
- Implementation resides in
src/tabpfn/preprocessing/ensemble.py, with configuration defined insrc/tabpfn/preprocessing/configs.py. - Configure subsampling via classifier/regressor parameters, using
constant_feature_countorimportance_top_k_countfor method-specific tuning.
Frequently Asked Questions
What happens if my dataset has fewer features than max_features_per_estimator?
When the dataset dimensionality is below the budget, the subsampling logic returns None and each estimator trains on the full feature set. No subsampling occurs, and the FeatureSubsamplingMethod has no practical effect on that specific fit.
How does CONSTANT_AND_BALANCED differ from GINI_FEATURE_IMPORTANCE?
CONSTANT_AND_BALANCED mechanically keeps the first N features (indices 0 to constant_feature_count-1) regardless of their predictive value, making it suitable when feature order carries semantic meaning. GINI_FEATURE_IMPORTANCE dynamically selects features based on LightGBM Gini gain scores calculated from the training data, prioritizing predictive power over position.
When does AUTO select GINI_FEATURE_IMPORTANCE over BALANCED?
The AUTO resolver selects GINI_FEATURE_IMPORTANCE only when two conditions are met: the dataset actually requires subsampling (needs_subsampling is True), and the number of samples exceeds AUTO_FEATURE_SUBSAMPLING_IMPORTANCE_MIN_SAMPLES. If either condition fails, it defaults to BALANCED to avoid unstable importance estimates.
Can I use feature subsampling with TabPFNRegressor?
Yes. Both TabPFNClassifier and TabPFNRegressor in src/tabpfn/classifier.py and src/tabpfn/regressor.py respectively accept the feature_subsampling_method parameter and pass it to the shared preprocessing pipeline in src/tabpfn/preprocessing/ensemble.py. The behavior is identical for regression and classification tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →