Specialized TabPFN Checkpoints: Large-Features, Large-Samples, and Real-Data Finetuned Models

TabPFN provides specialized checkpoint files on Hugging Face optimized for high-dimensional data (up to ~1,000 features), large training sets (~30,000+ samples), and real-world tabular distributions, accessible via the model_path parameter in TabPFNClassifier and TabPFNRegressor.

The PriorLabs/TabPFN repository ships with multiple pre-trained transformer variants that extend beyond the default synthetic training distribution. These specialized TabPFN checkpoints are automatically downloaded from the Hugging Face hub on first instantiation and are stored in ~/.cache/tabpfn/ by default, unless overridden by the TABPFN_MODEL_CACHE_DIR environment variable.

Large-Features Checkpoints for High-Dimensional Data

When working with datasets containing hundreds of features, the default checkpoints may underperform. TabPFN provides large-features variants that maintain accuracy up to higher dimensionalities.

Classifier Variants

  • tabpfn-v2.5-classifier-v2.5_large-features-L.ckpt — Optimized for datasets with up to approximately 500 features while maintaining modest sample sizes (typically < 5,000 samples).

  • tabpfn-v2.5-classifier-v2.5_large-features-XL.ckpt — Extends capacity to approximately 1,000 features when used with max_features_per_estimator=1000 in the estimator configuration.

Real-Data Finetuned Large-Features

  • tabpfn-v2.5-classifier-v2.5_real-large-features.ckpt — Combines real-world data finetuning with high-dimensional capability, though performance degrades on very large samples (> 10,000 instances).

Large-Samples Checkpoints for Big Training Sets

Standard TabPFN models are trained on relatively small synthetic datasets. For production workloads with extensive training data, use the large-samples specialization.

  • tabpfn-v2.5-classifier-v2.5_large-samples.ckpt — Specifically tailored for large training sets of approximately 30,000+ samples, maintaining calibration and accuracy where standard models would saturate.

Note that the checkpoint tabpfn-v2.5-classifier-v2.5_real-large-samples-and-features.ckpt is identical to the default real-data checkpoint and exists for naming completeness rather than distinct functionality.

Real-Data Finetuned Checkpoints

Checkpoints marked with 🌍 indicate models finetuned on real-world tabular datasets rather than purely synthetic prior-fitting networks. These generally provide superior performance on practical business and scientific data.

Classification

  • tabpfn-v2.5-classifier-v2.5_real.ckpt — The default recommended checkpoint for most classification tasks, finetuned on diverse real-world tabular distributions.

Regression

  • tabpfn-v2.5-regressor-v2.5_real.ckpt — The best-performing real-data finetuned regression checkpoint, recommended over synthetic-only alternatives.

  • tabpfn-v2.5-regressor-v2.5_real-variant.ckpt — An alternative real-data finetuned variant with slightly different performance characteristics, useful when the primary real checkpoint underperforms on specific benchmarks.

Specialized Regression Checkpoints

Beyond real-data variants, the regressor supports several task-specific optimizations:

  • tabpfn-v2.5-regressor-v2.5_low-skew.ckpt — Optimized for low-skew target distributions, though generally less robust than the standard model.

  • tabpfn-v2.5-regressor-v2.5_quantiles.ckpt — Intended for quantile and distribution estimation tasks rather than point predictions.

  • tabpfn-v2.5-regressor-v2.5_small-samples.ckpt — Slightly better performance on very small datasets (< 3,000 samples).

  • tabpfn-v2.5-regressor-v2.5_variant.ckpt — A generic variant without a clear specialty, occasionally advantageous on specific benchmark configurations.

Loading Specialized Checkpoints in Python

In src/tabpfn/classifier.py and src/tabpfn/regressor.py, the model_path parameter accepts either a local file path or a Hugging Face hub reference. The ModelVersion enum defined in src/tabpfn/constants.py provides version constants for programmatic selection.

Loading a Large-Features Classifier

from tabpfn import TabPFNClassifier

# Load the XL variant supporting ~1,000 features

clf = TabPFNClassifier(
    model_path="~/.cache/tabpfn/tabpfn-v2.5-classifier-v2.5_large-features-XL.ckpt"
)
clf.fit(X_train, y_train)
predictions = clf.predict(X_test)

Loading a Real-Data Finetuned Regressor

from tabpfn import TabPFNRegressor

# Best real-data regression checkpoint

reg = TabPFNRegressor(
    model_path="~/.cache/tabpfn/tabpfn-v2.5-regressor-v2.5_real.ckpt"
)
reg.fit(X_train, y_train)
predictions = reg.predict(X_test)

Using Factory Methods for Default Selection

from tabpfn import TabPFNClassifier, ModelVersion

# Automatically selects the appropriate default for V2.5 (real-data finetuned classifier)

clf = TabPFNClassifier.create_default_for_version(ModelVersion.V2_5)
clf.fit(X_train, y_train)

Switching Checkpoints After Initialization

from tabpfn import TabPFNClassifier, ModelVersion

clf = TabPFNClassifier.create_default_for_version(ModelVersion.V2_5)
clf.set_params(
    model_path="~/.cache/tabpfn/tabpfn-v2.5-classifier-v2.5_large-samples.ckpt"
)
clf.fit(X_train, y_train)

Checkpoint Caching and Environment Configuration

The utility functions in src/tabpfn/model_loading.py handle checkpoint persistence. By default, models cache to ~/.cache/tabpfn/, but you can redirect this for headless or containerized environments:

import os

os.environ["TABPFN_MODEL_CACHE_DIR"] = "/path/to/custom/cache"
os.environ["TABPFN_TOKEN"] = "hf_..."  # If accessing gated models

from tabpfn import TabPFNClassifier

clf = TabPFNClassifier(model_path="tabpfn-v2.5-classifier-v2.5_real.ckpt")

Summary

  • Large-features checkpoints (large-features-L.ckpt and large-features-XL.ckpt) support up to 500 and 1,000 features respectively, available for both standard and real-data finetuned classifiers.

  • Large-samples checkpoints (large-samples.ckpt) handle training sets of 30,000+ instances where standard models degrade.

  • Real-data finetuned checkpoints (marked 🌍 in documentation) include tabpfn-v2.5-classifier-v2.5_real.ckpt and tabpfn-v2.5-regressor-v2.5_real.ckpt, providing superior performance on practical tabular data compared to synthetic-only training.

  • All checkpoints are specified via the model_path parameter in TabPFNClassifier or TabPFNRegressor, with automatic downloading from Hugging Face on first use.

Frequently Asked Questions

Which checkpoint should I use for a dataset with 800 features?

Use tabpfn-v2.5-classifier-v2.5_large-features-XL.ckpt (or the real-data variant tabpfn-v2.5-classifier-v2.5_real-large-features.ckpt if you have < 10,000 samples). Set max_features_per_estimator=1000 when initializing the estimator to ensure the model allocates sufficient capacity for the high dimensionality.

What is the difference between synthetic and real-data finetuned checkpoints?

Synthetic checkpoints (including the original default) are trained purely on prior distributions and synthetic tabular data generation. Real-data finetuned checkpoints (indicated by "real" in the filename) undergo additional training on actual tabular datasets from sources like OpenML, making them more robust to real-world noise, feature correlations, and distribution shifts. According to the PriorLabs/TabPFN source code, the real-data classifier (tabpfn-v2.5-classifier-v2.5_real.ckpt) is now the default recommendation for most tasks.

Can I use large-features checkpoints with large sample sizes simultaneously?

Generally no—specialized checkpoints optimize for specific regimes. The tabpfn-v2.5-classifier-v2.5_real-large-features.ckpt specifically degrades when samples exceed 10,000 instances. For datasets with both high dimensionality and large sample counts (> 10,000), test the standard real-data checkpoint or the dedicated large-samples variant, though you may need to experiment to find the optimal trade-off between feature capacity and sample capacity.

How do I persist custom checkpoints in a specific directory?

Set the TABPFN_MODEL_CACHE_DIR environment variable before importing the library. The load_fitted_tabpfn_model and save_fitted_tabpfn_model functions in src/tabpfn/model_loading.py respect this path for both pre-trained foundation weights and user-saved model states. This is essential for offline or air-gapped deployments where automatic Hugging Face downloads are not possible.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →