TabPFN Dataset Size Limitations: Understanding ignore_pretraining_limits and Model Constraints

TabPFN enforces strict dataset size limits—500 to 2,000 features and 10,000 to 50,000 samples depending on model version—and setting ignore_pretraining_limits=True bypasses these validations to allow larger datasets at the risk of degraded performance and disabled CPU safeguards.

TabPFN, the transformer-based tabular data model developed by PriorLabs, imposes maximum sample and feature constraints derived directly from its pretraining regime. Understanding these TabPFN dataset size limitations is essential for production deployments, as exceeding them raises a TabPFNValidationError unless you explicitly override the built-in safeguards.

Default Dataset Size Limits by Model Version

TabPFN validates dataset dimensions against hardcoded constants defined in src/tabpfn/inference_config.py. These limits vary by pretrained model version.

Model v2 Limits

The base model v2 supports datasets up to:

  • 500 features (MAX_NUMBER_OF_FEATURES = 500)
  • 10,000 samples (MAX_NUMBER_OF_SAMPLES = 10_000)

These constants are defined at lines 1088-1091 and 1094-1097 in src/tabpfn/inference_config.py.

Model v2_5 and Newer Limits

Updated model versions (v2_5 and later) expand capacity to:

  • 2,000 features (MAX_NUMBER_OF_FEATURES = 2000)
  • 50,000 samples (MAX_NUMBER_OF_SAMPLES = 50_000)

You can find these definitions at lines 1335-1337 and 1345-1347 in the same configuration file.

How Dataset Validation Works

Before fitting or predicting, TabPFN checks dataset dimensions through a validation chain implemented in src/tabpfn/validation.py.

The function _validate_num_samples_and_features performs the actual limit checking:

def _validate_num_samples_and_features(..., ignore_pretraining_limits: bool = False):
    if ignore_pretraining_limits:
        return                      # ← skip the checks

    # … raise TabPFNValidationError if limits are exceeded

This function is invoked by validate_dataset_size, which is called from ensure_compatible_fit_inputs—the central entry point used by both TabPFNClassifier and TabPFNRegressor at lines 111-138 in src/tabpfn/validation.py.

The Effect of ignore_pretraining_limits

The ignore_pretraining_limits parameter controls whether TabPFN enforces its pretrained size constraints.

Default Behavior (False)

When ignore_pretraining_limits=False (the default), the validation logic compares your dataset against the model's limits. Exceeding either threshold raises a TabPFNValidationError with a clear message:


TabPFNValidationError: Number of samples `12000` in the input data is greater than the maximum number of samples `10000` officially supported by TabPFN. Set `ignore_pretraining_limits=True` to override this error!

Bypassing Limits (True)

Setting ignore_pretraining_limits=True triggers three critical changes:

  • Validation skipped: The size checks (num_samples > max_num_samples and num_features > max_num_features) are completely bypassed
  • Performance degradation risk: The model will execute but may produce suboptimal results because the transformer was never exposed to datasets of that scale during pretraining
  • CPU safeguards disabled: The flag also propagates to CPU-specific large-dataset override logic at lines 146-154 in src/tabpfn/validation.py

Practical Code Examples

Example 1: Default Behavior with Exceeded Limits

Attempting to fit a dataset that exceeds v2 limits without override:

from tabpfn import TabPFNRegressor
import numpy as np

X = np.random.randn(12_000, 600)   # 12,000 samples, 600 features (exceeds both limits)

y = np.random.randn(12_000)

try:
    reg = TabPFNRegressor(device="cpu")  # ignore_pretraining_limits=False (default)

    reg.fit(X, y)
except Exception as e:
    print(e)   # -> TabPFNValidationError about sample count

Example 2: Overriding Limits with ignore_pretraining_limits

Fitting the same oversized dataset by disabling safeguards:

reg = TabPFNRegressor(device="cpu", ignore_pretraining_limits=True)
reg.fit(X, y)  # Succeeds, but note that model was not pretrained on such a large dataset

Example 3: Leveraging Newer Model Versions

Using v2_5 to accommodate larger datasets within official limits:

from tabpfn import TabPFNClassifier

clf = TabPFNClassifier(
    device="cpu",
    model_version="v2_5",               # selects the v2_5 inference config

    ignore_pretraining_limits=False,
)

X_small = np.random.randn(45_000, 1_800)   # Within v2_5 limits (50,000 samples, 2,000 features)

y_small = np.random.randint(0, 2, size=45_000)

clf.fit(X_small, y_small)   # Works without needing ignore_pretraining_limits

Key Source Files

Understanding the implementation requires familiarity with these specific files:

Summary

  • TabPFN enforces strict dataset size limits based on pretraining data: 500 features / 10,000 samples for v2 models, and 2,000 features / 50,000 samples for v2_5+
  • The validation occurs in _validate_num_samples_and_features within src/tabpfn/validation.py before any fitting or prediction
  • Setting ignore_pretraining_limits=True skips size validation and CPU safeguards, allowing larger datasets to run but risking performance degradation
  • Prefer upgrading to model_version="v2_5" rather than disabling limits when your data exceeds v2 constraints but fits within v2_5 thresholds

Frequently Asked Questions

What happens if I set ignore_pretraining_limits=True on a dataset that exceeds the limits?

The model will execute without raising a TabPFNValidationError, but you may experience degraded predictive performance. The transformer was pretrained exclusively on datasets within the official limits, so it has learned no patterns for larger scales. Additionally, CPU-specific safeguards for large datasets are disabled when this flag is active, as implemented in src/tabpfn/validation.py.

Can I use ignore_pretraining_limits with any TabPFN model version?

Yes. The ignore_pretraining_limits parameter works across all model versions (v2, v2_5, etc.), bypassing the specific limits defined in src/tabpfn/inference_config.py for each. However, newer versions like v2_5 natively support larger datasets (up to 50,000 samples and 2,000 features), making the override unnecessary for moderately large data that fits within the expanded v2_5 constraints.

Where does TabPFN check dataset size limits?

The check occurs in src/tabpfn/validation.py inside the _validate_num_samples_and_features function, which is called through validate_dataset_size and ensure_compatible_fit_inputs. This validation chain runs before the model begins fitting or predicting, ensuring constraints are enforced early in the pipeline.

What is the maximum dataset size TabPFN can handle without ignore_pretraining_limits?

Without overriding limits, the maximum depends on your model version. Model v2 supports up to 10,000 samples and 500 features, while model v2_5 and newer support up to 50,000 samples and 2,000 features. These constants are defined in src/tabpfn/inference_config.py and enforced at runtime by the validation logic in src/tabpfn/validation.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →