How FullSupportBarDistribution Computes Predictions in TabPFN Regression
FullSupportBarDistribution transforms discrete transformer logits into a continuous piecewise-uniform probability distribution over fixed target intervals, computing final predictions as the distribution mean while enabling quantile extraction and ensemble aggregation.
TabPFN treats regression as a distribution-over-bins problem rather than direct value prediction. The FullSupportBarDistribution—implemented in tabpfn/architectures/base/bar_distribution.py—converts raw logits into a calibrated probability distribution spanning the entire target support. This architecture enables probabilistic regression with well-defined uncertainty estimates while maintaining full differentiability through the inference pipeline.
Logits to Probabilities: The Softmax Transformation
During inference, the TabPFN transformer emits a vector of logits representing log-probabilities for each predefined interval (bar) of the target space. The BarDistribution class first applies a softmax normalization to convert these logits into valid probability masses.
In tabpfn/architectures/base/bar_distribution.py, this conversion occurs via:
import torch.nn.functional as F
# logits shape: (batch_size, num_bars)
probs = F.softmax(logits, dim=-1)
Each resulting probability value represents the mass assigned to a specific bar defined by the border set (e.g., intervals spanning [-1, 2] for the default regression checkpoint).
Constructing the Piecewise-Uniform Density
Once probabilities are obtained, FullSupportBarDistribution assumes a uniform density within each bar interval. The probability mass is spread evenly across the bar's width, calculated as borders[i+1] - borders[i]. This creates a piecewise-uniform probability density function (PDF) over the continuous target space.
The class provides get_probs_for_different_borders to remap distributions onto alternative binning schemes while preserving total probability mass. This method rescales probabilities when downstream components require different interval granularities:
# Map to a different border set while maintaining probability mass
new_probs = dist.get_probs_for_different_borders(new_borders)
From CDF to Point Predictions
Building the Cumulative Distribution Function
The cumulative distribution function (CDF) is constructed by cumulatively summing the uniform densities across bars. This enables probability queries such as "what is the probability the target ≤ x?" The BarDistribution.cdf method implements this functionality, allowing quantile extraction and uncertainty quantification.
Calculating the Mean Prediction
For point predictions, TabPFN regression returns the mean of the piecewise-uniform distribution. The BarDistribution.mean method computes this as a weighted sum of bar midpoints, using the softmax probabilities as weights:
# Inside TabPFNRegressor.predict()
mean_prediction = dist.mean(logits)
This mean value serves as the primary regression output, minimizing squared error while accounting for the full predictive distribution.
Advanced Distribution Operations
Quantiles and Median Extraction
Beyond the mean, FullSupportBarDistribution supports extraction of arbitrary quantiles and the median via inverse-CDF operations. The BarDistribution.median method provides the 50th percentile, while cdf enables computation of any quantile:
# Extract median and 25th percentile
median_val = dist.median(logits)
q25 = dist.cdf(logits, torch.tensor([0.25]))
Ensemble Averaging
For ensemble predictions, multiple BarDistribution objects can be merged using average_bar_distributions_into_this. This method averages the bar probabilities across ensemble members before normalization, producing a single consolidated distribution:
# Average multiple distributions into the current instance
dist.average_bar_distributions_into_this(other_distributions_list)
Why "Full Support" Matters
The "Full Support" designation indicates that the border set spans the entire target range encountered during training (e.g., [-1, 2] for default checkpoints). This guarantees that the resulting distribution has no zero-probability regions outside the training range, ensuring the CDF remains well-defined for any input value and preventing edge-case failures during extrapolation.
Practical Implementation Example
The following example demonstrates accessing the FullSupportBarDistribution internals for custom analysis:
from tabpfn import TabPFNRegressor
from tabpfn.architectures.base import bar_distribution
# Fit model and obtain predictions
reg = TabPFNRegressor()
reg.fit(X_train, y_train)
# Standard point prediction (uses BarDistribution.mean internally)
y_pred = reg.predict(X_test)
# Access raw distribution for uncertainty quantification
logits = reg._raw_predict_logits(X_test[:1]) # Raw logits per bar
borders = reg.model_.borders # Training border set
dist = bar_distribution.BarDistribution(borders=borders)
# Compute distribution statistics
mean_val = dist.mean(logits).item()
median_val = dist.median(logits).item()
quantile_90 = dist.cdf(logits, torch.tensor([0.9])).item()
Summary
- Logit transformation: FullSupportBarDistribution applies
softmaxto transformer outputs to obtain bar probabilities intabpfn/architectures/base/bar_distribution.py. - Density modeling: Probability mass is uniformly distributed within each bar interval, creating a piecewise-uniform PDF.
- Point prediction: The
BarDistribution.meanmethod computes predictions as the weighted average of bar midpoints. - Uncertainty quantification: The
cdfandmedianmethods enable extraction of arbitrary quantiles and confidence intervals. - Flexibility:
get_probs_for_different_borderssupports alternative binning schemes, whileaverage_bar_distributions_into_thisenables ensemble aggregation. - Full support guarantee: Borders span the complete training range (e.g.,
[-1, 2]), ensuring well-behaved distributions across the entire target space.
Frequently Asked Questions
How does FullSupportBarDistribution handle values outside the training range?
The distribution maintains probability mass across the full support defined during training (typically [-1, 2] for TabPFN regression checkpoints). Because the border set spans the entire range seen during training, the CDF remains well-defined for any query point, preventing zero-probability assignments to out-of-distribution values while preserving differentiability.
What is the difference between using predict() and accessing the distribution directly?
The TabPFNRegressor.predict() method internally calls BarDistribution.mean to return point estimates optimized for squared error loss. Accessing the distribution directly via BarDistribution objects allows extraction of medians, arbitrary quantiles, or full probability densities for uncertainty quantification and custom loss functions.
How does ensemble averaging work with BarDistribution?
The average_bar_distributions_into_this method merges multiple distributions by averaging their bar probabilities before renormalizing. This occurs in the probability space rather than the logit space, ensuring that ensemble predictions properly aggregate uncertainty estimates across multiple model checkpoints or ensemble members.
Can I change the number of bins or border locations after training?
Yes, the get_probs_for_different_borders method remaps the original distribution onto a new border set while preserving total probability mass. This allows downstream components to query predictions at different granularities or align distributions from different models without retraining the transformer backbone.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →