How PostHog Experiments Statistical Analysis Calculates Significance: A Bayesian Deep Dive

PostHog determines experiment significance using a Bayesian approach that combines win-probability calculations with expected-loss thresholds, evaluating count-based, continuous, and funnel metrics through posterior sampling rather than traditional p-values.

PostHog's experiments statistical analysis calculates significance using a sophisticated Bayesian framework implemented across three specialized modules in the PostHog/posthog repository. Unlike frequentist A/B testing that relies on p-values, the system employs probabilistic modeling to quantify both the likelihood that a variant is best and the potential cost of choosing incorrectly. This methodology provides product teams with intuitive, risk-aware significance determinations for any metric type.

The Bayesian Significance Framework

The PostHog experiments engine processes three distinct metric families, each with specialized statistical models:

Count and Rate Metrics

For events-per-exposure calculations (e.g., "how many purchases per user"), PostHog uses a Gamma-Poisson conjugate prior model. The implementation in posthog/hogql_queries/experiments/trends_statistics_v2_count.py treats each variant's event rate as a Gamma-distributed random variable, allowing for posterior updating that accounts for both the count of events and exposure volume.

Continuous Metrics

For revenue, duration, or other unbounded continuous values, the system employs a Normal-Inverse-Gamma posterior on the log-mean, effectively modeling the data as log-normal. The source in posthog/hogql_queries/experiments/trends_statistics_v2_continuous.py handles skewed distributions common in monetary metrics by working in log-space before transforming back to the original scale.

Funnel Conversion Metrics

For conversion-rate experiments, PostHog implements a Beta-Binomial model in posthog/hogql_queries/experiments/funnels_statistics_v2.py. This approach treats conversion rates as Beta-distributed probabilities updated by binomial observations (successes vs. failures).

The 7-Step Significance Calculation Process

Regardless of metric type, the experiments statistical analysis follows a rigorous seven-step pipeline to calculate significance:

1. Exposure Validation

Before statistical evaluation, each variant must satisfy minimum sample size requirements. The system checks against FF_DISTRIBUTION_THRESHOLD (defined in the experiment configuration) in the are_results_significant_* functions:

If any variant falls below the threshold, the function returns ExperimentSignificanceCode.NOT_ENOUGH_EXPOSURE.

2. Posterior Sampling and Win Probabilities

The engine draws SAMPLE_SIZE samples from each variant's posterior distribution to estimate the probability that each variant is the best:

  • Count: Uses gamma.rvs to draw from the Gamma posterior (calculate_probabilities_v2_count, lines 71–93)
  • Continuous: Draws from the t-distribution posterior in log-space (calculate_probabilities_v2_continuous, lines 71–93)
  • Funnel: Uses betabinom.rvs for Beta-Binomial sampling (calculate_probabilities_v2, lines 24–48)

The resulting probabilities array contains the fraction of draws where each variant outperforms all others.

3. Probability Threshold Check

If the highest win probability is ≥ 0.9 (MIN_PROBABILITY_FOR_SIGNIFICANCE), the experiment passes the first significance gate. Otherwise, are_results_significant_* returns LOW_WIN_PROBABILITY (count: lines 146–152, continuous: lines 69–71, funnel: lines 84–88).

4. Best Variant Identification

The system identifies the "best" variant based on empirical rates:

  • Count: Highest count/exposure ratio
  • Continuous: Highest raw mean
  • Funnel: Highest conversion rate (success_count / (success_count + failure_count))

5. Expected Loss Calculation

PostHog calculates the expected loss—the average penalty incurred if the chosen best variant is actually suboptimal:

  • Count: calculate_expected_loss_v2_count (lines 56–71) draws from the best variant's posterior and competitors' posteriors, computing the mean positive difference
  • Continuous: calculate_expected_loss_v2_continuous (lines 85–106) performs the calculation in log-space using the t-distribution posterior
  • Funnel: calculate_expected_loss_v2 (lines 20–38) averages the surplus of competitors' conversion rates over the target's

This produces a loss value between 0 and 1 (or monetary units for continuous metrics).

6. Loss Threshold Application

If the expected loss exceeds 0.05 (EXPECTED_LOSS_SIGNIFICANCE_LEVEL), the result is HIGH_LOSS. Otherwise, the experiment achieves SIGNIFICANT status:

7. Significance Code Return

The final output is an ExperimentSignificanceCode enum value (defined in posthog/schema.py), which the frontend consumes to display status badges like "Significant" or "High Loss".

Implementation Examples by Metric Type

Count-Based Analysis (Events per Exposure)

from posthog.schema import ExperimentVariantTrendsBaseStats
from posthog.hogql_queries.experiments.trends_statistics_v2_count import (
    calculate_probabilities_v2_count,
    are_results_significant_v2_count,
)

control = ExperimentVariantTrendsBaseStats(
    key="control", count=120, exposure=1, absolute_exposure=5000
)
test_a = ExperimentVariantTrendsBaseStats(
    key="test_a", count=150, exposure=1, absolute_exposure=5000
)

# Calculate win probabilities via Gamma posterior sampling

probabilities = calculate_probabilities_v2_count(control, [test_a])

# Returns: [0.12, 0.88] (control wins 12%, test_a wins 88%)

# Determine significance with expected loss check

code, loss = are_results_significant_v2_count(control, [test_a], probabilities)

# Returns: (ExperimentSignificanceCode.SIGNIFICANT, 0.03)

This example demonstrates the Gamma-Poisson model in trends_statistics_v2_count.py, where count represents total events and absolute_exposure represents the denominator for rate calculation.

Continuous Metric Analysis (Revenue)

from posthog.schema import ExperimentVariantTrendsBaseStats
from posthog.hogql_queries.experiments.trends_statistics_v2_continuous import (
    calculate_probabilities_v2_continuous,
    are_results_significant_v2_continuous,
)

control = ExperimentVariantTrendsBaseStats(
    key="control", count=5000.0, exposure=1, absolute_exposure=2000
)  # count stores total revenue

test = ExperimentVariantTrendsBaseStats(
    key="test", count=5600.0, exposure=1, absolute_exposure=2000
)

# Log-normal posterior sampling

prob = calculate_probabilities_v2_continuous(control, [test])
code, loss = are_results_significant_v2_continuous(control, [test], prob)

# Returns: (SIGNIFICANT, 0.07) 

# Interpretation: $0.07 expected loss per user if wrong variant chosen

The trends_statistics_v2_continuous.py implementation treats count as the sum of continuous values (like total revenue), modeling the per-user average via Normal-Inverse-Gamma conjugacy.

Funnel Conversion Analysis

from posthog.schema import ExperimentVariantFunnelsBaseStats
from posthog.hogql_queries.experiments.funnels_statistics_v2 import (
    calculate_probabilities_v2,
    are_results_significant_v2,
)

control = ExperimentVariantFunnelsBaseStats(
    key="control", success_count=200, failure_count=800
)
test = ExperimentVariantFunnelsBaseStats(
    key="test", success_count=260, failure_count=740
)

# Beta-Binomial posterior sampling

prob = calculate_probabilities_v2(control, [test])
code, loss = are_results_significant_v2(control, [test], prob)

# Returns: (SIGNIFICANT, 0.018)

# Interpretation: 1.8% conversion point expected loss if wrong

As implemented in funnels_statistics_v2.py, the Beta-Binomial model directly models uncertainty in conversion rates using success and failure counts.

Summary

  • PostHog employs a Bayesian + expected-loss framework that avoids p-value limitations by calculating both the probability a variant is best and the cost of potential errors.
  • Three specialized modules handle distinct metric types: Gamma-Poisson for counts (trends_statistics_v2_count.py), Normal-Inverse-Gamma for continuous values (trends_statistics_v2_continuous.py), and Beta-Binomial for conversions (funnels_statistics_v2.py).
  • Significance requires dual thresholds: a win probability ≥ 0.9 (MIN_PROBABILITY_FOR_SIGNIFICANCE) and an expected loss ≤ 0.05 (EXPECTED_LOSS_SIGNIFICANCE_LEVEL).
  • Exposure validation occurs before statistical testing, ensuring variants meet minimum sample size (FF_DISTRIBUTION_THRESHOLD) to prevent premature conclusions.
  • Results are deterministic rather than probabilistic—each function returns a concrete ExperimentSignificanceCode enum value consumed by the PostHog UI.

Frequently Asked Questions

How does PostHog's Bayesian approach differ from traditional frequentist A/B testing?

Traditional frequentist testing relies on p-values and null hypothesis significance testing, which can lead to misinterpretation of "significance" as the probability of a variant being best. PostHog's Bayesian experiments statistical analysis calculates significance by directly computing the probability that each variant is the best (win probability) and the expected loss from choosing the wrong variant. This provides intuitive, actionable metrics: you know exactly how confident you should be and what it costs to be wrong, rather than just rejecting a null hypothesis.

What happens if my experiment doesn't meet the sample size requirements?

If any variant falls below FF_DISTRIBUTION_THRESHOLD exposures (or sample size for funnels), the system immediately returns ExperimentSignificanceCode.NOT_ENOUGH_EXPOSURE from the are_results_significant_* functions without performing posterior sampling. This safeguard prevents false positives from insufficient data and is enforced in trends_statistics_v2_count.py (lines 136–143), trends_statistics_v2_continuous.py (lines 58–66), and funnels_statistics_v2.py (lines 78–82).

Why does PostHog use expected loss instead of just win probability?

Win probability alone doesn't capture the magnitude of potential error. A variant might have a 95% chance of being best, but if the alternative is only 0.01% worse, the risk is minimal. Conversely, a 90% win probability with a 50% performance downside represents significant risk. By calculating expected loss in calculate_expected_loss_v2_count, calculate_expected_loss_v2_continuous, and calculate_expected_loss_v2, PostHog ensures experiments are only marked significant when both the confidence is high (≥90%) and the potential downside is low (≤5% of the metric value).

Can I adjust the significance thresholds for my experiments?

The thresholds MIN_PROBABILITY_FOR_SIGNIFICANCE (default 0.9) and EXPECTED_LOSS_SIGNIFICANCE_LEVEL (default 0.05) are constants defined in the statistical modules. While these defaults are hardcoded in the current implementation within posthog/hogql_queries/experiments/, they represent industry-standard Bayesian decision theory thresholds that balance discovery speed with decision safety. The ExperimentSignificanceCode enum in posthog/schema.py provides the standardized result codes that the frontend interprets to display significance states.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →