# How to Run A/B Test Analysis with Statistical Significance Validation in pm-skills

> Learn to run A/B test analysis with statistical significance validation using the pm-skills repository. Our guide shows you how to compute conversion lifts, p-values, and confidence intervals.

- Repository: [Pawel Huryn/pm-skills](https://github.com/phuryn/pm-skills)
- Tags: how-to-guide
- Published: 2026-07-07

---

**Running A/B test analysis with statistical significance validation in the pm-skills repository requires invoking the `/analyze-test` command to parse experiment data, which then applies the `ab-test-analysis` skill to compute conversion lifts, p-values, and confidence intervals while validating sample size and guardrail metrics.**

The **pm-skills** repository provides a structured workflow for A/B test analysis with statistical significance validation, enabling product managers to evaluate controlled experiments through automated statistical calculations. According to the source code in `phuryn/pm-skills`, the analysis combines a high-level command interface with a comprehensive skill definition that enforces rigorous validation across six critical stages.

## Overview of the A/B Test Analysis Architecture

The workflow separates data input from statistical computation. First, the command layer in [`pm-data-analytics/commands/analyze-test.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md) accepts multiple input formats. Then, the skill layer in [`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md) executes the statistical validation checklist and generates the final recommendation.

This architecture ensures that every analysis includes sample-size power calculations, randomization checks, and guardrail metric monitoring before producing a Ship, Extend, Stop, or Investigate verdict.

## Step 1: Invoke the Analyze-Test Command

The `/analyze-test` command serves as the entry point for all A/B test evaluations. Located in [`pm-data-analytics/commands/analyze-test.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md), this command parser accepts four distinct input types:

- **Summary statistics** (conversion rates and sample sizes)
- **Raw CSV files** (user-level conversion data)
- **Screenshots** (of experiment dashboards)
- **Textual descriptions** (natural language experiment summaries)

When you submit data via the chat interface, the command standardizes the input and triggers the `ab-test-analysis` skill.

## Step 2: Apply the AB-Test-Analysis Skill

The `ab-test-analysis` skill, defined in [`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md), executes a six-step validation framework that ensures statistical significance before generating recommendations.

### Understand the Experiment Context

The skill captures the hypothesis, variant definitions, primary metric, test duration, and traffic split percentages. This contextual data prevents misinterpretation of results by anchoring statistical outputs to specific business objectives.

### Validate the Test Setup

Before calculating outcomes, the skill checks for **sample-size adequacy** using power formulas, verifies **business-cycle duration** to avoid seasonality bias, detects **sample-ratio mismatch** (SRM) to catch randomization errors, and accounts for **novelty effects** that might skew early results.

### Calculate Statistical Significance

The skill computes **conversion rates** for each variant, **relative lift** percentages, **two-tailed z-test** p-values, and **95% confidence intervals**. These calculations use the standard error formula for binomial proportions:

```python
se = ((cr_control * (1 - cr_control) / len(control)) +
      (cr_variant * (1 - cr_variant) / len(variant))) ** 0.5

```

### Monitor Guardrail Metrics

The analysis checks secondary metrics such as revenue, latency, or engagement to ensure the variant does not introduce adverse effects that could damage overall user experience despite improving the primary conversion metric.

### Interpret Results and Recommendations

The skill maps quantitative outcomes to concrete product decisions:
- **Ship**: Statistically significant lift with no guardrail violations
- **Extend**: Underpowered results requiring additional sample
- **Stop**: Negative or flat results with significant p-values
- **Investigate**: Anomalous patterns requiring manual review

### Generate the Markdown Report

Finally, the skill produces a structured markdown document containing the full analysis, statistical tables, and next-step recommendations, creating an auditable record for stakeholders.

## Practical Implementation Examples

### Analyzing Summary Statistics via Chat Interface

For quick evaluations with existing metrics, invoke the command directly with summary statistics:

```text
/analyze-test Control: 4.2% conversion (n=5000), Variant: 4.8% conversion (n=5100)

```

The system returns a formatted analysis following the template in the skill file:

```markdown

## A/B Test Analysis: MySignupExperiment

**Hypothesis**: New signup flow improves conversion
**Duration**: 14 days | **Sample**: 5000 control / 5100 variant

| Variant | Sample | Conversion | 95% CI |
|---------|--------|------------|--------|
| Control | 5000   | 4.2%       | 4.0%–4.4% |
| Variant | 5100   | 4.8%       | 4.6%–5.0% |

**Relative lift**: +14.3% (95% CI +8.2% – +20.4%)
**p‑value**: 0.0012 → **Significant**
**Recommendation**: **SHIP** – lift is statistically and practically significant, no guardrail issues.

```

### Processing Raw CSV Data with Python

For user-level data analysis, create a CSV file with individual conversion records:

```csv
user_id,variant,converted
1,control,0
2,variant,1
3,control,1

```

Upload the file and invoke:

```text
/analyze-test [upload ab_test.csv]

```

The command generates a temporary Python script using `pandas` and `scipy.stats` to compute the statistics. The underlying calculation logic follows this implementation:

```python
import pandas as pd
from scipy import stats

df = pd.read_csv('ab_test.csv')
control = df[df.variant == 'control']
variant = df[df.variant == 'variant']

def conversion_rate(group):
    return group.converted.mean()

cr_control = conversion_rate(control)
cr_variant = conversion_rate(variant)

lift = (cr_variant - cr_control) / cr_control
se = ((cr_control * (1 - cr_control) / len(control)) +
      (cr_variant * (1 - cr_variant) / len(variant))) ** 0.5
z = lift / se
p = 2 * (1 - stats.norm.cdf(abs(z)))
ci_low, ci_high = stats.norm.interval(0.95, loc=lift, scale=se)

print(f'Control CR: {cr_control:.4%}')
print(f'Variant CR: {cr_variant:.4%}')
print(f'Lift: {lift:.2%} (95% CI: {ci_low:.2%} – {ci_high:.2%})')
print(f'p-value: {p:.4f}')

```

Running this script produces the exact metrics embedded in the skill's markdown report.

## Key Statistical Methods in the Workflow

The **pm-skills** repository implements industry-standard statistical tests for binomial outcomes:

- **Two-tailed z-test**: Compares conversion rates between variants assuming normal approximation of the binomial distribution
- **Confidence interval calculation**: Uses `scipy.stats.norm.interval` to compute the 95% range for the true lift
- **Power analysis**: Validates that sample sizes achieve sufficient statistical power (typically 80%) to detect minimum detectable effects
- **Sample Ratio Mismatch (SRM) detection**: Chi-square test comparing observed traffic split to expected randomization ratios

## Summary

- The `/analyze-test` command in [`pm-data-analytics/commands/analyze-test.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md) accepts multiple input formats to initiate A/B test analysis
- The `ab-test-analysis` skill in [`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md) enforces a six-step validation including setup verification, statistical calculation, and guardrail monitoring
- Statistical significance is determined using two-tailed z-tests with 95% confidence intervals calculated via standard error formulas
- Recommendations are categorized as Ship, Extend, Stop, or Investigate based on statistical and practical significance thresholds
- The workflow generates reproducible markdown reports suitable for stakeholder review and audit trails

## Frequently Asked Questions

### What input formats does the analyze-test command support?

The command accepts summary statistics (conversion rates and sample sizes), raw CSV files containing user-level data, screenshots of experiment dashboards, and textual descriptions. This flexibility allows product managers to analyze experiments regardless of whether they have access to raw data warehouses or just summary reports.

### How does the skill determine if results are statistically significant?

The skill calculates a two-tailed z-test p-value using the standard error of the difference between conversion rates. If the p-value falls below the alpha threshold (typically 0.05) and the 95% confidence interval for the lift does not include zero, the result is flagged as statistically significant. The calculation uses `scipy.stats.norm.cdf` for precise probability estimation.

### What is the difference between statistical and practical significance in this context?

Statistical significance indicates that the observed difference between variants is unlikely due to random chance (p < 0.05), while practical significance requires that the relative lift exceeds a minimum threshold that justifies implementation costs. The skill flags "practical significance" separately to prevent shipping technically significant but business-negligible improvements.

### Can I customize the confidence level for the analysis?

While the default implementation uses 95% confidence intervals (alpha = 0.05), the Python script generated by the command exposes the confidence level parameter in `stats.norm.interval(0.95, loc=lift, scale=se)`. Advanced users can modify this value in the generated code or extend the skill definition to accept custom confidence levels as command parameters.