How to Run A/B Test Analysis with Statistical Significance Validation in pm-skills

Running A/B test analysis with statistical significance validation in the pm-skills repository requires invoking the /analyze-test command to parse experiment data, which then applies the ab-test-analysis skill to compute conversion lifts, p-values, and confidence intervals while validating sample size and guardrail metrics.

The pm-skills repository provides a structured workflow for A/B test analysis with statistical significance validation, enabling product managers to evaluate controlled experiments through automated statistical calculations. According to the source code in phuryn/pm-skills, the analysis combines a high-level command interface with a comprehensive skill definition that enforces rigorous validation across six critical stages.

Overview of the A/B Test Analysis Architecture

The workflow separates data input from statistical computation. First, the command layer in pm-data-analytics/commands/analyze-test.md accepts multiple input formats. Then, the skill layer in pm-data-analytics/skills/ab-test-analysis/SKILL.md executes the statistical validation checklist and generates the final recommendation.

This architecture ensures that every analysis includes sample-size power calculations, randomization checks, and guardrail metric monitoring before producing a Ship, Extend, Stop, or Investigate verdict.

Step 1: Invoke the Analyze-Test Command

The /analyze-test command serves as the entry point for all A/B test evaluations. Located in pm-data-analytics/commands/analyze-test.md, this command parser accepts four distinct input types:

  • Summary statistics (conversion rates and sample sizes)
  • Raw CSV files (user-level conversion data)
  • Screenshots (of experiment dashboards)
  • Textual descriptions (natural language experiment summaries)

When you submit data via the chat interface, the command standardizes the input and triggers the ab-test-analysis skill.

Step 2: Apply the AB-Test-Analysis Skill

The ab-test-analysis skill, defined in pm-data-analytics/skills/ab-test-analysis/SKILL.md, executes a six-step validation framework that ensures statistical significance before generating recommendations.

Understand the Experiment Context

The skill captures the hypothesis, variant definitions, primary metric, test duration, and traffic split percentages. This contextual data prevents misinterpretation of results by anchoring statistical outputs to specific business objectives.

Validate the Test Setup

Before calculating outcomes, the skill checks for sample-size adequacy using power formulas, verifies business-cycle duration to avoid seasonality bias, detects sample-ratio mismatch (SRM) to catch randomization errors, and accounts for novelty effects that might skew early results.

Calculate Statistical Significance

The skill computes conversion rates for each variant, relative lift percentages, two-tailed z-test p-values, and 95% confidence intervals. These calculations use the standard error formula for binomial proportions:

se = ((cr_control * (1 - cr_control) / len(control)) +
      (cr_variant * (1 - cr_variant) / len(variant))) ** 0.5

Monitor Guardrail Metrics

The analysis checks secondary metrics such as revenue, latency, or engagement to ensure the variant does not introduce adverse effects that could damage overall user experience despite improving the primary conversion metric.

Interpret Results and Recommendations

The skill maps quantitative outcomes to concrete product decisions:

  • Ship: Statistically significant lift with no guardrail violations
  • Extend: Underpowered results requiring additional sample
  • Stop: Negative or flat results with significant p-values
  • Investigate: Anomalous patterns requiring manual review

Generate the Markdown Report

Finally, the skill produces a structured markdown document containing the full analysis, statistical tables, and next-step recommendations, creating an auditable record for stakeholders.

Practical Implementation Examples

Analyzing Summary Statistics via Chat Interface

For quick evaluations with existing metrics, invoke the command directly with summary statistics:

/analyze-test Control: 4.2% conversion (n=5000), Variant: 4.8% conversion (n=5100)

The system returns a formatted analysis following the template in the skill file:


## A/B Test Analysis: MySignupExperiment

**Hypothesis**: New signup flow improves conversion
**Duration**: 14 days | **Sample**: 5000 control / 5100 variant

| Variant | Sample | Conversion | 95% CI |
|---------|--------|------------|--------|
| Control | 5000   | 4.2%       | 4.0%–4.4% |
| Variant | 5100   | 4.8%       | 4.6%–5.0% |

**Relative lift**: +14.3% (95% CI +8.2% – +20.4%)
**p‑value**: 0.0012 → **Significant**
**Recommendation**: **SHIP** – lift is statistically and practically significant, no guardrail issues.

Processing Raw CSV Data with Python

For user-level data analysis, create a CSV file with individual conversion records:

user_id,variant,converted
1,control,0
2,variant,1
3,control,1

Upload the file and invoke:

/analyze-test [upload ab_test.csv]

The command generates a temporary Python script using pandas and scipy.stats to compute the statistics. The underlying calculation logic follows this implementation:

import pandas as pd
from scipy import stats

df = pd.read_csv('ab_test.csv')
control = df[df.variant == 'control']
variant = df[df.variant == 'variant']

def conversion_rate(group):
    return group.converted.mean()

cr_control = conversion_rate(control)
cr_variant = conversion_rate(variant)

lift = (cr_variant - cr_control) / cr_control
se = ((cr_control * (1 - cr_control) / len(control)) +
      (cr_variant * (1 - cr_variant) / len(variant))) ** 0.5
z = lift / se
p = 2 * (1 - stats.norm.cdf(abs(z)))
ci_low, ci_high = stats.norm.interval(0.95, loc=lift, scale=se)

print(f'Control CR: {cr_control:.4%}')
print(f'Variant CR: {cr_variant:.4%}')
print(f'Lift: {lift:.2%} (95% CI: {ci_low:.2%} – {ci_high:.2%})')
print(f'p-value: {p:.4f}')

Running this script produces the exact metrics embedded in the skill's markdown report.

Key Statistical Methods in the Workflow

The pm-skills repository implements industry-standard statistical tests for binomial outcomes:

  • Two-tailed z-test: Compares conversion rates between variants assuming normal approximation of the binomial distribution
  • Confidence interval calculation: Uses scipy.stats.norm.interval to compute the 95% range for the true lift
  • Power analysis: Validates that sample sizes achieve sufficient statistical power (typically 80%) to detect minimum detectable effects
  • Sample Ratio Mismatch (SRM) detection: Chi-square test comparing observed traffic split to expected randomization ratios

Summary

  • The /analyze-test command in pm-data-analytics/commands/analyze-test.md accepts multiple input formats to initiate A/B test analysis
  • The ab-test-analysis skill in pm-data-analytics/skills/ab-test-analysis/SKILL.md enforces a six-step validation including setup verification, statistical calculation, and guardrail monitoring
  • Statistical significance is determined using two-tailed z-tests with 95% confidence intervals calculated via standard error formulas
  • Recommendations are categorized as Ship, Extend, Stop, or Investigate based on statistical and practical significance thresholds
  • The workflow generates reproducible markdown reports suitable for stakeholder review and audit trails

Frequently Asked Questions

What input formats does the analyze-test command support?

The command accepts summary statistics (conversion rates and sample sizes), raw CSV files containing user-level data, screenshots of experiment dashboards, and textual descriptions. This flexibility allows product managers to analyze experiments regardless of whether they have access to raw data warehouses or just summary reports.

How does the skill determine if results are statistically significant?

The skill calculates a two-tailed z-test p-value using the standard error of the difference between conversion rates. If the p-value falls below the alpha threshold (typically 0.05) and the 95% confidence interval for the lift does not include zero, the result is flagged as statistically significant. The calculation uses scipy.stats.norm.cdf for precise probability estimation.

What is the difference between statistical and practical significance in this context?

Statistical significance indicates that the observed difference between variants is unlikely due to random chance (p < 0.05), while practical significance requires that the relative lift exceeds a minimum threshold that justifies implementation costs. The skill flags "practical significance" separately to prevent shipping technically significant but business-negligible improvements.

Can I customize the confidence level for the analysis?

While the default implementation uses 95% confidence intervals (alpha = 0.05), the Python script generated by the command exposes the confidence level parameter in stats.norm.interval(0.95, loc=lift, scale=se). Advanced users can modify this value in the generated code or extend the skill definition to accept custom confidence levels as command parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →