# How to Use the pm-data-analytics Skill for A/B Test Result Analysis

> Automate A/B test result analysis with the pm-data-analytics skill. Validate design, calculate significance, and get reports instantly. Simplify experimentation today.

- Repository: [Pawel Huryn/pm-skills](https://github.com/phuryn/pm-skills)
- Tags: how-to-guide
- Published: 2026-07-10

---

**The pm-data-analytics skill automates end-to-end A/B test analysis through the `/analyze-test` command, validating experimental design, calculating statistical significance with scipy.stats, and generating decision-ready reports without manual coding.**

The `phuryn/pm-skills` repository provides a comprehensive toolkit for product management workflows, with the **pm-data-analytics** skill offering specialized capabilities for rigorous A/B test evaluation. Encapsulating statistical methodology into a structured pipeline, this skill processes everything from raw CSV exports to summary statistics while enforcing experimental best practices. The workflow is defined in [[`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md)](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md) and invoked via the command interface documented in [[`pm-data-analytics/commands/analyze-test.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md)](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md).

## Running the /analyze-test Command

The `/analyze-test` shortcut serves as the primary entry point for A/B test analysis, accommodating multiple input formats through a unified interface.

### Quick Inline Analysis with Summary Statistics

For rapid evaluation when summary statistics are already available, pass the data directly as arguments:

```text
/analyze-test Control: 4.2% conversion (n=5000), Variant: 4.8% conversion (n=5100)

```

The system parses the conversion percentages and sample sizes, executes the statistical routine defined in lines 33-38 of the skill definition, and returns a Markdown report with a **SHIP** recommendation if the lift passes significance and guardrail checks.

### Processing Raw Experiment Data

When analyzing granular user-level data, upload a CSV file containing the experiment results:

```text
/analyze-test [upload a CSV of test results]

```

The expected schema includes `user_id`, `variant`, `converted`, and `timestamp` columns:

| user_id | variant | converted | timestamp |
|---------|---------|-----------|-----------|
| 12345   | control | 0         | 2024-01-01|
| 12346   | variant | 1         | 2024-01-01|

Upon detecting raw data, the skill auto-generates and executes a Python script (see lines 41-42 of the command file) that leverages `pandas` and `scipy` for statistical computation:

```python
import pandas as pd
from scipy import stats

df = pd.read_csv("test_results.csv")
control = df[df.variant == "control"].converted
variant = df[df.variant == "variant"].converted

p = stats.proportions_ztest([variant.sum(), control.sum()],
                            [len(variant), len(control)])[1]
print(f"P‑value: {p}")

```

The resulting statistics are automatically injected into the final report template defined in lines 57-94 of the skill definition.

## The Six-Step Analytical Workflow

The **ab-test-analysis** skill implements a rigorous six-phase pipeline that mirrors industry-standard experimental analysis:

1. **Input handling** – Accepts raw CSVs, screenshots, or concise summary statistics and normalizes them for processing.
2. **Experiment validation** – Checks sample size adequacy, test duration, randomization integrity (SRM detection), and external confounding factors using the power-analysis methodology described in the *Validate the test setup* section.
3. **Statistical calculation** – Computes conversion rates, relative lift, two-tailed p-values, and 95% confidence intervals (lines 33-38).
4. **Guardrail assessment** – Evaluates secondary metrics for degradation that might contradict the primary metric success (lines 43-46).
5. **Decision matrix** – Maps statistical outcomes to one of four recommended actions: **Ship**, **Extend**, **Stop**, or **Investigate** (the decision table starting at line 49).
6. **Report generation** – Emits a Markdown summary containing the hypothesis, test duration, sample breakdown, results table, business-impact estimate, and follow-up suggestions (template lines 57-94).

## Programmatic Integration with invoke_skill

For automated pipelines or custom applications, the skill can be invoked programmatically using the toolkit's helper function:

```python
from pm_toolkit import invoke_skill

result = invoke_skill(
    "ab-test-analysis",
    arguments="Quarterly checkout flow experiment",
    data_path="checkout_ab_test.csv"
)
print(result.markdown_report)

```

The `invoke_skill` helper loads the skill definition from [`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md), executes the full validation and analysis pipeline, and returns a structured object containing the formatted Markdown report.

## Key Source Files and Architecture

Understanding the implementation requires familiarity with these repository components:

- **Skill definition**: [[`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md)](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md) – Contains the step-by-step methodology, statistical formulas, and decision matrix logic.
- **Command wrapper**: [[`pm-data-analytics/commands/analyze-test.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md)](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md) – Exposes the skill via the `/analyze-test` CLI shortcut and documents the user workflow.
- **Runtime Python helper**: Auto-generated during execution when raw data is supplied, executing calculations using `pandas` and `scipy.stats`.
- **Analytics suite overview**: [[`pm-data-analytics/README.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/README.md)](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/README.md) – Provides context on additional skills including cohort analysis and SQL query generation.

## Summary

- The **pm-data-analytics** skill delivers automated A/B test analysis through the `ab-test-analysis` workflow, accessible via the `/analyze-test` command.
- Input flexibility allows processing of both summary statistics and raw CSV data, with automatic Python script generation using `scipy.stats.proportions_ztest` for statistical validation.
- A comprehensive six-step pipeline ensures experimental rigor through setup validation, significance testing, guardrail assessment, and decision mapping to **Ship**, **Extend**, **Stop**, or **Investigate**.
- All statistical methodology resides in [`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md) while the user interface is defined in [`pm-data-analytics/commands/analyze-test.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md).
- The `invoke_skill` Python helper enables seamless integration into automated reporting systems and CI/CD pipelines.

## Frequently Asked Questions

### What input formats does the pm-data-analytics skill accept for A/B tests?

The skill accepts three primary formats: concise text summaries (e.g., "Control: 4.2% (n=5000), Variant: 4.8% (n=5100)"), raw CSV files with user-level columns including `user_id`, `variant`, and `converted`, and screenshot uploads of experiment dashboards. The input handler automatically detects the format and routes data to the appropriate parsing and calculation logic.

### How does the skill calculate statistical significance?

The skill computes two-tailed p-values and 95% confidence intervals using `scipy.stats.proportions_ztest` for binary conversion metrics. When processing raw data, it auto-generates a Python script (referenced in lines 41-42 of the command file) that executes the z-test against the control and variant arrays. The specific statistical formulas and thresholds are documented in lines 33-38 of [`pm-data-analytics/skills/ab-test-analysis/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md).

### What happens if secondary metrics degrade during analysis?

The guardrail assessment phase (lines 43-46 of the skill definition) evaluates secondary metrics for statistically significant degradation. If negative movement is detected in guardrail metrics, the decision matrix may override a primary metric win, changing the final recommendation from **Ship** to **Investigate** or **Stop** to prevent harmful business impact despite statistical significance in the primary metric.

### Can I integrate this skill into an automated reporting pipeline?

Yes, the skill supports programmatic invocation through the `invoke_skill` helper from the `pm_toolkit` module. By calling `invoke_skill("ab-test-analysis", arguments=..., data_path=...)` within Python scripts, you can execute the full analytical workflow, capture the Markdown report output, and integrate results into automated decision systems, Slack notifications, or scheduled analytics jobs.