How to Use the pm-data-analytics Skill for A/B Test Result Analysis

The pm-data-analytics skill automates end-to-end A/B test analysis through the /analyze-test command, validating experimental design, calculating statistical significance with scipy.stats, and generating decision-ready reports without manual coding.

The phuryn/pm-skills repository provides a comprehensive toolkit for product management workflows, with the pm-data-analytics skill offering specialized capabilities for rigorous A/B test evaluation. Encapsulating statistical methodology into a structured pipeline, this skill processes everything from raw CSV exports to summary statistics while enforcing experimental best practices. The workflow is defined in [pm-data-analytics/skills/ab-test-analysis/SKILL.md](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/skills/ab-test-analysis/SKILL.md) and invoked via the command interface documented in [pm-data-analytics/commands/analyze-test.md](https://github.com/phuryn/pm-skills/blob/main/pm-data-analytics/commands/analyze-test.md).

Running the /analyze-test Command

The /analyze-test shortcut serves as the primary entry point for A/B test analysis, accommodating multiple input formats through a unified interface.

Quick Inline Analysis with Summary Statistics

For rapid evaluation when summary statistics are already available, pass the data directly as arguments:

/analyze-test Control: 4.2% conversion (n=5000), Variant: 4.8% conversion (n=5100)

The system parses the conversion percentages and sample sizes, executes the statistical routine defined in lines 33-38 of the skill definition, and returns a Markdown report with a SHIP recommendation if the lift passes significance and guardrail checks.

Processing Raw Experiment Data

When analyzing granular user-level data, upload a CSV file containing the experiment results:

/analyze-test [upload a CSV of test results]

The expected schema includes user_id, variant, converted, and timestamp columns:

user_id variant converted timestamp
12345 control 0 2024-01-01
12346 variant 1 2024-01-01

Upon detecting raw data, the skill auto-generates and executes a Python script (see lines 41-42 of the command file) that leverages pandas and scipy for statistical computation:

import pandas as pd
from scipy import stats

df = pd.read_csv("test_results.csv")
control = df[df.variant == "control"].converted
variant = df[df.variant == "variant"].converted

p = stats.proportions_ztest([variant.sum(), control.sum()],
                            [len(variant), len(control)])[1]
print(f"P‑value: {p}")

The resulting statistics are automatically injected into the final report template defined in lines 57-94 of the skill definition.

The Six-Step Analytical Workflow

The ab-test-analysis skill implements a rigorous six-phase pipeline that mirrors industry-standard experimental analysis:

  1. Input handling – Accepts raw CSVs, screenshots, or concise summary statistics and normalizes them for processing.
  2. Experiment validation – Checks sample size adequacy, test duration, randomization integrity (SRM detection), and external confounding factors using the power-analysis methodology described in the Validate the test setup section.
  3. Statistical calculation – Computes conversion rates, relative lift, two-tailed p-values, and 95% confidence intervals (lines 33-38).
  4. Guardrail assessment – Evaluates secondary metrics for degradation that might contradict the primary metric success (lines 43-46).
  5. Decision matrix – Maps statistical outcomes to one of four recommended actions: Ship, Extend, Stop, or Investigate (the decision table starting at line 49).
  6. Report generation – Emits a Markdown summary containing the hypothesis, test duration, sample breakdown, results table, business-impact estimate, and follow-up suggestions (template lines 57-94).

Programmatic Integration with invoke_skill

For automated pipelines or custom applications, the skill can be invoked programmatically using the toolkit's helper function:

from pm_toolkit import invoke_skill

result = invoke_skill(
    "ab-test-analysis",
    arguments="Quarterly checkout flow experiment",
    data_path="checkout_ab_test.csv"
)
print(result.markdown_report)

The invoke_skill helper loads the skill definition from pm-data-analytics/skills/ab-test-analysis/SKILL.md, executes the full validation and analysis pipeline, and returns a structured object containing the formatted Markdown report.

Key Source Files and Architecture

Understanding the implementation requires familiarity with these repository components:

Summary

  • The pm-data-analytics skill delivers automated A/B test analysis through the ab-test-analysis workflow, accessible via the /analyze-test command.
  • Input flexibility allows processing of both summary statistics and raw CSV data, with automatic Python script generation using scipy.stats.proportions_ztest for statistical validation.
  • A comprehensive six-step pipeline ensures experimental rigor through setup validation, significance testing, guardrail assessment, and decision mapping to Ship, Extend, Stop, or Investigate.
  • All statistical methodology resides in pm-data-analytics/skills/ab-test-analysis/SKILL.md while the user interface is defined in pm-data-analytics/commands/analyze-test.md.
  • The invoke_skill Python helper enables seamless integration into automated reporting systems and CI/CD pipelines.

Frequently Asked Questions

What input formats does the pm-data-analytics skill accept for A/B tests?

The skill accepts three primary formats: concise text summaries (e.g., "Control: 4.2% (n=5000), Variant: 4.8% (n=5100)"), raw CSV files with user-level columns including user_id, variant, and converted, and screenshot uploads of experiment dashboards. The input handler automatically detects the format and routes data to the appropriate parsing and calculation logic.

How does the skill calculate statistical significance?

The skill computes two-tailed p-values and 95% confidence intervals using scipy.stats.proportions_ztest for binary conversion metrics. When processing raw data, it auto-generates a Python script (referenced in lines 41-42 of the command file) that executes the z-test against the control and variant arrays. The specific statistical formulas and thresholds are documented in lines 33-38 of pm-data-analytics/skills/ab-test-analysis/SKILL.md.

What happens if secondary metrics degrade during analysis?

The guardrail assessment phase (lines 43-46 of the skill definition) evaluates secondary metrics for statistically significant degradation. If negative movement is detected in guardrail metrics, the decision matrix may override a primary metric win, changing the final recommendation from Ship to Investigate or Stop to prevent harmful business impact despite statistical significance in the primary metric.

Can I integrate this skill into an automated reporting pipeline?

Yes, the skill supports programmatic invocation through the invoke_skill helper from the pm_toolkit module. By calling invoke_skill("ab-test-analysis", arguments=..., data_path=...) within Python scripts, you can execute the full analytical workflow, capture the Markdown report output, and integrate results into automated decision systems, Slack notifications, or scheduled analytics jobs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →