A/B Testing & Experimentation: AI Skill for Product Management
A/B Testing & Experimentation is an AI skill that designs controlled experiments with statistical discipline: pre-registered hypotheses, sample size up front, one primary metric plus guardrails, and no peeking at early results.
npx skills add Uxcel-Lab/product-skills --skill pm-experimentation-ab
What this skill does
Helps an AI assistant design controlled experiments that produce trustworthy decisions instead of laundering a guess into data.
- Checks first that an A/B test is the right tool for the traffic and the question.
- Writes falsifiable if-then-because hypotheses with success and failure criteria.
- Calculates sample size from baseline, minimum effect, and confidence before launch.
- Runs tests through a full business cycle instead of stopping at the first green p-value.
- Pairs one primary metric with guardrails that catch hidden damage.
- Reads segments and long-term effects, since an overall lift can hide a mobile drop.
- Sets iterate, pivot, or persevere thresholds before the test runs.
- Documents every result, including failures, in a searchable learning record.
When to use it
Use it when you need an AI assistant to:
- set up or review an A/B test;
- design a controlled product experiment;
- pick experiment metrics and guardrails;
- decide sample size, significance, and duration;
- interpret test results before acting on them;
- decide whether a change is ready to roll out.
Decisions this skill helps you make
| Decision | Options |
|---|---|
| Confidence threshold | 95% default, 99% for critical flows, 90% only for cheap reversible calls |
| Traffic split | 50/50 for fast reads, 90/10 when the variant carries real downside |
| A/B vs. multivariate | A/B by default, multivariate only on high-traffic surfaces |
| Metric type | A small standardized set plus a custom metric tied to the hypothesis |
| Leading vs. long-term reads | Leading indicators for early signal, holdout groups to confirm durability |
| Test prioritization | PIE scoring (Potential, Importance, Ease) to sequence a test backlog |
How the skill works
1
Check the tool fits
The skill confirms an A/B test is right: enough traffic, a behavioral change, one clear decision.
2
Pre-register the test
It writes a falsifiable hypothesis, calculates sample size, and sets decision thresholds before launch.
3
Read it honestly
It interprets results past the winner: segments, guardrails, long-term reads, and a documented decision.
