A/B Test Playbook
Test, learn, improve
A library of structured experiment briefs. Each scenario names the problem, the one variable that changes, the metric that decides the result and the guardrails that must hold.
- Find the evidence
- Frame the experiment
- Read the result

Evidence
Start with evidence, then narrow by page.
Use analytics, research or customer feedback to locate the problem first. Then narrow the library to that part of the experience and choose a scenario that matches what you observed.
Hypothesis
Write the hypothesis before building the variant.
State the observed problem, the change you expect to help and the outcome you expect to move. Then change one variable and keep the rest of the experience stable.
The only thing that changes:Coupon code field
Coupon field open in the cart, a directly visible box
Coupon field hidden behind an “I have a discount code” link
The rest of the cart is identical on both sides.
211 scenarios in the library come with both sides written like this
Guardrails
Decide what must not get worse.
A lift is not a win if it damages margin, refund rate, coupon usage, accessibility or another important part of the experience. Set those guardrails before reading the result.
- Don't remove the coupon field entirely - users with a code will be frustrated.
- Don't leave an invalid-code error ambiguous.
- Don't launch or end an active campaign during the test.
- Don't hide the coupon field so far it becomes unfindable.
- Don't change both the position and the copy in the same test.
The library
Find a scenario that matches the problem you observed.
Filter by category and page. Each entry keeps the hypothesis, the changing variable, the primary KPI and the guardrails together.
Changed:field label position
Above vs Left-Aligned Form Label A/B Test Example
Decided by Form completion rate
Open scenarioChanged:primary CTA offer
Adding a Request Demo CTA Beside Start Free A/B Test
Decided by Qualified opportunity rate
Open scenarioChanged:primary CTA position
Moving the Main CTA Above the Fold A/B Test
Decided by Primary action completion rate
Open scenarioChanged:number of product images
One Product Image vs Four Angles A/B Test
Product Detail Page
Decided by Purchase conversion rate (CR)
Open scenarioChanged:coupon code field
Visible vs Hidden Coupon Code Field A/B Test
Cart & Checkout
Decided by Revenue per visitor (RPV)
Open scenarioChanged:number of pricing plans
How Many Pricing Plans Should You Show? Test Idea
Decided by Revenue per visitor (RPV)
Open scenarioHow it works
Move from evidence to a decision.
Three steps from a real observation to a result you can act on.
Find the evidence
Start with the behavior, research finding or customer feedback that points to a problem.
8 scenarios on this page
Frame the experiment
Write the hypothesis, change one variable, choose one primary KPI and set the guardrails.
Coupon field hidden behind an “I have a discount code” link
control-vs-treatment · element
Read the result
Use the planned sample, statistics and guardrails together. A large-looking lift can still be noise.
5.00%
250 / 5,000
5.80%
290 / 5,000
- Relative uplift
- +16.0%
- p-value
- 0.0768
two-proportion z-test
A worked example from this site's own Significance calculator: a +16% lift that doesn't clear the bar.
The rules
Five rules for a valid test
Change one variable
Every variant pair changes exactly one thing. Ask for a multivariate test and it gets split into separate ones - insist, and the output says plainly that no one will know which change produced the result.
One primary metric
The first KPI in the list decides the winner. Presenting five metrics as equally important is the easiest way to call a losing test a win.
Always define a guardrail
Every scenario ships with at least one metric that must not degrade - margin, refund rate, speed, support tickets. If a change could affect accessibility, that's a guardrail candidate too.
Don't test away security steps
CAPTCHA, identity or age verification, two-factor login, legal consent steps - never proposed as friction to remove, even if asked. Those exist for protection, not conversion; the plugin says so and generates nothing.
State how strong the evidence is
Every suggestion says how strong the evidence behind it is - the user's own data, an archive precedent, an industry pattern, or a hunch. A weak-evidence idea can still be offered, but never dressed up as certain.
Install
- 1
Add the plugin to Claude Code
/plugin marketplace add ali-demirbas/ab-test-playbook/plugin install ab-test-playbook@ab-test-playbook - 2
Read the repository, or try the demo
Questions
Frequently asked questions
Answers drawn from the plugin's own methodology docs, not external citations.
What should I A/B test first?
Rank candidates with ICE (Impact x Confidence x Ease), not gut feeling. A low-effort test on a high-traffic page beats an ambitious test on a low-traffic one.
How many visitors do I need for an A/B test?
Not a rule-of-thumb number - it's computed from your actual baseline conversion rate and the minimum effect size you care about detecting. Without real traffic data, no duration or sample-size promise is made.
Can I peek at results early and stop when they look significant?
No - repeatedly checking a test and stopping the moment it looks significant inflates the false-positive rate well above 5%, even with no real difference. Decide sample size or duration up front, look once. The one exception: a guardrail metric visibly breaking mid-test.
Why run a test for at least two full weeks?
Not a statistical-power requirement - a coverage requirement. Weekday/weekend behavior and payday effects need to be represented in the data, even if the sample-size target is hit in three days.
What are common A/B testing mistakes that invalidate a result?
Changing more than one variable at once; declaring a winner from the first days of data; reading conversion rate alone on a price test (revenue per visitor can drop even as CR rises); redesigning "Variant A" instead of testing the real page as-is; running overlapping tests on the same page.
Is this playbook the right tool for every product?
No, and it says so. It fits B2C e-commerce, consumer mobile apps and self-serve SaaS with real weekly traffic. It fits poorly for low-traffic enterprise sales pages, long sales cycles, or heavily regulated flows - for those, it points to qualitative methods instead of forcing a split test where it doesn't belong.
If this work overlaps with yours, let's talk.
Send me a note about growth, CRM, measurement or one of the projects here. A question, a counterpoint or a simple hello all work.
