Ali Demirbaş

A/B Test Playbook

Test, learn, improve

A library of structured experiment briefs. Each scenario names the problem, the one variable that changes, the metric that decides the result and the guardrails that must hold.

  • Find the evidence
  • Frame the experiment
  • Read the result
A/B test library examples, each with the changed element, primary KPI and a guardrail.

Evidence

Start with evidence, then narrow by page.

Use analytics, research or customer feedback to locate the problem first. Then narrow the library to that part of the experience and choose a scenario that matches what you observed.

Scenarios by page14 · 211
Product page37
Home & landing30
Forms & signup23
UI elements23
Category listing23
Checkout14
SaaS & B2B13
Pricing12
Mobile app11
Cart8
Filters5
Search5
Thank you4
Dashboard3

Hypothesis

Write the hypothesis before building the variant.

State the observed problem, the change you expect to help and the outcome you expect to move. Then change one variable and keep the rest of the experience stable.

Control vs variant

The only thing that changes:Coupon code field

AControl

Coupon field open in the cart, a directly visible box

BVariant

Coupon field hidden behind an “I have a discount code” link

The rest of the cart is identical on both sides.

211 scenarios in the library come with both sides written like this

Guardrails

Decide what must not get worse.

A lift is not a win if it damages margin, refund rate, coupon usage, accessibility or another important part of the experience. Set those guardrails before reading the result.

What not to do
  1. Don't remove the coupon field entirely - users with a code will be frustrated.
  2. Don't leave an invalid-code error ambiguous.
  3. Don't launch or end an active campaign during the test.
  4. Don't hide the coupon field so far it becomes unfindable.
  5. Don't change both the position and the copy in the same test.
1,055guardrail rules across the library

The library

Find a scenario that matches the problem you observed.

Filter by category and page. Each entry keeps the hypothesis, the changing variable, the primary KPI and the guardrails together.

Forms & signup

Changed:field label position

Above vs Left-Aligned Form Label A/B Test Example

Decided by Form completion rate

Open scenario
SaaS & B2B

Changed:primary CTA offer

Adding a Request Demo CTA Beside Start Free A/B Test

Decided by Qualified opportunity rate

Open scenario
Home & landing

Changed:primary CTA position

Moving the Main CTA Above the Fold A/B Test

Decided by Primary action completion rate

Open scenario
Product page

Changed:number of product images

One Product Image vs Four Angles A/B Test

Product Detail Page

Decided by Purchase conversion rate (CR)

Open scenario
Cart

Changed:coupon code field

Visible vs Hidden Coupon Code Field A/B Test

Cart & Checkout

Decided by Revenue per visitor (RPV)

Open scenario
Pricing

Changed:number of pricing plans

How Many Pricing Plans Should You Show? Test Idea

Decided by Revenue per visitor (RPV)

Open scenario

How it works

Move from evidence to a decision.

Three steps from a real observation to a result you can act on.

01

Find the evidence

Start with the behavior, research finding or customer feedback that points to a problem.

Page
Product page37
Home & landing30
Forms & signup23
UI elements23
Category listing23
Cart8

8 scenarios on this page

02

Frame the experiment

Write the hypothesis, change one variable, choose one primary KPI and set the guardrails.

Hypothesis

Coupon field hidden behind an “I have a discount code” link

Primary KPIRevenue Per Visitor (RPV)
GuardrailCoupon Usage Rate must not collapse

control-vs-treatment · element

03

Read the result

Use the planned sample, statistics and guardrails together. A large-looking lift can still be noise.

AControl

5.00%

250 / 5,000

BVariant

5.80%

290 / 5,000

Relative uplift
+16.0%
p-value
0.0768
Significant at 95%?No, keep running

two-proportion z-test

A worked example from this site's own Significance calculator: a +16% lift that doesn't clear the bar.

The rules

Five rules for a valid test

1

Change one variable

Every variant pair changes exactly one thing. Ask for a multivariate test and it gets split into separate ones - insist, and the output says plainly that no one will know which change produced the result.

2

One primary metric

The first KPI in the list decides the winner. Presenting five metrics as equally important is the easiest way to call a losing test a win.

3

Always define a guardrail

Every scenario ships with at least one metric that must not degrade - margin, refund rate, speed, support tickets. If a change could affect accessibility, that's a guardrail candidate too.

4

Don't test away security steps

CAPTCHA, identity or age verification, two-factor login, legal consent steps - never proposed as friction to remove, even if asked. Those exist for protection, not conversion; the plugin says so and generates nothing.

5

State how strong the evidence is

Every suggestion says how strong the evidence behind it is - the user's own data, an archive precedent, an industry pattern, or a hunch. A weak-evidence idea can still be offered, but never dressed up as certain.

Install

  1. 1

    Add the plugin to Claude Code

    /plugin marketplace add ali-demirbas/ab-test-playbook
    /plugin install ab-test-playbook@ab-test-playbook
  2. 2

    Read the repository, or try the demo

Questions

Frequently asked questions

Answers drawn from the plugin's own methodology docs, not external citations.

What should I A/B test first?

Rank candidates with ICE (Impact x Confidence x Ease), not gut feeling. A low-effort test on a high-traffic page beats an ambitious test on a low-traffic one.

How many visitors do I need for an A/B test?

Not a rule-of-thumb number - it's computed from your actual baseline conversion rate and the minimum effect size you care about detecting. Without real traffic data, no duration or sample-size promise is made.

Can I peek at results early and stop when they look significant?

No - repeatedly checking a test and stopping the moment it looks significant inflates the false-positive rate well above 5%, even with no real difference. Decide sample size or duration up front, look once. The one exception: a guardrail metric visibly breaking mid-test.

Why run a test for at least two full weeks?

Not a statistical-power requirement - a coverage requirement. Weekday/weekend behavior and payday effects need to be represented in the data, even if the sample-size target is hit in three days.

What are common A/B testing mistakes that invalidate a result?

Changing more than one variable at once; declaring a winner from the first days of data; reading conversion rate alone on a price test (revenue per visitor can drop even as CR rises); redesigning "Variant A" instead of testing the real page as-is; running overlapping tests on the same page.

Is this playbook the right tool for every product?

No, and it says so. It fits B2C e-commerce, consumer mobile apps and self-serve SaaS with real weekly traffic. It fits poorly for low-traffic enterprise sales pages, long sales cycles, or heavily regulated flows - for those, it points to qualitative methods instead of forcing a split test where it doesn't belong.

If this work overlaps with yours, let's talk.

Send me a note about growth, CRM, measurement or one of the projects here. A question, a counterpoint or a simple hello all work.