A/B Test Calculator
Two variants, head to head: the pooled two-proportion z-test with its p-value, the difference interval, and the lift with a log-scale interval — significance and size on one panel, so a win is never just a p-value.
A/B Test Calculator
Results recalculate instantly on every keystroke. Nothing you type is transmitted.
What this result does not account for
- One test, two arms — no multi-variant (A/B/n) corrections
- Normal-approximation z: tiny conversion counts should move to an exact test (out of scope)
In short: Variant A: 132 of 1,200 converted (11%). Variant B: 174 of 1,200 (14.5%). The pooled rate is 306/2,400 = 0.127500 and the test says z = 2.570430, two-sided p = 0.010157 — the difference clears the 5% bar. The lift is 31.818182%, with a 95% interval of (6.678334%, 62.882494%) built on the log scale; the absolute difference is 3.5 points with interval (0.008349, 0.061651), which sits entirely above zero. Read the pair: the test says the difference is real, the intervals say it could still be anywhere from a whisper to a shout — ship it, and keep the interval, not just the trophy.
Formula
z = (p̂₂−p̂₁)/√(p̂(1−p̂)(1/n₁+1/n₂)) · CIₑₓₚ = (p̂₂−p̂₁) ± z*·√(p̂₁q̂₁/n₁ + p̂₂q̂₂/n₂)
The test pools both arms under the no-difference hypothesis; the interval unpools, because once the difference is the quantity being estimated each arm carries its own variance. The lift interval lives on the log scale — ratios are skewed — using the same SE the risk-ratio page derives: √(q̂₁/x₁ + q̂₂/x₂).
Worked Example
- Enter visitors and conversions for both variants — whole counts, one arm each.
- Read the test: z and its two-sided p against your alpha (5% is the convention, typed in your head).
- Read BOTH intervals before celebrating: the difference interval for size, the lift interval for the headline.
- If the rates tie, the page says so honestly — no detected difference is a finding, not a failure.
Defaults: 11% vs 14.5% — z = 2.570430, p = 0.010157, difference interval (0.008349, 0.061651), lift 31.818182% with interval (6.678334%, 62.882494%). A tie drive (132 of 1,200 both arms) prints the honest no-difference card.
Strengths & Limits Of This Model
Where this engine is strong
- Lift interval on the log scale beside the difference interval
- Tie printed as a finding with the power pointer
Where it stops
- No sequential/peeking correction
- No Bayesian posterior variant
Practical Use Cases
Product
landing-page and pricing variants
Marketing
subject lines, bids, creative
Ops
two processes on two weeks of counts
Methodology & Editorial Standards
Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.
This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.
Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.
A/B Test Calculator — 8 Expert FAQs
8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.
Why does the test pool the proportions but the interval does not?
Different jobs. Under the no-difference hypothesis both arms estimate ONE shared rate, so the pooled p̂ gives the test its sharpest standard error. The interval makes no such hypothesis — it estimates the difference itself, so each arm keeps its own variance. Same data, two honest arithmetic.
The p-value is 0.010157 — how big is the win?
That is exactly the wrong question, and this page refuses to let it stay unasked: significance is about whether, the interval is about how much. Here the lift interval runs 6.678334% to 62.882494% — real, and still wildly loose. Quote both or quote neither.
Why is the lift interval built on the log scale?
Ratios are skewed: a lift can fall to −100% below but rise unbounded above, so a symmetric normal band around the raw ratio lies by construction. Logging bends the scale symmetric, the normal approximation goes back to being honest, and the result is exponentiated home — the same move the odds-ratio and risk-ratio pages make.
How many visitors do I need before the test?
More than the estimate needs — detecting a difference is harder than describing one proportion, which is why this n dwarfs the sample-size page's. The power-analysis page runs that question exactly: baseline, target effect, and alpha in; n per arm out. For these defaults it prices about 1,425 per arm for 80% power.
Can I peek at the test daily and stop when p dips under 0.05?
That is the peeking trap: repeated testing at a 5% bar is not a 5% procedure — the false positive rate balloons with every look. Fix the horizon before starting, or use a sequential design built for peeking. This page prices ONE test on ONE finished sample.
The rates tie — is the test broken?
No: identical rates give z = 0 and p = 1, the honest answer that this sample detected nothing. The card prints that as a finding — with the reminder that absence of evidence at a small n is a power statement, not proof of equivalence: the power-analysis page says how deaf your test was.
Why refuse an arm with zero visitors?
A proportion over nobody has no denominator, and the pooled SE collapses with it. One visitor is the honest floor; below that the page refuses rather than print z from an empty arm.
Is a significant lift always worth shipping?
The test prices randomness, not value: a significant 0.008 lift may cost more to implement than it returns, and a non-significant 3-point gap may merit a bigger rerun. The verdict card hands you the statistics; the shipping decision also owns cost, risk, and reversibility — numbers this page deliberately does not invent.