Statistics

A/B Test Calculator

Two variants, head to head: the pooled two-proportion z-test with its p-value, the difference interval, and the lift with a log-scale interval — significance and size on one panel, so a win is never just a p-value.

A/B Test Calculator

Results recalculate instantly on every keystroke. Nothing you type is transmitted.

Variant A (control)
Variant B (treatment)
The lift
—
The test — z and p—
Difference interval—
Lift interval (log scale)—
Significance is not a size—

What this result does not account for

  • One test, two arms — no multi-variant (A/B/n) corrections
  • Normal-approximation z: tiny conversion counts should move to an exact test (out of scope)
● Zero-Server Execution Updated 11 Aug 2026 Reviewed by Sana Khalid IEEE-754 Double Precision

In short: Variant A: 132 of 1,200 converted (11%). Variant B: 174 of 1,200 (14.5%). The pooled rate is 306/2,400 = 0.127500 and the test says z = 2.570430, two-sided p = 0.010157 — the difference clears the 5% bar. The lift is 31.818182%, with a 95% interval of (6.678334%, 62.882494%) built on the log scale; the absolute difference is 3.5 points with interval (0.008349, 0.061651), which sits entirely above zero. Read the pair: the test says the difference is real, the intervals say it could still be anywhere from a whisper to a shout — ship it, and keep the interval, not just the trophy.

Formula

z = (p̂₂−p̂₁)/√(p̂(1−p̂)(1/n₁+1/n₂)) · CIₑₓₚ = (p̂₂−p̂₁) ± z*·√(p̂₁q̂₁/n₁ + p̂₂q̂₂/n₂)

The test pools both arms under the no-difference hypothesis; the interval unpools, because once the difference is the quantity being estimated each arm carries its own variance. The lift interval lives on the log scale — ratios are skewed — using the same SE the risk-ratio page derives: √(q̂₁/x₁ + q̂₂/x₂).

Worked Example

  1. Enter visitors and conversions for both variants — whole counts, one arm each.
  2. Read the test: z and its two-sided p against your alpha (5% is the convention, typed in your head).
  3. Read BOTH intervals before celebrating: the difference interval for size, the lift interval for the headline.
  4. If the rates tie, the page says so honestly — no detected difference is a finding, not a failure.

Defaults: 11% vs 14.5% — z = 2.570430, p = 0.010157, difference interval (0.008349, 0.061651), lift 31.818182% with interval (6.678334%, 62.882494%). A tie drive (132 of 1,200 both arms) prints the honest no-difference card.

Strengths & Limits Of This Model

Where this engine is strong

  • Lift interval on the log scale beside the difference interval
  • Tie printed as a finding with the power pointer

Where it stops

  • No sequential/peeking correction
  • No Bayesian posterior variant

Risk & accuracy notice. The A/B test is the most powerful ritual in the modern toolbox and the most commonly faked: peeking, stopping on significance, and testing twenty variants against the lucky one all manufacture wins from noise. This page prices ONE honest test — the discipline around it is yours, and the interval cards are there precisely so a bare p-value cannot end the conversation.

Practical Use Cases

Product

landing-page and pricing variants

Marketing

subject lines, bids, creative

Ops

two processes on two weeks of counts

Methodology & Editorial Standards

Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.

This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.

Sana Khalid Principal Front-End Engineer · ApexConverter

Statistical inference, experiment design and numerical stability. Last reviewed: 11 August 2026.

Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.


A/B Test Calculator — 8 Expert FAQs

8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.

Why does the test pool the proportions but the interval does not?

Different jobs. Under the no-difference hypothesis both arms estimate ONE shared rate, so the pooled p̂ gives the test its sharpest standard error. The interval makes no such hypothesis — it estimates the difference itself, so each arm keeps its own variance. Same data, two honest arithmetic.

The p-value is 0.010157 — how big is the win?

That is exactly the wrong question, and this page refuses to let it stay unasked: significance is about whether, the interval is about how much. Here the lift interval runs 6.678334% to 62.882494% — real, and still wildly loose. Quote both or quote neither.

Why is the lift interval built on the log scale?

Ratios are skewed: a lift can fall to −100% below but rise unbounded above, so a symmetric normal band around the raw ratio lies by construction. Logging bends the scale symmetric, the normal approximation goes back to being honest, and the result is exponentiated home — the same move the odds-ratio and risk-ratio pages make.

How many visitors do I need before the test?

More than the estimate needs — detecting a difference is harder than describing one proportion, which is why this n dwarfs the sample-size page's. The power-analysis page runs that question exactly: baseline, target effect, and alpha in; n per arm out. For these defaults it prices about 1,425 per arm for 80% power.

Can I peek at the test daily and stop when p dips under 0.05?

That is the peeking trap: repeated testing at a 5% bar is not a 5% procedure — the false positive rate balloons with every look. Fix the horizon before starting, or use a sequential design built for peeking. This page prices ONE test on ONE finished sample.

The rates tie — is the test broken?

No: identical rates give z = 0 and p = 1, the honest answer that this sample detected nothing. The card prints that as a finding — with the reminder that absence of evidence at a small n is a power statement, not proof of equivalence: the power-analysis page says how deaf your test was.

Why refuse an arm with zero visitors?

A proportion over nobody has no denominator, and the pooled SE collapses with it. One visitor is the honest floor; below that the page refuses rather than print z from an empty arm.

Is a significant lift always worth shipping?

The test prices randomness, not value: a significant 0.008 lift may cost more to implement than it returns, and a non-significant 3-point gap may merit a bigger rerun. The verdict card hands you the statistics; the shipping decision also owns cost, risk, and reversibility — numbers this page deliberately does not invent.

Related Statistics Engines