Conversion Significance Calculator
One conversion rate against a benchmark: the one-proportion z-test that answers “is my rate actually above (or below) the number everyone quotes?” — with the interval that says how far it could still be.
Conversion Significance Calculator
Results recalculate instantly on every keystroke. Nothing you type is transmitted.
What this result does not account for
- One rate, one benchmark — two-variant questions belong to the A/B page
- Wald interval: boundary rates (0 or all) deserve the prp page's Wilson treatment
In short: 96 conversions on 2,400 visitors is a 4% rate against a 3.5% benchmark: z = (0.040000−0.035000)/√(0.035000×0.965000/2,400) = 1.332840, two-sided p = 0.182584 — the benchmark-beating is NOT established at the 5% bar, even though the point estimate sits above it. The 95% Wald interval for the true rate is (0.032160, 0.047840), and the benchmark 0.035000 lies INSIDE it: your data cannot yet distinguish this rate from the benchmark. That is the honest read — “above the benchmark” is a claim about noise, and 2,400 visitors do not yet buy it. A drive to 120 conversions (5%) takes z to 3.998519 and p below 0.0001: then, and only then, the claim clears.
Formula
z = (p̂−p₀)/√(p₀(1−p₀)/n) · CI = p̂ ± z*·√(p̂(1−p̂)/n)
The test's standard error uses the BENCHMARK p₀ — under the no-difference hypothesis the benchmark is the truth being tested against. The interval, which makes no such hypothesis, uses the observed p̂. One rate against a fixed number is a different question from two rates against each other — that is the A/B page's job.
Worked Example
- Enter visitors and conversions for the ONE thing you measured.
- Enter the benchmark and pick the claim — two-sided is the conservative default; one-sided halves the p-value but bets the question in advance.
- Read the verdict with the interval: is the benchmark inside the span your data can honestly defend?
- A “not yet” is a power statement — more visitors shrink the interval until the benchmark either falls inside forever or out.
Defaults: 4% vs 3.5% — z = 1.332840, two-sided p = 0.182584, Wald interval (0.032160, 0.047840) containing the benchmark: not established. Driving x = 120: z = 3.998519, p below 0.0001 — established, both sides.
Strengths & Limits Of This Model
Where this engine is strong
- Benchmark-SE test and rate-SE interval kept consistent with their own logics
- One-sided vs two-sided made an explicit, named choice
Where it stops
- No sequential monitoring correction
- No exact binomial tail at tiny counts
Practical Use Cases
Marketing
campaign rate vs industry benchmark
Product
checkout rate vs last quarter's bar
QA
error rate vs an agreed threshold
Methodology & Editorial Standards
Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.
This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.
Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.
Conversion Significance Calculator — 8 Expert FAQs
8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.
This vs the A/B page — which do I need?
Count your rates. ONE rate against a FIXED number (industry benchmark, last quarter, an agreed threshold) is this page: the SE leans on the benchmark p₀. TWO variants measured side-by-side are the A/B page: pooled, two arms, lift. Testing your 4% against a 3.5% benchmark here and calling it an A/B win is the category error the pair of pages exists to prevent.
The rate is above the benchmark — why isn't that significant?
Because the claim is not about your SAMPLE, it is about the process that generated it, and 2,400 visitors carry a ±0.00784-point wobble — wider than the 0.005 gap you are claiming. The interval containing the benchmark IS the honest sentence: with this much data, the true rate and the benchmark are still mutually consistent.
One-sided vs two-sided — why does the choice matter?
One-sided spends all 5% of the error budget on ONE direction, so it clears faster (p halved) but promises in advance you will never claim the other direction with this data. Two-sided hedges. Pick BEFORE looking at the result — choosing “beats it” after seeing a positive z is exactly the half-the-p-value trick this page exists to make visible.
Why does the test use the benchmark in the SE but the interval uses my rate?
The test asks: IF the benchmark were the truth, how weird is my sample? So its SE lives at p₀. The interval makes no such assumption and describes where the true rate plausibly sits given YOUR p̂. Mixing the two — benchmark-SE intervals — is a classic quiet inconsistency this page refuses to ship.
My conversions are zero — what does the page say?
It still runs, because 0 conversions is real information: the rate is below the benchmark and the test says by how much, with the interval refusing to pretend precision at the boundary. What it will not do is dress the zero up as a small number — at n = 2,400 an observed zero is a strong statement about the process.
Why a Wald interval and not Wilson here?
The population-proportion page owns the deeper interval story (Wald beside Wilson, with the boundary collapses named). This page prints the plain Wald because the verdict — is the benchmark inside the span — matches the z test on the same SE at moderate rates. Near 0% or 100%, take the question to the prp page.
What does “not established” let me claim?
Exactly this: at this sample size, the data do not distinguish your rate from the benchmark. It does NOT license “the same as” or “no difference” — that would be claiming the null, and only an equivalence test (a different design) can argue for sameness. The interval says everything claimable; the verdict card respects it.
How many more visitors until it's significant?
Roughly: the SE shrinks with √n, so quadrupling the sample halves the wobble — and the gap to clear here is 0.005. The power-analysis page prices it exactly: baseline, target rate, alpha, and the n per arm falls out. The pattern to respect: decide the n BEFORE the campaign, not after the first dip.