Power Analysis Calculator
The pre-flight check for a two-variant test: given a baseline, a target rate and your alpha, how much power does your n actually buy — and how many per arm would buy 80%.
Power Analysis Calculator
Results recalculate instantly on every keystroke. Nothing you type is transmitted.
What this result does not account for
- Two-proportion normal-approximation power — no continuous-outcome (t-test) power here
- Equal allocation assumed (n per arm)
In short: Baseline 11%, target 14.5% (the canonical 3.5-point lift), α = 5% two-sided, n = 1,200 per arm: the test's power is 0.729502 — under the conventional 80% floor, so even a real 3.5-point effect escapes detection better than one time in four. The same arithmetic run backwards prices the fix: 1,425 per arm buys 80% power (zβ = 0.841621 at the 0.8 convention). The underline: the sample-size page's n estimates ONE proportion cheaply; DETECTING a difference needs more — the promise the ssz page makes about A/B testing, priced here to the digit.
Formula
power = Φ((√n·|p₂−p₁| − zα/₂·√((p₁+p₂)(1−p̄))/√(p₁q₁+p₂q₂)) · n = (zα/₂√(2p̄q̄) + zβ√(p₁q₁+p₂q₂))²/(p₂−p₁)²
Both z quantiles are computed by bisection on the erf-series normal curve — zα/₂ for your alpha and zβ for the 80% convention — never read from a table. The n formula is the closed form; the power formula is R's power.prop.test structure. One-sided tests swap zα/₂ for zα, buying power by promising away one direction.
Worked Example
- Enter the baseline rate and the target rate — the smallest effect you would hate to miss.
- Set alpha and the sidedness; 5% two-sided is the convention and the default.
- Read power at your n: under 80% means the test is more deaf than the convention tolerates.
- Read the n-for-80% card and budget the test properly BEFORE running it.
Defaults: 11% vs 14.5%, α = 5% two-sided, 1,200/arm — power 0.729502, n-for-80% = 1,425 per arm. Driving n = 1,425 gives power ≈ 0.800052; halving the effect to 12.75% sends the requirement soaring — effects price their own samples.
Strengths & Limits Of This Model
Where this engine is strong
- Power and n-for-80% on one panel, z quantiles bisected live
- The ssz page's A/B promise priced to the digit
Where it stops
- No continuity-corrected variant
- No unequal-allocation option
Practical Use Cases
Experiment planning
the n before the test, not after
Post-mortems
how deaf a null result actually was
Teaching
type II error made concrete
Methodology & Editorial Standards
Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.
This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.
Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.
Power Analysis Calculator — 8 Expert FAQs
8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.
What does “power 0.729502” actually promise?
That IF the true rates are 11% and 14.5%, and you run this exact test at α = 5%, you detect the difference about 73% of the time — and miss it (a false negative, beta) about 27% of the time. Power is a property of the DESIGN plus a claimed effect size: it is a pre-flight number, and the honest way to use it is before the test, not as an autopsy weapon after.
Why is the 80% line the convention?
It sets beta = 0.20 against alpha = 0.05 — a 4-to-1 asymmetry reflecting how much more society historically feared false positives than false negatives. It is a convention, and the page prices any level you like: type in the effect you truly care about and let the numbers argue with the tradition.
Why does detecting a difference need more n than estimating one proportion?
Estimating one rate at ±3% needs about a thousand; detecting a 3.5-point GAP between two rates needs about 1,425 PER ARM. The difference is the burden of proof: a difference must beat the noise of TWO arms, and the test pays for certainty it does not use. The sample-size page's FAQ promises exactly this; here is the invoice.
Can I run the test with low power anyway?
You can; know what you bought. An under-powered test that comes back null tells you almost nothing — “no detected difference” at 73% deafness is weak evidence of anything. And the significant results from low-power hunts overstate effects, because only the lucky-large draws clear the bar. Power protects both of your errors.
What if the true effect is bigger than my target?
Power rises steeply with the effect: the same n that gives 73% at a 3.5-point lift gives far more at a 6-point one. This is why honest planning states the SMALLEST effect worth detecting — the test is priced to the effect you cannot afford to miss, and bigger effects ride free.
One-sided or two-sided — does it change power?
At the same alpha, one-sided moves the rejection threshold in (1.644854 vs 1.959964), buying power — by promising before the data that you will never claim the other direction. It is a real discount with a real contract attached; choosing it AFTER seeing the direction is the trick this page makes visible but cannot stop.
Why is post-hoc power (computed from the observed effect) frowned on?
Because once the test is run, the observed effect and the p-value carry the same information — a significant result always shows high post-hoc power, a null one low, so the number just restates the p-value with extra steps. Power's honest life is BEFORE the test: design the n, then run the study. This page prices the design question, which is the only question power answers cleanly.
Where do the z quantiles come from?
Computed, not recited: both zα/₂ and zβ are bisected from the erf-series normal curve at runtime — the same engine the confidence-interval page uses for z*. The printed 1.959964 and 0.841621 are outputs of that arithmetic, which is why typing a custom alpha reprices everything honestly.