Statistics

Power Analysis Calculator

The pre-flight check for a two-variant test: given a baseline, a target rate and your alpha, how much power does your n actually buy — and how many per arm would buy 80%.

Power Analysis Calculator

Results recalculate instantly on every keystroke. Nothing you type is transmitted.

The effect
The design
Power at your n
—
n per arm for 80% power—
The read—
Power, honestly—

What this result does not account for

  • Two-proportion normal-approximation power — no continuous-outcome (t-test) power here
  • Equal allocation assumed (n per arm)
● Zero-Server Execution Updated 11 Aug 2026 Reviewed by Sana Khalid IEEE-754 Double Precision

In short: Baseline 11%, target 14.5% (the canonical 3.5-point lift), α = 5% two-sided, n = 1,200 per arm: the test's power is 0.729502 — under the conventional 80% floor, so even a real 3.5-point effect escapes detection better than one time in four. The same arithmetic run backwards prices the fix: 1,425 per arm buys 80% power (zβ = 0.841621 at the 0.8 convention). The underline: the sample-size page's n estimates ONE proportion cheaply; DETECTING a difference needs more — the promise the ssz page makes about A/B testing, priced here to the digit.

Formula

power = Φ((√n·|p₂−p₁| − zα/₂·√((p₁+p₂)(1−p̄))/√(p₁q₁+p₂q₂)) · n = (zα/₂√(2p̄q̄) + zβ√(p₁q₁+p₂q₂))²/(p₂−p₁)²

Both z quantiles are computed by bisection on the erf-series normal curve — zα/₂ for your alpha and zβ for the 80% convention — never read from a table. The n formula is the closed form; the power formula is R's power.prop.test structure. One-sided tests swap zα/₂ for zα, buying power by promising away one direction.

Worked Example

  1. Enter the baseline rate and the target rate — the smallest effect you would hate to miss.
  2. Set alpha and the sidedness; 5% two-sided is the convention and the default.
  3. Read power at your n: under 80% means the test is more deaf than the convention tolerates.
  4. Read the n-for-80% card and budget the test properly BEFORE running it.

Defaults: 11% vs 14.5%, α = 5% two-sided, 1,200/arm — power 0.729502, n-for-80% = 1,425 per arm. Driving n = 1,425 gives power ≈ 0.800052; halving the effect to 12.75% sends the requirement soaring — effects price their own samples.

Strengths & Limits Of This Model

Where this engine is strong

  • Power and n-for-80% on one panel, z quantiles bisected live
  • The ssz page's A/B promise priced to the digit

Where it stops

  • No continuity-corrected variant
  • No unequal-allocation option

Risk & accuracy notice. Power is the sentence you buy BEFORE the experiment or forfeit after: an under-powered null is uninterpretable, and a power number computed post hoc from the observed effect is the p-value in a costume. Plan with the smallest effect that matters, run to the n, and let this page's arithmetic be the reason — not the rationalization.

Practical Use Cases

Experiment planning

the n before the test, not after

Post-mortems

how deaf a null result actually was

Teaching

type II error made concrete

Methodology & Editorial Standards

Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.

This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.

Sana Khalid Principal Front-End Engineer · ApexConverter

Statistical inference, experiment design and numerical stability. Last reviewed: 11 August 2026.

Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.


Power Analysis Calculator — 8 Expert FAQs

8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.

What does “power 0.729502” actually promise?

That IF the true rates are 11% and 14.5%, and you run this exact test at α = 5%, you detect the difference about 73% of the time — and miss it (a false negative, beta) about 27% of the time. Power is a property of the DESIGN plus a claimed effect size: it is a pre-flight number, and the honest way to use it is before the test, not as an autopsy weapon after.

Why is the 80% line the convention?

It sets beta = 0.20 against alpha = 0.05 — a 4-to-1 asymmetry reflecting how much more society historically feared false positives than false negatives. It is a convention, and the page prices any level you like: type in the effect you truly care about and let the numbers argue with the tradition.

Why does detecting a difference need more n than estimating one proportion?

Estimating one rate at ±3% needs about a thousand; detecting a 3.5-point GAP between two rates needs about 1,425 PER ARM. The difference is the burden of proof: a difference must beat the noise of TWO arms, and the test pays for certainty it does not use. The sample-size page's FAQ promises exactly this; here is the invoice.

Can I run the test with low power anyway?

You can; know what you bought. An under-powered test that comes back null tells you almost nothing — “no detected difference” at 73% deafness is weak evidence of anything. And the significant results from low-power hunts overstate effects, because only the lucky-large draws clear the bar. Power protects both of your errors.

What if the true effect is bigger than my target?

Power rises steeply with the effect: the same n that gives 73% at a 3.5-point lift gives far more at a 6-point one. This is why honest planning states the SMALLEST effect worth detecting — the test is priced to the effect you cannot afford to miss, and bigger effects ride free.

One-sided or two-sided — does it change power?

At the same alpha, one-sided moves the rejection threshold in (1.644854 vs 1.959964), buying power — by promising before the data that you will never claim the other direction. It is a real discount with a real contract attached; choosing it AFTER seeing the direction is the trick this page makes visible but cannot stop.

Why is post-hoc power (computed from the observed effect) frowned on?

Because once the test is run, the observed effect and the p-value carry the same information — a significant result always shows high post-hoc power, a null one low, so the number just restates the p-value with extra steps. Power's honest life is BEFORE the test: design the n, then run the study. This page prices the design question, which is the only question power answers cleanly.

Where do the z quantiles come from?

Computed, not recited: both zα/₂ and zβ are bisected from the erf-series normal curve at runtime — the same engine the confidence-interval page uses for z*. The printed 1.959964 and 0.841621 are outputs of that arithmetic, which is why typing a custom alpha reprices everything honestly.

Related Statistics Engines